Search SpacerrApps

Find an app or a write-up by title

All posts

MDReduce’s File-Cleaning Bet for Lower LLM Costs

A proposed web tool that removes document and image noise before AI processing, but is not yet a finished product.

Written by
SpacerrApps
Reviewed by
Spacerr Team
Published
Reading time
4 min read

A four-page PDF can contain the same footer four times. A screenshot can include a battery icon, browser controls, and other interface elements that have nothing to do with the question you want an AI to answer. If an LLM receives the whole file, that irrelevant material becomes part of the input.

That is the problem MDReduce is aimed at. It proposes cleaning documents and images before they reach an AI, with the goal of preserving useful meaning while removing material that adds processing weight. The important qualification is that MDReduce is not live yet. Its current page is collecting feedback about which use case should be built first.

The proposed job is file preparation

MDReduce is intended to handle PDFs, Word files, Excel files, PowerPoint files, images, and screenshots. The cleaning process described by its developer would remove repeated headers and footers, boilerplate, HTML tags, irrelevant pixels, and interface chrome. The output would be a cleaner file for an LLM to process.

That puts the product before the model in a workflow. Instead of sending an untouched PDF or screenshot directly to an AI, a user would first pass it through MDReduce, then send the resulting file or extracted content onward. In practical terms, it is a token reduction tool for people who want to reduce LLM token cost by changing the input rather than changing the model.

The intended users are fairly clear: people who pay for LLM API calls, teams processing documents in bulk, and anyone who regularly sends screenshots or scans to an AI. The proposed web format also means the product is not presented as a desktop utility or a local command-line tool. For users who need files to stay on their own machines, that distinction may matter once the product exists.

The numbers come from specific examples

MDReduce’s central claim is that it measures reductions on real files rather than offering a general estimate. The examples supplied are narrow but concrete.

A standard four-page PDF reportedly used 23% fewer tokens after repeated boilerplate was removed. The stated reason is straightforward: the same legal footer appeared on each page, so removing those repetitions reduced the amount of text sent to the model.

For PNG screenshots, the claimed reduction reaches up to 3.7 times. The description attributes that result to stripping user-interface elements and changing from a vision model to a text model for extraction.

Those figures are useful as illustrations of the intended approach, not as a guarantee for every file. One PDF and one screenshot do not establish a typical reduction across contracts, spreadsheets, scans, or slide decks. Results would depend on how much repeated text, interface material, or other noise a file contains. The product’s own wording recognises this by asking users which workflow matters most before committing to a full build.

It is still testing the problem

The most important fact about MDReduce is its stage of development. The developer says the product is not live and that the current page exists to collect demand rather than sell a finished service. Whichever use case attracts the strongest response is supposed to determine what gets built and shipped first.

That makes MDReduce better understood as a product proposal and validation effort than as a tool you can add to a production pipeline today. There is no basis yet for treating the stated reductions as independently verified performance, or for assuming that every listed file type will be supported in the first release.

The plan is also deliberately broad. PDFs, office documents, images, and screenshots can require different kinds of processing. Removing a repeated footer from a text-heavy PDF is a different task from deciding which pixels in a screenshot are irrelevant. Supporting all of these use cases well could require different workflows, even if the common goal is smaller, cleaner input.

Privacy is part of the proposal

The developer says the eventual product is planned without an account, email, or cookies. The current page is also described as having none of those requirements. That is a useful direction for a file-cleaning service, since users may be handling documents they do not want tied to a new account.

It is still only a plan for the product itself. The description does not explain how files would be handled, how long they would be retained, or whether processing would happen locally or on a remote service. Anyone dealing with confidential documents would need those details before relying on MDReduce.

Who should pay attention

MDReduce is aimed at anyone who pays per token for LLM API calls, processes documents in bulk, or sends screenshots and scans to AI systems regularly. Its strongest case is for workflows where files contain obvious repetition or visual clutter and where reducing input size has a direct cost benefit.

It is not for someone who needs a working file-processing service today, a confirmed reduction rate across varied documents, or clear data-handling guarantees already in place. For now, it is a focused proposal with two claimed real-file examples and an open question about which audience it should serve first. That makes feedback part of the product, not a finished feature.

MDReduce

Strip the noise. Keep the meaning.

Visit MDReduce