Why your scanned PDF is 40 MB (diagnosis walkthrough)

By Filemynt Editorial · Last updated July 22, 2026

A step-by-step way to find out exactly why one scanned document ballooned to 40 megabytes, before you reach for a compressor.

Start with a question, not a compressor

It is tempting to see '40 MB' and immediately hit Maximum compression. That skips the diagnosis step, and diagnosis usually tells you something worth knowing: whether the size is a real problem, a fixable habit, or both.

This walkthrough treats the file like a patient, not a nuisance. Five checks, done in order, will tell you almost exactly where the bytes went — before you decide what to do about it.

Check one: how many pages, and what kind

Open the file and look at the page count next to the size. A 3-page document at 40 MB is a completely different problem from a 300-page document at 40 MB. Divide megabytes by pages: over roughly 1 MB per page usually points to image-heavy content; well under that usually means the file is mostly text.

Scroll through the thumbnails. If most pages look like photographs of paper — visible paper texture, slight skew, uneven lighting — you are looking at a scan. Scanners turn every page into a picture, even a page that is entirely typed text, unless OCR has added a separate text layer.

Check two: resolution and color depth

Scanning software often defaults to settings meant for photographs, not documents: high DPI and full color, even for a black-and-white letter. A one-page text letter scanned at 600 DPI in full color can easily be ten times larger than the same page scanned at 300 DPI in grayscale, with no visible difference in readability.

If you control the scanner or scanning app, check its default profile. Many offer a 'document' or 'text' preset that already targets grayscale and a document-appropriate resolution. Switching the default fixes this at the source instead of compressing every file after the fact.

Check three: blank backs and duplicate pages

Automatic document feeders frequently scan both sides of every sheet, even single-sided ones, producing a blank image page for every real page. Those blanks are still full-resolution images and still cost bytes, even though they show nothing.

Flip through the whole document once, not just the first few pages. It is common to find five or ten blank backs, an accidental duplicate scan of the same sheet, or a cover page scanned twice. Removing those with Remove PDF Pages is free size reduction with zero quality loss — do it before you compress anything.

Check four: was this actually several scans merged together?

If the file grew over several sessions — you scanned page one Monday, added exhibits Wednesday, appended a signature page Friday — different scanning settings may have been used each time. One badly-configured session (accidentally left at high DPI and full color) can account for most of the total weight even if it is a small fraction of the pages.

Spot-check file size contribution by eye: pages that look noticeably crisper or more saturated than the rest were probably captured differently, and are worth a closer look.

Check five: is this actually a photo stack disguised as a PDF

Some 'scans' are really a stack of full-resolution phone camera photos assembled into a PDF with JPG to PDF or similar. Phone cameras often capture at far higher resolution than a document needs, and shadows or glare add visual noise that compresses poorly.

If that is the source, the fix starts at capture: flatten the page, use even lighting, and let a scanning app crop tightly to the page edges before assembly, rather than compressing away noise after the fact.

Putting the diagnosis into action

Junk pages found? Remove them first. Full color where grayscale would do, and the source is still available? Rescan with a document preset if that is realistic. Neither of those apply, and the file is simply what it is? Now compression is the right next step — start with Balanced, not Maximum, and check a signature and a paragraph of small print before you trust the result.

The full mechanics of what a compressor actually changes are covered in Why PDFs become large. Once you have removed the junk and picked a sensible starting point, run Compress PDF and compare before/after sizes with your own eyes rather than trusting a percentage alone.

When 40 MB is simply correct

Not every large scan is a mistake. A 150-page exhibit binder scanned at a reasonable resolution can legitimately be tens of megabytes, and crushing it to fit an arbitrary target can make signatures and stamps unreadable. If the size reflects genuine content rather than blank pages or an overcooked scanner setting, the honest fix is splitting the file or using a link for delivery, not destroying quality to hit a number.

Related tools

Keep reading

Next steps