Research notes

Do images come out of a PDF at full quality?

We put three photos into PDFs six different ways, took them out again and compared the files byte by byte.

Updated 24 September 2026

We put a 13.8 MB photo of the Golden Gate Bridge into a PDF with four different tools, pulled it back out, and got the same file four times: same size, same SHA-256 hash, and the same GPS position the camera wrote, 37°48′31.82″ N, 122°28′30.90″ W. The PDF had carried the whole JPEG, location included, untouched. A fifth tool gave back a different photo a third the size. Whether a picture comes out of a PDF at full quality depends almost entirely on the app that made the PDF, not on the tool you use to get it out.

Six ways to make a PDF, three photos, one question

On 24 September 2026 we took three public JPEGs from Wikimedia Commons: the Golden Gate photo (5388 × 3368, a Canon EOS 7D Mark II, with GPS), Hopetoun Falls (3072 × 2048, a Canon EOS 10D) and the Eiffel Tower (2900 × 5367, no camera details at all). Each went into a PDF six ways: our own JPG to PDF, the pdf-lib 1.17.1 library called directly, Chromium 153 printing a page with the photo on it, Python’s reportlab 5.0.1, img2pdf 0.6.3, and macOS’s sips command, which uses the system’s own PDF writer. Then we took the pictures out with our Extract images from PDF tool and compared them byte by byte with the files that went in.

The Golden Gate JPEG (13.8 MB) after a trip through a PDF, extracted with WipeTheAI in Chromium 153, 24 September 2026. The other two photos gave the same pattern; the Eiffel Tower photo had no details to remove (n=3 photos, 18 PDFs)
Made withPicture that came outCamera details and GPSPDF size
pdf-lib 1.17.1Identical fileKept13.8 MB
Chromium 153, print to PDFIdentical fileKept13.9 MB
img2pdf 0.6.3Identical fileKept13.9 MB
reportlab 5.0.1Identical fileKept17.3 MB
WipeTheAI JPG to PDFIdentical picture data, details removedRemoved (colour profile kept)13.8 MB
macOS sipsSaved again, 4.70 MBRemoved4.71 MB

Four of the six generators copy the JPEG into the PDF as it is, and it comes out exactly as it went in. That includes the camera model, the lens (a Canon EF 24-70mm f/2.8L II USM) and the GPS position. Our own JPG to PDF keeps every byte of the compressed picture but takes out the EXIF, IPTC and XMP blocks first, so the photo comes out looking identical with no camera details or location. macOS’s writer is the odd one out: it decoded each photo and compressed it again, so what came out was a new, smaller JPEG with none of the camera details. The Golden Gate photo shrank from 13.8 MB to 4.70 MB, and the Eiffel Tower, which had no details to lose, from 5.28 MB to 4.57 MB.

reportlab surprised us. Its PDF was 25% bigger than the photo inside it, 17.3 MB for 13.8 MB, because it stores the JPEG wrapped in ASCII85, a text encoding that makes binary data safe to print but costs a quarter more space. The picture inside is still byte for byte the original.

The details you can’t see from outside

Running exiftool on the PDF made by pdf-lib reported no camera, no lens and no GPS. Those details were there: a plain text search of the PDF found “Canon EOS 7D Mark II” in it, and exiftool on the extracted JPEG showed the full location. Checking a PDF’s own metadata doesn’t check the photos inside it.

Public PDFs carry the same kind of passengers. The NIST AI Risk Management Framework, typeset with LaTeX, holds nine pictures: two JPEGs and seven compressed bitmaps that came out as PNG. Its JPEGs still carry XMP naming the apps that made them, Adobe Illustrator 26.0 on Windows for one and Photoshop CC 2017 on a Mac for the other. The same 379 × 379 “Check for updates” badge sits in both NIST reports we tested and came out with the same hash from each. Apart from that badge, the NIST Cybersecurity Framework, exported from Word, gave three JPEGs with no metadata at all and nine PNGs, which fits an app that saves pictures again on the way in, though we can’t see which step did it.

What “full quality” means for the ones that come out as PNG

A picture stored in a PDF as JPEG can be handed back byte for byte. Most other pictures are stored with lossless compression, the same kind PNG uses, so unpacking them and saving a PNG keeps every pixel: nothing is lost, even though the file isn’t the one someone originally dropped into their document. The arXiv copy of “Attention Is All You Need” and the 226-page C2PA specification hold only pictures of that kind, and all of them came out as PNG. See-through areas are kept, because a PDF stores them as a separate mask that our tool puts back as the PNG’s transparency.

Two cases lose something. A JPEG with a see-through mask comes out as the JPEG alone, since JPEG has no transparency; the NIST framework’s cover photo is one. And print files in CMYK colours come out converted to screen colours, approximately. Black-and-white scans stored in the fax formats JBIG2 or CCITT aren’t saved at all; PDF to PNG saves the whole page instead.

If you are about to send a PDF made of phone or camera photos, assume the location is inside unless you know the app strips it. Making the PDF with our JPG to PDF takes the details out while keeping the picture data byte for byte. In a PDF you already have, Remove PDF metadata clears the document’s own details but not the photos’; Compress PDF clears those too for every photo it saves again smaller.

Sources

  1. Wikimedia Commons: Golden Gate Bridge as seen from Battery East
  2. Wikimedia Commons: Hopetoun falls
  3. Wikimedia Commons: Tour Eiffel Wikimedia Commons
  4. NIST AI 100-1: AI Risk Management Framework
  5. NIST CSWP 29: The NIST Cybersecurity Framework (CSF) 2.0
  6. arXiv: Attention Is All You Need (1706.03762v7)
  7. C2PA Technical Specification 2.1
  8. pdf-lib
  9. reportlab
  10. img2pdf