On 8 January 2019, Paul Manafort’s lawyers filed a ten-page response in federal court in Washington with whole paragraphs blacked out. Reporters copied the black bars into a text editor and read what was under them. We downloaded that filing from CourtListener’s archive of the public court record on 24 September 2026 and ran it through the same code our PDF to text tool uses. It gave back 325 words that the black bars were supposed to hide.
What came out of the Manafort filing
The filing is document 471 in United States v. Manafort, case 1:17-cr-00201 in the District of Columbia. The same day, the court record got a corrected version, document 472. Both are ten pages and look the same on screen, black bars and all. The difference is underneath.
| Page | Document 471 (first filing) | Document 472 (refiled) | Words only in 471 |
|---|---|---|---|
| 5 | 327 | 204 | 123 |
| 6 | 329 | 251 | 78 |
| 7 | 323 | 293 | 30 |
| 9 | 325 | 231 | 94 |
| All ten pages | 3,033 | 2,708 | 325 |
In document 471 the bars are 39 filled rectangles drawn in the page content, on top of text that is still there in full. A rectangle hides letters from your eyes and nothing else. Text extraction, search and copy work from the text itself and never look at what is painted over it. The recovered passages include the ones the Columbia Journalism Review reported at the time: that Manafort had discussed a Ukraine peace plan with his associate Konstantin Kilimnik more than once, and that he was accused of sharing 2016 campaign polling data with him. Our extraction also turned up a meeting between the two in Madrid. None of the words “polling”, “Madrid” or “peace plan” appear in the text of document 472; all three are in 471.
The file’s own details told a second story. Its title field reads “Attachment B - REDACTED Response to OSC Breach Submission.docx”: it was written in Word, turned into a PDF with Nuance PDF Create and then edited with iText, and its author field holds a staff member’s login name. Removing that is a separate job (see what your PDF says about you), but it is the same mistake: assuming that what you can’t see isn’t in the file.
The same mistake, made on purpose
We built a three-page test letter with an invented customer name, email address, card number and phone number, then drew black rectangles exactly over each of them, the way a rectangle or highlight tool in a PDF editor does. On screen, all four were gone. Our PDF to text code returned every one of them, on the first try. So did pypdf and pdfminer, two widely used Python libraries.
Then we redacted the original letter with our Redact PDF tool: a search for the name, then the Email address, Long number and Phone number buttons, and one box drawn by hand on page 2. Pages with a box on them were redrawn as pictures with the boxes painted in, and their text, fonts and drawing commands were thrown away. Afterwards pdf.js, pypdf and pdfminer all found zero characters of text on those pages, and a search through every byte of the file, with every compressed stream unpacked, found none of the four values. A bookmark titled with the customer’s name was renamed “Redacted”, because bookmarks, form fields and accessibility tags can repeat page text somewhere the page itself never shows.
What real redaction costs
A redacted page is a picture, so its text can no longer be selected, searched or read aloud by a screen reader, and it weighs more. Pages without a box keep their text: when we redacted page 1 of the 48-page NIST AI Risk Management Framework, the other 47 pages came out with exactly the same text as before.
| File | Original | Standard (150 dpi) | Sharp (200 dpi, default) | Print (300 dpi) |
|---|---|---|---|---|
| IRS Form W-9, 6 pages | 141 KB | 400 KB | 562 KB | 912 KB |
| Manafort filing 471, 10 pages | 90 KB | 247 KB | 327 KB | 515 KB |
| NIST AI RMF 1.0, 48 pages | 1.95 MB | 1.36 MB | 1.44 MB | 1.62 MB |
The NIST report got smaller because saving drops data the original carried but never used, which outweighed one page becoming a picture. Short documents grow: one redacted page of the W-9 added about 420 KB at the default setting. We think that is the right trade. A redaction that keeps the page small by keeping the text is not a redaction.
Search boxes need care too. Our tool places a box from where the PDF says each letter sits, then checks the drawn page and widens any box that cuts through a letter. That second step exists because of a measurement: on justified text in a NIST PDF, the positions pdf.js reports were 2.6 points away from where it drew the letters, enough to leave the edge of a “d” showing. We then checked 600 sample words from seven PDFs against the letter positions another library, pdfminer, reports. No box stopped more than 0.3 points short of its word’s first or last letter, except three words ending in a full stop, where the dot sat just outside. Look at the preview before you save anyway: it shows exactly what will be blacked out.
To black out something for real, use Redact PDF, then check the result with PDF to text: whatever it can read, anyone can. A password doesn’t help here, for reasons covered in what an owner password really stops.
Sources
- United States v. Manafort, 1:17-cr-00201 (D.D.C.), Document 471, filed 8 January 2019 (CourtListener RECAP copy)
- United States v. Manafort, Document 472, filed 8 January 2019 (CourtListener RECAP copy)
- Columbia Journalism Review: Thank you to everyone who can’t redact documents properly
- NIST AI 100-1: Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- IRS Form W-9 (Rev. March 2024)