Copy the positional-encoding formula out of the “Attention Is All You Need” PDF and you get P E(pos,2i) = sin(pos/100002i/dmodel ). The original says PE, with 10000 raised to the power 2i/dmodel. The exponent has dropped onto the line and fused with the number, the subscript has lost its place, and PE has split in two. Nothing went wrong in the copy. That is what the PDF holds.
A PDF stores where each run of characters sits on the page, not which words form a sentence or which numbers are exponents. Any tool that pulls text out, ours included, has to guess the structure back from positions. We ran eight PDFs through our PDF to text tool on 24 September 2026 to see which guesses fail, and how often.
What came out, file by file
A scan gave nothing
Three pages we had turned into 300 dpi pictures, the way a scanner app saves them, produced zero characters. A scanned page is a photo of text, and reading it needs OCR, which PDF to text doesn’t do. If a PDF won’t let you select a word in your reader, this is why.
Our Image to text tool does: drop the scanned PDF in and it reads each page as a picture. A PDF of two scanned pages from an 1892 book came back with 4 and 2 characters in 1,000 wrong, at just over a second a page. Receipts, faded print and two pages scanned side by side do far worse, as our OCR accuracy tests show.
Maths lost its exponents
In the Transformer paper, every superscript came down to the baseline: the complexity table gives O(n2 · d) for n squared times d, and O(k · n · d2) for the convolutional layer. You can repair a few by hand if you know what they meant. You can’t repair a page of them. For a paper on arXiv, download the TeX source instead; the formulas are written out there in full.
Tables turned into lines of words
The same table came out as rows of words separated by single spaces, with a two-line header split so that “Operations”, the second line of one column title, landed on a line of its own under the whole header:
Layer Type Complexity per Layer Sequential Maximum Path Length Operations Self-Attention O(n2 · d) O(1) O(1) Recurrent O(n · d2) O(n) O(n)
Where a cell holds more than one word, as in “Self-Attention (restricted)”, nothing tells you where it ends. Our tool doesn’t try to rebuild tables; for a table you need in a spreadsheet, copying it column by column from the PDF reader is slower but right.
Hyphenated words stayed split
The NIST AI Risk Management Framework is typeset with hyphenation, and 238 of its 1,710 extracted lines end in a hyphen: “mea-” on one line and “sure” on the next. Not every one is a split word, since some compound words really do break at a hyphen, so a blind search-and-replace joins those wrongly. Page numbers also came out glued to the last sentence of the page, as in “benefits for people, organizations, and ecosystems. 5”.
The two-column form was fine
The IRS W-9 prints its instructions in two columns, and the text came out one column after the other, lists and numbered items intact. Our tool reads text in the order the PDF stores it, and the software that made the form stored it in reading order. A two-column PDF from other software can store it line by line across both columns, and then the sentences interleave. We didn’t have such a file in this test.
How long it takes
Not long. The 226-page C2PA specification gave 458,580 characters in under a fifth of a second on a laptop, and the whole test set took less than half a second. Speed was never the problem in our test; the structure was.
Which PDFs are safe to extract
Plain reports and letters, like our Word-made NIST report and the memo we printed from a browser, came out well: paragraphs in order, lists intact, a stray page number to delete. Forms like the W-9 came out well too. Expect to clean up anything with maths, tables or hyphenation, and expect nothing at all from scans. When you want the pages to look exactly as they do, saving them as images is the other route; the text just stops being text.