We converted Wikipedia’s article on the PDF format to Markdown and the fact box at the top of the page, the one listing the file extension and media types, fell apart. It came out as a table with a one-column header over two-column rows. Shown as Markdown, the table stopped after three rows, and “.pdf” and every media type were gone from it.
That was our own HTML to Markdown tool, on 24 September 2026. We fixed it the same day, and the test that found it is a fair picture of what survives when you turn a web page, or formatted text you copied, into Markdown.
Why tables break
Markdown tables are simple: one line per row, a pipe between cells, every row with the same number of cells. Web tables are not. The Wikipedia fact box has a title cell stretched across both columns and cells holding bulleted lists. The converter wrote the stretched title as a single cell, so the header had one column, and the GitHub Markdown rules drop any cell beyond the header’s count. The lists were written over several lines with blank lines between them, and a blank line ends a Markdown table on the spot.
The tool now gives every row one cell per column, leaving empty cells where a web table merged them, and turns line breaks inside a cell into <br>, which GitHub and most Markdown viewers show as a new line. After the fix, all ten tables on the page came through: six as Markdown tables, four left as HTML because they have no header row, which the Markdown rules allow. The four are the boxes of related-article links and “citation needed” banners you probably didn’t want anyway.
The rest of the page, counted
| What | In the page | After Markdown | Notes |
|---|---|---|---|
| Size | 488 KB of HTML | 151 KB of Markdown | Converted in 55 ms |
| Words | 11,122 | 10,663 (96%) | We didn’t trace where the other 4% went |
| Links | 1,526 | 1,527 | Every link kept |
| Pictures | 8 | 8 | Kept as links to Wikimedia’s servers; 4 of the 6 in the text had no description |
| Tables | 10 | 10 | 6 as Markdown, 4 as HTML |
| Citation markers | 118 | 118 | Each one a link like [1] in the text |
Pictures need a warning. Markdown can’t hold an image, only a link to one, so the converted file still loads every picture from Wikimedia. Take it offline and the pictures go. Four of them had no description at all, so the Markdown shows ![] followed by an address: screen readers and plain-text exports get nothing. The page wrote those addresses without https: at the front, which works in a browser and points nowhere in a file opened from disk, so the tool now adds it.
The citations are the clutter. All 118 of them come across as links, [\[1\]](./PDF#cite_note-4) in the raw text, spread over 65 lines. If you want the article as notes or as text for an AI tool to read, select just the paragraphs you need in your browser, copy and paste them into the tool, and skip the fact box, the references and the link boxes. You get cleaner Markdown than a whole page ever gives.
Going the other way, Markdown to HTML and Markdown to PDF use the same GitHub rules, so a table that looks right in one of our tools looks the same in the other. For spreadsheet data, CSV to Markdown table writes tables that line up in the raw text too.