Research notes

Turning a web page into Markdown: what survives the trip

A 488 KB Wikipedia page through our HTML to Markdown tool, with every link, picture and table counted, and the table bug it found.

Updated 24 September 2026

We converted Wikipedia’s article on the PDF format to Markdown and the fact box at the top of the page, the one listing the file extension and media types, fell apart. It came out as a table with a one-column header over two-column rows. Shown as Markdown, the table stopped after three rows, and “.pdf” and every media type were gone from it.

That was our own HTML to Markdown tool, on 24 September 2026. We fixed it the same day, and the test that found it is a fair picture of what survives when you turn a web page, or formatted text you copied, into Markdown.

Why tables break

Markdown tables are simple: one line per row, a pipe between cells, every row with the same number of cells. Web tables are not. The Wikipedia fact box has a title cell stretched across both columns and cells holding bulleted lists. The converter wrote the stretched title as a single cell, so the header had one column, and the GitHub Markdown rules drop any cell beyond the header’s count. The lists were written over several lines with blank lines between them, and a blank line ends a Markdown table on the spot.

The tool now gives every row one cell per column, leaving empty cells where a web table merged them, and turns line breaks inside a cell into <br>, which GitHub and most Markdown viewers show as a new line. After the fix, all ten tables on the page came through: six as Markdown tables, four left as HTML because they have no header row, which the Markdown rules allow. The four are the boxes of related-article links and “citation needed” banners you probably didn’t want anyway.

The rest of the page, counted

Wikipedia's 'PDF' article (HTML from the Wikipedia API) through WipeTheAI HTML to Markdown in Chromium 153, 24 September 2026, then shown again as HTML to count what came back (n=1 page)
WhatIn the pageAfter MarkdownNotes
Size488 KB of HTML151 KB of MarkdownConverted in 55 ms
Words11,12210,663 (96%)We didn’t trace where the other 4% went
Links1,5261,527Every link kept
Pictures88Kept as links to Wikimedia’s servers; 4 of the 6 in the text had no description
Tables10106 as Markdown, 4 as HTML
Citation markers118118Each one a link like [1] in the text

Pictures need a warning. Markdown can’t hold an image, only a link to one, so the converted file still loads every picture from Wikimedia. Take it offline and the pictures go. Four of them had no description at all, so the Markdown shows ![] followed by an address: screen readers and plain-text exports get nothing. The page wrote those addresses without https: at the front, which works in a browser and points nowhere in a file opened from disk, so the tool now adds it.

The citations are the clutter. All 118 of them come across as links, [\[1\]](./PDF#cite_note-4) in the raw text, spread over 65 lines. If you want the article as notes or as text for an AI tool to read, select just the paragraphs you need in your browser, copy and paste them into the tool, and skip the fact box, the references and the link boxes. You get cleaner Markdown than a whole page ever gives.

Going the other way, Markdown to HTML and Markdown to PDF use the same GitHub rules, so a table that looks right in one of our tools looks the same in the other. For spreadsheet data, CSV to Markdown table writes tables that line up in the raw text too.

Sources

  1. Wikipedia: PDF (the test article)
  2. GitHub Flavored Markdown spec: tables
  3. Turndown, the HTML to Markdown converter our tool uses