Research notes

Curly quotes, em dashes and ellipses: when plain text needs them gone

We ran curly quotes, long dashes, odd spaces and zero-width characters through Python, Node, JSON, CSV and the SMS alphabet, and counted them in six public-domain books.

Updated 24 September 2026

Paste a sentence with a curly apostrophe into an old form or an email template, and the reader can end up with’ where the apostrophe was. That is the three bytes of one character, ’, read one byte at a time as Windows-1252. We generated every common case on 24 September 2026 in Chromium 153 and they all start the same way, with â€.

Our view is simple. Curly quotes, long dashes and the ellipsis belong in text people read on a page: a book, an article, a printed letter. They should come out of text a program reads: code, JSON, CSV, database fields, text messages and anything bound for a system that only knows plain ASCII. Below is what we measured for each of those, and what our straight quotes tool does about it.

What the garbled versions look like

Each character's UTF-8 bytes, decoded as Windows-1252 with TextDecoder in Chromium 153, 24 September 2026
CharacterUTF-8 bytesShows up asOur tool returns
’ right single quote (apostrophe)E2 80 99’'
“ left double quoteE2 80 9C“"
” right double quoteE2 80 9D†plus an invisible control character"
— em dashE2 80 94—-
– en dashE2 80 93–-
… ellipsisE2 80 A6…...
Non-breaking spaceC2 A0Â followed by a spacea space
Zero-width spaceE2 80 8B​ (the last one is ‹)removed
Byte order markEF BB BFremoved
éC3 A9éé (repaired, then kept)

The right double quote is the awkward one. Its last byte, 9D, has no character in Windows-1252, so decoders show a control character that takes no space on screen, and some programs drop it altogether. What the reader sees is†on its own. Our repair needs that last byte to be there: when it has been thrown away, the tool can no longer tell a lost ” from other damage, and it leaves †alone rather than guess.

The repair only acts where the letters line up into valid UTF-8. The words façade, naïve and Ärger went through our test suite untouched, and so did the pair ßü, whose bytes sit next to each other but don’t form a valid sequence. We also hit a snag worth knowing about if you script this yourself: in Node 20.19, new TextDecoder('windows-1252') turned bytes 80 and 99 into invisible control characters instead of € and ™, the way ISO-8859-1 does. Chromium gave the right answer, so our test file maps the bytes by hand.

Code, JSON and CSV refuse them, or read them wrongly

This is where most of the damage happens, and it is rarely loud. We pasted the same characters into Python 3.14, Node 20.19 and Chromium 153 on 24 September 2026.

Curly quotes around a string are a syntax error in both languages. Python says invalid character '“' (U+201C), which at least names the problem. JSON is stricter still: the spec, RFC 8259, allows only the plain double quote around strings and only space, tab, line feed and carriage return between values. So JSON.parse rejected curly quotes, a non-breaking space after a colon, a zero-width space and a byte order mark at the start. Python’s json.loads rejected the curly quotes and the non-breaking space too.

The quiet failures are worse. In JavaScript, a non-breaking space between tokens is legal, so code pasted from a web page runs, and then the same space inside a JSON file breaks. A zero-width joiner is allowed inside a JavaScript name: we declared total and total followed by a joiner, got two separate variables holding 1 and 2, and no warning at all. Tidying text doesn’t help either. JavaScript’s trim() removed a non-breaking space and a byte order mark but kept a zero-width space, and Python’s strip() kept it too. Number('42' + zero-width space) came back NaN, and Python’s float() raised an error on the same string.

CSV breaks by miscounting. RFC 4180 wraps a field that holds a comma in plain double quotes. Python’s csv module read Ana,"Lead, Design",2019 as three fields, as it should. With curly quotes, the same row came back as four: “Lead and Design” became two columns, and every column after them moved one place to the right.

One apostrophe can cost a text message

SMS has its own alphabet. Under 3GPP TS 23.038, a message written in the GSM 7-bit default alphabet can hold up to 160 characters, and one written in UCS2 holds 140 bytes, which is 70 characters. That alphabet has the plain quote and apostrophe, and a hyphen. It has no curly quotes, no en or em dash, no ellipsis, no bullet and no non-breaking space; we checked each against the default alphabet and its extension table in version 19.0.0 of the spec.

So a 125-character booking reminder with it’s in it can’t be sent as GSM 7-bit, and at 125 characters it is well over the 70 that UCS2 allows in one message. With a straight apostrophe the same text fits in one. We worked this out from the spec rather than sending messages through carriers, and some phones and messaging services swap characters on their own before sending, so treat it as the rule the network starts from.

How many there are in real books

Characters our tool changed in six Project Gutenberg books (text between Gutenberg's start and end markers), counted with its own code in Node 20.19, 24 September 2026. Bytes are the UTF-8 size of that text
BookWordsCurly “ ”Curly ‘ ’Em dashesBytes beforeBytes after
Pride and Prejudice127,3603,8067850752,487743,305
Moby-Dick212,7963,0942,9361,7281,256,4531,244,194
The Adventures of Sherlock Holmes104,5065,0891,485191587,736574,545
Frankenstein75,042773187124428,903426,943
The Great Gatsby48,2082,9111,371417286,681277,955
Alice's Adventures in Wonderland26,5252,231753263154,483148,357

Every book used curly quotes throughout. The dashes depend on who transcribed it: the Gutenberg Pride and Prejudice has no em dashes at all, but 498 double hyphens, the old typewriter way of writing one. The Great Gatsby was the only one with ellipsis characters (54) and with hair spaces, 12 of them, each sitting between nested quote marks like ’ ” so they don’t touch. Our tool turned those into a normal space.

Frankenstein has 480 opening double quotes and only 293 closing ones. Much of it is told in letters and long speeches, and a quotation that runs over several paragraphs opens each one with a quote mark and closes only the last, which would account for the gap. We didn’t check all 187 unmatched quotes one by one.

Plain text is smaller too, since a curly quote takes 3 bytes in UTF-8 and a straight one takes 1. Alice came out 4.0% smaller. For a whole book that hardly matters; for a database column with a byte limit, it can decide whether a row fits. The largest book, Moby-Dick at 1.24 million characters, took 56 ms.

What the tool leaves in

Accented letters and symbols stay, and the tool lists them under the result. Moby-Dick kept 23 æ, 4 £ and a short Hebrew word with the invisible direction marks around it; removing those marks can scramble right-to-left text, so the tool keeps them in any text that has right-to-left letters. It also keeps a zero-width joiner inside an emoji, where removing it splits a family emoji into three separate faces, and inside words in Persian, Hindi and other scripts that need it to shape the letters.

One more source of odd spaces: number formatting. In Chromium 153, formatting 1234567.5 for French gave 1 234 567,5 with a narrow non-breaking space (U+202F) between the groups. Copy that number into a spreadsheet or a parser expecting a plain space, and it may not be read as a number at all.

If the text is Markdown, clean it first and then use Markdown to plain text to drop the formatting marks, or CSV to Markdown table once the quotes in a CSV are straight again.

Sources

  1. ETSI TS 123 038 V19.0.0 (3GPP TS 23.038): SMS alphabets and message lengths
  2. RFC 8259: The JSON data format (whitespace and quotation mark)
  3. RFC 4180: The CSV format
  4. Project Gutenberg: Pride and Prejudice (ebook 1342)
  5. Project Gutenberg: Moby-Dick (ebook 2701)
  6. Project Gutenberg: The Adventures of Sherlock Holmes (ebook 1661)
  7. Project Gutenberg: Frankenstein (ebook 84)
  8. Project Gutenberg: The Great Gatsby (ebook 64317)
  9. Project Gutenberg: Alice's Adventures in Wonderland (ebook 11)