Rarely, and when they do, the character usually belongs there. We ran 220,375 ChatGPT answers from WildChat, a public research dataset of real conversations, through our hidden character remover on 24 September 2026. 6,320 of them (2.87%) held at least one invisible character or unusual space. Only 705 (0.32%) held one the remover takes out with its default settings. The rest were joiners inside Persian words and emoji, the kind of character that has to stay or the text breaks.
Our position: a stray invisible character in a chatbot answer is a formatting defect, not a fingerprint. Clean it for the same reasons you would clean any pasted text, because it trips up code, forms, search and word counts. Removing it tells nobody anything about who wrote the words, and it does nothing to the statistical watermarks some AI companies put into text. More on that below.
What we counted
WildChat was collected by researchers who ran a free chatbot on OpenAI’s models and published the transcripts of users who opted in. It is not the ChatGPT app or website, and it has no copy button in between: the records carry the model’s own text, with fields such as token usage that come from OpenAI’s service. We took 3 of the 86 files of the 4.8M release, one each from April 2023, October 2024 and July 2025, and counted every character in every turn with the remover’s own code.
| Model | Month | Answers | Any hidden or odd space | Removable | Most common removable |
|---|---|---|---|---|---|
| gpt-3.5-turbo-0301 | Apr 2023 | 85,127 | 490 | 95 (0.11%) | zero-width space, 47 answers |
| gpt-4-0314 | Apr 2023 | 23,976 | 227 | 44 (0.18%) | zero-width non-joiner, 16 |
| gpt-4o-2024-08-06 | Oct 2024 | 34,681 | 1,091 | 63 (0.18%) | zero-width non-joiner, 55 |
| gpt-4o-mini-2024-07-18 | Oct 2024 | 9,026 | 68 | 3 (0.03%) | zero-width non-joiner, 3 |
| o1-mini-2024-09-12 | Oct 2024 | 1,976 | 74 | 20 (1.01%) | narrow no-break space, 18 |
| gpt-4.1-mini-2025-04-14 | Jul 2025 | 65,589 | 4,370 | 480 (0.73%) | Arabic letter mark, 230 |
The big numbers in the fourth column are mostly harmless. GPT-4o’s answers held 17,087 zero-width non-joiners in 999 answers, and 383 of the first 400 of those answers were in Persian, where the character sits inside words to stop two letters joining. GPT-4.1 mini used the emoji style selector (U+FE0F, the invisible character after ❤ that makes it a red emoji) in 3,293 answers. The remover keeps both unless you untick “Keep the ones that are needed”.
Where the removable ones came from
The only pattern that looks like the model’s own habit is o1-mini’s narrow no-break space (U+202F). It showed up in 18 of 1,976 answers, between numbers and the words or symbols around them: Value = 1.25 × y₂, 7 million new shares, 14 TB disks, even 1 + 1 equals 2, written here with plain spaces where the answers had narrow ones. On screen the two look almost the same. That matches what the education company Rumi reported on 20 April 2025 about o3 and o4-mini. OpenAI told Rumi the characters were “a quirk of large-scale reinforcement learning”, not a watermark, and on 23 April Rumi reported they had stopped appearing. A user on OpenAI’s developer forum reported the same character in GPT-5 answers in October 2025. We had no way to test either model ourselves.
GPT-4.1 mini’s removable characters were mostly copied from the question. Many of the answers that held them were replies in group chats, addressed to people by @name, and those names are full of decoration. 191 of the 244 answers with an Arabic letter mark (U+061C) had it inside an @name, and so did 140 of the 220 with a non-breaking space. The model repeated the name exactly as it was given, invisible parts and all.
GPT-3.5’s zero-width spaces often came in pairs just before a word in Portuguese, Dutch and German answers, as in responsáveis pela with two of them after the space. We can’t say where those came from.
People’s own messages carried more
The same files hold the users’ side of each conversation, 218,535 messages. In April 2023, 0.42% held a removable character, about four times the rate in GPT-3.5’s answers. By July 2025 the figure was 26.43%; many of the messages we looked at were long chat logs pasted in whole. Non-breaking spaces turned up in 13,822 of them, em spaces in 13,154 and Arabic letter marks in 6,119. Whatever lands in a chatbot answer, text that people copy from other apps brings in far more.
What copying and pasting does to them
If a character is in the answer, does it reach your document? We put each of 18 hidden characters between two letters and moved it through the paths text usually takes, in Chromium 153 on 24 September 2026.
| Path | Non-breaking space | The other 17 |
|---|---|---|
| Clipboard pasted into a text box or one-line field | kept | kept |
| Clipboard pasted into a rich text editor | became a space | kept |
| Selected on a web page and copied | became a space | kept |
| JSON.stringify in the browser | kept, unescaped | kept, unescaped |
The other 17 were the narrow no-break, thin, em and ideographic spaces, the zero-width space, both joiners, the word joiner, the byte order mark, the soft hyphen, a left-to-right mark, a right-to-left override, the line separator, the emoji style selector, a tag character, the Hangul filler and the Mongolian vowel separator. Only the plain non-breaking space gets turned into a space on the way out of a page, and the narrow one, which is what the o-series reports were about, gets through.
Rich text editors also make non-breaking spaces of their own. Typing a, two spaces and b into an editable area stored a non-breaking space followed by a normal one, because HTML would otherwise collapse the pair into one. Copying it back out turned it into two plain spaces again. Code sees these characters differently too: in Python 3.14, json.dumps wrote a zero-width space as the visible text \u200b, and split() split on a non-breaking space but not on a zero-width one.
Which ones you can see
| Character | Inter | Arial | Times New Roman | Menlo |
|---|---|---|---|---|
| Non-breaking space U+00A0 | 1.00 | 1.00 | 1.00 | 1.00 |
| Narrow no-break space U+202F | 0.50 | 0.50 | 0.80 | 1.00 |
| Hair space U+200A | 0.22 | 0.30 | 0.33 | 1.00 |
| Em space U+2003 | 3.56 | 3.60 | 4.00 | 1.00 |
| Hangul filler U+3164 | 3.01 | 3.05 | 4.00, drew a mark | 1.41 |
| Blank Braille cell U+2800 | 2.43 | 2.46 | 2.73 | 1.14 |
| 12 more: zero-width space, both joiners, word joiner, byte order mark, soft hyphen, left-to-right mark, right-to-left override, emoji style selector, a tag character, Mongolian vowel separator, invisible times | 0 | 0 | 0 | 0 |
Helvetica Neue gave the same zero widths. Two traps sit at opposite ends of this table. The zero-width group can’t be seen at all, and the soft hyphen only shows as a hyphen when a line happens to break at it. The Hangul filler and the blank Braille cell are wider than a space yet blank, so a name or a line made of them looks empty. Our remover shows each one as a labelled box so you don’t have to guess.
Hidden messages are a separate problem
Two blocks of Unicode can carry whole messages. Tag characters (U+E0000 to U+E007F) mirror the keyboard’s letters, digits and punctuation and are drawn as nothing; their legitimate use is inside a few flag emoji, such as England’s. Johann Rehberger showed in January 2024 that chatbots read instructions written in them. Variation selectors are the other: Paul Butler showed in February 2025 that 256 of them can encode any byte, so a run of them after an emoji can hold a message of any length.
We found neither in WildChat. None of the 438,910 turns, questions or answers, held a tag or variation selector run that spelled out readable text. Early versions of our decoder did flag 805 user messages, all of them the emoji style selector repeated, from twice to more than sixty times, after one heart or smiley, so it now ignores repeats of one selector. When the remover does find a message, it shows the decoded text above the table before taking it out.
For curly quotes, long dashes and garbled ’ characters in the same text, the straight quotes tool runs the same code with every punctuation rule on, and our guide to plain punctuation covers what those characters break.
Source for the dataset: WildChat-4.8M, ODC-BY licence.
Sources
- WildChat-4.8M dataset card (Allen Institute for AI, ODC-BY)
- Zhao et al., WildChat: 1M ChatGPT Interaction Logs in the Wild (ICLR 2024)
- Rumi: New ChatGPT models seem to leave watermarks on text (20 April 2025, with updates)
- OpenAI Developer Community: GPT-5 outputs U+202F instead of normal spaces (October 2025)
- Embrace The Red: hiding and finding text with Unicode tags (January 2024)
- Paul Butler: Smuggling arbitrary data through an emoji (February 2025)
- Unicode Technical Standard #51: emoji tag sequences and presentation selectors
- Dathathri et al., Scalable watermarking for identifying large language model outputs, Nature 634 (2024)
- Google AI for Developers: SynthID Text