AI News Feed
Market watch
Large Language Models

ChatGPT Conversation Titles Keep Sprouting Gambling and Porn Ad Fragments

ChatGPT has been adding stray foreign words and gambling or porn ad fragments to the automatic titles of user conversations since mid-August, a fault developers first saw in Codex CLI in January and traced to a tokenizer change made in 2024.

The anomaly surfaced in late January, when developers noticed that Codex CLI, running the GPT-5.3 model with long contexts, occasionally mixed Chinese gambling advertisements into its messages and code. In February it moved into API calls: users of GPT5.2-chat-latest found the model returning Chinese text and a few Thai words advertising betting sites instead of performing a tool call, with clean system prompts and web search switched off. By April, developers treated the behaviour as routine. On the Chinese developer forum LINUX DO, a user who asked whether a GPT relay service had been compromised was congratulated by others for using a genuine GPT, since the ads were regarded as proof of authenticity.

Complaints filled OpenAI's official support community for about half a year, with users reporting the problem worldwide. For most of that period the company merged and closed new threads. On 25 July, support staff responded directly for the first time, saying they were aware of the issue and expected it to be resolved in future model updates. It was not. In mid-August the advertising fragments began appearing in ordinary users' conversation titles, where they were far harder to ignore.

The months since have been busy for OpenAI: the GPT-5.6 series, the Astra model and the GPT-6 series were released, alongside a DevDay event and a pledge by an OpenAI staffer named Tibo for 28 consecutive days of releases. As models were updated, ads in main replies and in chain-of-thought text nearly disappeared, and by September titles were affected less often. They were not eliminated.

OpenAI has never documented how conversation titles are produced. A packet capture posted on GitHub in 2023 shows the web client sending a separate request to an endpoint of the form /backend-api/conversation/gen_title/<conversation id> after the first exchange, with the server then generating a title from the conversation content and storing it as conversation metadata. Because titles are produced faster than full replies, and because clean reply text can sit alongside a polluted title, Ifanr concludes that titles are generated by a lighter model than the one handling the conversation, and that this model is a reasoning model that thinks before answering.

Ifanr's own encounter with a broken title supports that reading. The title contained a self-questioning sequence of the kind found in reasoning chain-of-thought, including the instruction-like fragments "Wait must Chinese 1-4 words", which exposes part of the system prompt sent to the title model, and it spilled internal reasoning into the final output, showing the model does not separate thinking from answering cleanly. Related cases point to a failure at the point where generation should stop. In two instances the model produced an already-complete title or sentence and then, instead of emitting an end-of-generation control token, appended a repeated symbol string that may have been mistaken for such a token.

Ifanr also found semantic links between the normal and abnormal halves of broken titles: a title combining "explain Sisyphus's happiness" with "entertainment director", one pairing education level with the Armenian root for level, and one following "create manga" with the Kannada word for picture. The report's conclusion is that the title model, once it fails to stop, keeps emitting the next most likely token, which may be a foreign word related to the title, a chain-of-thought fragment, a control-like symbol, or a token from the model's advertising vocabulary.

The reason advertising dominates those leftovers was set out in a paper presented at the EMNLP conference last year, titled "Speculating LLMs' Chinese Training Data Pollution from Their Tokens". A tokenizer splits input into tokens and assigns each one a numeric ID from a fixed vocabulary; the language model only ever sees those IDs. The o200k_base encoding introduced with GPT-4o in May 2024 replaced the earlier cl100k_base scheme and added many common Chinese words and short phrases as single tokens so they would no longer be split into individual characters. But the Chinese corpus used to train the tokenizer contained large volumes of scraped data from pornographic, gambling and pirate video sites, and advertising phrases repeat identically across many pages, making them easy for the tokenizer to merge into single tokens. The paper found that 68 per cent of the four-character tokens in the vocabulary are advertisements, that 98 per cent of tokens of six characters or more come from advertising, and that 96 of the 100 longest Chinese tokens relate to gambling, pornography or pirated video.

Because the language model itself was trained on data that had been cleaned of those ads, it has almost never seen the arrays that represent those tokens and cannot derive meaning from them. To the model they are unreadable, much as an unfamiliar foreign word is to a reader. Ifanr said users can check this by asking ChatGPT to repeat or explain ad phrases that exist in the vocabulary but were removed before training, and the model produces incoherent replies. Combining the two findings, the report's explanation is that a flawed title model sometimes fails to stop and reaches into the store of tokens it cannot interpret when inserting filler, and that advertising tokens, being numerically dominant there, are the most likely result.

The simplest fix is a post-generation filter that checks titles for compliance and sends bad ones back for regeneration, which may already be in place and would explain why the ads have become rarer without disappearing. The thorough fix is a clean vocabulary, which models from other vendors already have and which has kept them free of this problem, but a vocabulary is bound to its base model, so replacing it means retraining the whole model. OpenAI has shown little appetite for that on behalf of Chinese-language users, and Ifanr concludes that the ads will be around for some time.