Skip to content
AI Origin CheckRun a free check

Why the tells fail

The supposed signs that text was written by AI

A list of supposed AI writing tells circulates constantly: em dashes, the word 'delve', tidy tripartite structure, hidden Unicode. Each has a real observation behind it and each fails as a test. This guide explains why, and what remains once you discard them.

Updated

The em dash

The observation is real: language models use em dashes more often than the average internet commenter. The inference is not, because the comparison class is wrong. Models were trained heavily on edited prose — books, journalism, essays — where the em dash is a normal piece of punctuation.

So the em dash marks writing that resembles edited prose. Professional writers, academics, and anyone who reads a lot of books use it naturally. Casual chat writing does not. The test therefore separates formal from informal registers, not human from machine.

The real-world consequence has been people editing em dashes out of their own writing to avoid suspicion. That is a test corrupting the thing it claims to measure, and it is a good sign the test was never measuring what it claimed.

Hidden Unicode characters

This one is worth addressing directly because it feels the most technical and is the most wrong. The claim is that AI-generated text contains zero-width spaces or similar invisible markers inserted as a signature.

Language models generate ordinary text. The invisible characters people find come from the copy-and-paste path: PDF extraction inserts soft hyphens, word processors insert non-breaking spaces, web pages use zero-width spaces to control wrapping, and encoding conversion leaves byte order marks.

So a text full of invisible characters is a text that has been through a document pipeline, regardless of author. A text with none is a text typed into a plain editor, regardless of author. What the scan measures is real; what people conclude from it is not.

Vocabulary tells

Particular words get identified as machine markers — 'delve', 'tapestry', 'testament', 'navigate the complexities'. There is a real pattern underneath: models do overuse certain formal register words relative to casual speech.

But these words are also the ordinary vocabulary of academic writing, of professional communication, and notably of writers using English as a second language, who often learned from formal sources and reach for formal register naturally.

The documented consequence is that non-native English speakers are flagged as AI users at substantially higher rates than native speakers. A test whose errors concentrate on one population is not a neutral instrument, whatever its average accuracy looks like.

Structural tells

Tidy three-part structure, balanced paragraph lengths, a summary sentence at the end of each section, consistent hedging. Models do produce these patterns because they were trained on well-structured prose and tuned toward helpful clarity.

The problem is that these are also the properties of good writing, and specifically of writing produced under editorial guidance. Anyone taught to write clearly produces text with these features. Anyone who revises a draft produces text with these features.

The tell effectively measures editing effort. Punishing it means punishing the writers who worked hardest on their prose.

What actually survives scrutiny

Very little, at the level of the finished text — which is the honest and uncomfortable conclusion.

Factual errors of a specific kind are mildly informative: confidently stated, plausibly shaped, and wrong in the details. Invented citations, plausible-sounding but non-existent references, and fluent misstatements about niche subjects are more characteristic of generated text than of a careless human writer, who tends to be wrong in messier ways.

Absence of specificity is another weak signal: text that discusses a topic competently while never landing on a concrete detail, a real example, or a first-hand observation. But a great deal of human corporate writing has exactly this property.

Neither of these is a test. They are things that might prompt you to ask a question, which is a very different standard from evidence.

What to do instead when it matters

If you are assessing work and the answer has consequences for a person, the finished text is the wrong evidence to be looking at. Process evidence is both more reliable and fairer: drafts and version history, a conversation about the reasoning and the sources, or a short supervised writing sample for comparison.

If you are on the receiving end of an accusation, ask what specific mechanism the tool claims to measure and what its published false-positive rate is. Most cannot answer either, and a tool that cannot is not producing evidence.

If you are publishing, disclose. Regulatory direction is toward declaring AI involvement rather than detecting it afterwards, precisely because reliable after-the-fact detection of text does not exist.

Frequently asked questions

Do em dashes mean text was written by AI?

No. They indicate writing that resembles edited prose. Books, journalism, and academic writing use em dashes routinely, and models learned the habit from that material.

Do hidden Unicode characters indicate AI text?

No. They come from PDF extraction, word processors, web page copy, and encoding conversion. They reflect how a passage was moved between applications, not who wrote it.

Are AI writing detectors accurate?

Not reliably enough for decisions about individuals. Their false positives concentrate on formal, edited, and non-native writing, which is a documented and serious fairness problem.

Can I prove I wrote something myself?

Not from the finished text, but process evidence works: version history, drafts, notes, and the ability to discuss your sources and reasoning in detail.

Is there any reliable tell at all?

Nothing that functions as a test. Confidently-stated factual errors and invented citations are mildly informative, but they prompt a question rather than settle one.

Primary sources

The technical claims on this page follow the published specifications below rather than our own assertions.

Keep going