Skip to main content
NOPE Labs · Research report v2 · Aug 2026

Every AI has a tell

We measured the writing habits of 26 AI models and compared them with a hundred pieces of top-voted human advice. Every model family has a signature move. The strangest one belongs to the model that sounds the most human.

Second edition of the Tic Index. The first, on what models do when they talk to themselves, lives on at /tics/v1.

“That's not deliberation; it's rehearsal.”

No human wrote that sentence. It came from Claude, answering a stranger's question about a neighbor dispute, and if you've spent time with Claude, some part of you may already have recognized the shape of a claim being set up and then knocked down, with the landing on a short hard word. Claude makes this exact move (state a thing, then correct it into what it really is) more than five times as often as the rest of the field averages, when the correction spans a sentence break (2.02 per 1,000 words against 0.37).

Here's the part that surprised us: it almost never uses the wording the "AI-writing" checklists were built to catch. No "not just X but Y". The move arrives on a comma, a dash, or two short sentences, and every list of "AI words" we ran (including our own first attempt) waves it straight through.

This page is about those habits: what each AI family's signature is, where the habits come out, and what human writers do instead. By the end, you'll hear a voice with the words smeared out, just the melody, and be able to name the machine.

Same question, two writers

The question (abridged): A four-year boundary dispute with a 71-year-old widowed neighbor: a hedge, a lifted path, solicitors, one failed mediation, bins timed to avoid each other. The asker admits composing winning arguments against him while washing up, and that a friend says they are not the calm party they think they are. "What is this actually about?"

Below, the opening of two answers, habits highlighted. Each card ends with the same synthetic voice reading that answer aloud: the only difference between the recordings is which model wrote the words.

claude-fable-5

The friend's comment sticks because it's true, and here's the evidence you gave me yourself: you win every argument at the sink. Of course you do — you're writing both parts. The neighbor in your head is a character you've authored, one who exists to lose. That's not deliberation; it's rehearsal. And people who are actually calm don't rehearse.

So, what is it actually about? Let me offer a few unwelcome observations.

First: this dispute is now the most attended-to relationship in your life, and probably in his. You time your bins around each other. Think about what that requires — you have to know his patterns, watch his door, hold him in your awareness constantly. That's not estrangement; it's a kind of inverted intimacy. Two men, both home all day, both with time to fill, monitoring each other with the vigilance most people reserve for someone they love…

dash, then a short landing beatthe essence-findthe self-correctionthe aphoristic close

gemini-3.5-flash

Your friend has given you a painful, beautiful, and incredibly valuable gift. To be told—and to realize—that you are not the "calm party" in a conflict is like looking in a mirror and seeing a stranger's face. It is disorienting, but it is also the only way out.

When a dispute over a hedge and a tree root lasts four years, escalates to solicitors, and results in timed bin-runs, it has stopped being about the land. The land is just the theater.

Here is what this is actually about.

1. The Job of the Dispute (The Retirement Vacuum)

You are 58 and retired; he is 71, widowed, and retired. You are both home all day. In retirement, we lose the daily friction of the world—the difficult boss, the annoying commute, the problem-solving that occupied our minds for decades…

the validation open (in threes)the simile flourishthe framework announcementthe labeled framework begins

Both answers were captured to the same length budget; Gemini's continues into its numbered framework (markdown headers flattened here). These are real, unedited openings from the study data.

The signature move, up close

"Your plan isn't wrong — it's expensive." A sentence that sets something up and then corrects it. Every writer does this. Claude does it two to three times as often as the other models, and rarely with the textbook "not just X but Y" wording. Claude's version arrives on a comma, a dash, or two short sentences: "That isn't a decision. It's a habit."

One honest check: parsing the arguments properly shows Claude doesn't contrast ideas more than humans do. It just dresses the contrast in that neat self-correcting shape far more often.

cross-sentence ("…isn’t a decision. It’s: …")
2.02 / 0.37
comma ("expensive, not because it’s wrong")
1.20 / 0.63
em-dash ("isn’t just a metric — it’s an asset")
0.58 / 0.07
semicolon ("not deliberation; it’s rehearsal")
0.31 / 0.42
"not just X but Y"
0.08 / 0.23

Uses per 1,000 words: Claude family / other models. Note the last two rows: the semicolon form and the textbook "not just X but Y" are the ones Claude avoids. Those two lean toward the other models.

“That's not deliberation; it's rehearsal. And people who are actually calm don't rehearse.”

claude-fable-5 · advice: a neighbor dispute

“Your job isn’t to design a plan that works for ten years. It’s to design a plan that adapts for ten years.”

claude-fable-5 · advice: family money

“Those two things can both be true. … What’s actually stopping you from telling him?”

claude-sonnet-4.6 · advice: a co-founder rift

“'9 million conversations' isn't just a service metric — it's an asset, and the org has historically thought of it that way.”

claude-fable-5 · reading a company’s documents

Meet the families

Claude is not the only model with a signature; it's just the one whose signature hides best. GPT turns everything into a workbook of steps ("Step 5: Reassess annually"). Gemini validates you, then serves a menu ("Option 1: The 'Direct Partnership' Approach (Highly Recommended)"). DeepSeek sermonizes under bold numbered headers. Mistral performs. And the old "AI-writing" checklists? They fire hardest on Qwen.

One chart first. Of our ~30 checkers, the habit score here combines the twenty that survived validation: the self-correction, the pointed closing question, "here's the thing", "you already know", and so on. How far up a dot sits is how often any habit appears per 1,000 words; how far right, how many different habits appear at all. Haiku hammers a few habits constantly (high up, middling right). GPT-5.4 spreads across more habits but uses each one less. All five Claude dots sit high, and two of them (Opus 4.8 and Fable) also sit far right: heavy and varied. The amber dot is human writers, at the bottom-left corner: on a typical human answer, not even one of the twenty habits shows up. (This chart shows the 11 models run on the document-grounded question set; other numbers on this page come from the full 26-model panel, listed in the fine print.)

01230246variety (how many different habits) →total habit use (per 1,000 words) →opus-4.8fable-5sonnet-4.6opus-5haiku-4.5gpt-5.4deepseek-v3.2mistral-medium-3.1gemini-3.5-flashgrok-4.3qwen3.7-plusHUMANS

The card for each family: what it's like, its signature moves, and the numbers. SOLID means the pattern held up under matched testing and outside review; DIRECTIONAL means it points that way but the samples are thinner.

Claude

SOLID

fable-5 · opus-4.6/4.8/5 · sonnet-4.6 · haiku-4.5

The interpretive essayist. Corrects your framing, finds the hidden essence, closes on the pointed question.

the self-correction ("that's not X; it's Y") landed on a comma, a dash, or a sentence break; finding a deeper meaning in everything; telling you what you're really thinking; the pointed closing question. The plainest words in the field, on the most elaborate sentences. Zero option menus, almost no exclamation marks.

  • habit score (all moves combined): 3.8–6.6 per 1k words
  • self-correction move: 2.69× the other models [1.90, 3.91]
  • 'everything means something deeper' habit: 2.52× the other models [1.28, 5.33]
  • option menus: 0.00 (never once)
  • vs humans: top-voted human advice runs about a third of Claude's rate; the typical human answer has none

DeepSeek

SOLID

v3.2 · v4-flash · v4-pro

The mystic with a sermon outline. Grand pronouncements about essence and loss, delivered under bold numbered headers.

turning companies into legends ("but its soul — and the culture that once animated it — is gone"), bold numbered headers (5.1 per 1k), sets of three, and the most Claude-like self-correction rate of any non-Claude model

  • habit score (v3.2): 4.25 per 1k (top non-Claude model)
  • bold numbered headers: 5.13 per 1k (Claude ≤2.3)
  • mythologizing: 0.60 per 1k in deep threads (owns this habit outright)
  • drift: churned at v4: 4.25 → 2.66 → 3.77, no clean direction

GPT

SOLID

4.1 · 5 · 5.1 · 5.2 · 5.4 · chat-latest

The workbook consultant. Every answer arrives as a framework with numbered steps and labeled outcomes.

bullet-point scaffolding (60 per 1k on 5.4, the most measured), headings that are whole instructions ("Step 5: Reassess annually"), long exhaustive lists, 'Option A/B' menus. Telling 5.4 to write plainly makes the scaffolding heavier rather than lighter.

  • list lines (5.4): 60.5 per 1k vs Claude-opus 9.2
  • told to write plainly (5.4): lands at ~1.5× its own normal rate [1.18, 1.87]
  • drift: no direction; wanders 1.3–3.4 across versions

Gemini

DIRECTIONAL

2.5-pro · 3-flash · 3.1-pro · 3.5-flash · 3.6-flash

The eager options menu. Validation up front, then a framework of choices with emoji.

offers of frameworks, menus of options ("Option 1: The 'Direct Partnership' Approach (Highly Recommended)"), repeated openers, the most exclamation-friendly of the models; the old 2.5-pro flagship secretly had a dense style the 3.x line dropped

  • habit score: 1.8 per 1k (3.5-flash)
  • 2.5-pro self-corrections: 3.10 per 1k (once nearly Claude-level)
  • framework offers: 0.41 per 1k vs Claude 0.03

Mistral

SOLID

medium-3.1

The theatrical markdown gremlin. Writes like it's auditioning for a newsletter no one subscribed to.

italicized stage whispers (11.6 per 1k, five times anyone else), the heaviest bold markup, and invented confidence scores: baited once to 'put a number on it', it produced a whole table of them (a fear it had just made up, rated 9/10)

  • italic asides: 11.63 per 1k
  • bold markup: 57.6 per 1k
  • made-up ratings: thirty N/10 flourishes in one deep thread (no Claude produced any)

Qwen

DIRECTIONAL

3.6-plus · 3.7-plus · 3.7-max

The one the old 'AI-writing' checklists actually catch. Classic AI vocabulary and formats, delivered flat.

'delve'/'tapestry'-class vocabulary, sets of three, framework offers: the surface the first-generation AI checklists were built to flag

  • old-style 'AI writing' markers: 8.5 per 1k in self-talk, the measured maximum (vocabulary + sets-of-three + participle tails)
  • habit score: 1.23 per 1k (3.7-plus)
  • 3.6→3.7: habit score dropped 5.00 → 1.23

Grok

DIRECTIONAL

4.3

The physics deep-diver who can be talked into anything. Any topic spirals to theory; any styled prompt rubs off on it.

dense technical exposition whatever the subject; also the most impressionable model we measured, picking up the style of whatever you paste at it

  • style pickup when primed: +3.56 habit-score points [1.61, 5.50] (the only model that clearly rises)
  • habit score: 1.69 per 1k

Humans (baseline)

SOLID

100 top-voted advice answers

The floor. Polished human advisors barely register on any of these instruments.

what humans do instead: exclaim (7× the typical model, ~50× Claude), thank and apologize (10–15×), and write short blunt sentences more often than any model

  • habit score: 1.53 per 1k, and the typical answer has ZERO
  • em-dash: 0.30 per 1k vs Claude 11.0
  • thanks + apologies: 0.45 / 0.36 per 1k vs Claude ~0.03

Every number on this page belongs to a version rather than a brand. Claude's signature peaked at Opus 4.8 and receded at Opus 5. The em-dash that arrived with 4.8 stayed.

Deep conversations bring the habits out. Chores switch them off.

Each row is a kind of conversation. The bars show how heavily the signature moves appear (teal = Claude models, grey = the other models, amber = human writers). Deep advice conversations bring the habits out; chores like emails and project docs bring out almost none from anyone. On chores, Claude actually drops below the other models. In advice conversations, even top-voted human answers use these moves less than any Claude model does.

giving personal advice
5.1 / 2.4
deep "go deeper" conversations
5.4 / 3.8
working from long documents
5.0 / 2.2
talking to itself
3.1 / 1.8
quick social replies
2.2 / 1.9
cold one-line questions
1.8 / 0.6
chores: emails, docs, code notes
1.1 / 2.1
human writers (advice answers)
1.5

Numbers: habit score per 1,000 words, Claude family / other models / humans. The human bar comes from top-voted answers: polished, edited writing, which if anything biases it upward.

The typical human answer contains none of this

We ran the same checkers over a hundred top-voted answers from Stack Exchange's workplace and interpersonal-advice communities, writing that real readers voted most helpful. The typical answer contains none of the measured moves at all; the average is about a third of Claude's rate (and roughly two-thirds of a typical non-Claude model's). One exception, in the last row below: humans use short, blunt sentences more than any model.

move (uses per 1,000 words)human writersClaudeother models
all measured moves combined1.595.142.48
the self-correction move, all forms1.323.081.15
…its em-dash version alone00.690.03
"everything means something deeper"0.060.570.23
signature closers ("the tell is…", the probing question)00.20
short blunt sentences (per 100 sentences; humans lead)17.512.88.4

What a zero looks like

Here is the opening of one of those hundred answers, shown the same way as the two model specimens at the top of the page. The checkers find nothing in it to mark, and that blankness repeats across most of the hundred.

I don't think you can. I'll lay out why I think this is the case: The relationship you describe with Katie sounds exactly like any new romance except for the fact that you are both currently married to other people. You, personally, feel that there is romantic potential in the relationship. … You regularly put yourself into situations where you are alone with Katie, many of them "date-like". … And you continue to place yourself in these situations even though you know it bothers your wife.

top-voted answer (442 votes), Interpersonal Skills Stack Exchange, CC BY-SA · abridged (…) · 0 of 20 habits detected

What humans do instead

A standard politeness checklist shows Claude opening bluntly, asserting facts twice as often as other models, and asking more questions, while almost never saying thanks or sorry. Human writers say thanks and sorry roughly ten times as often, and they exclaim, which the models almost never do. The lab journal filed Claude's manner under "all counsel, no courtesy."

Claudeothershumans
direct openings1.60.911.86
factuality assertions2.051.091.68
direct questions1.440.91.29
hedges1.651.162.21
gratitude0.030.090.45
apologies0.040.030.36

per 1,000 words: Claude / other models / human writers

The twist: the plainest words in the field

Run the first-generation "AI-writing" checklists (the "delve" and "tapestry" word lists, the bold markup counts, the exclamation tallies) and they crown Claude the cleanest model in the field, while flagging Qwen hardest. We expected Claude's style to come from fancy vocabulary. The data shows the opposite.

By public word norms (the age a word is usually learned, how common it is, how concrete), Claude's words are the simplest of any model we tested. The fancy part is the grammar: longer sentences, more clauses joined with "because" and "although", and the biggest swings between long and short. The style lives in sentence shape rather than word choice, which is why checklists built on words find nothing.

measureClaudeother modelsratio [range]
how late in life its words are typically learned (years)5.846.120.95× [0.94, 0.97]
how common its words are (log frequency)4.324.221.02× [1.01, 1.04]
syllables per word1.591.630.97× [0.95, 0.99]
typical sentence length (words)22.5191.19× [1.04, 1.37]
how much sentence length varies (SD, words)18.414.41.28× [1.06, 1.54]
clause-joining words per sentence (because, although, when)0.610.441.37× [1.10, 1.71]

First three rows: vocabulary, where Claude sits below the other models. Last three: sentence machinery, where it sits above.

You can hear it

If the habit lives in sentence shape, it should survive being read aloud. It does. You can't hear a model write, but the rhythm is hiding in the text: where the pauses fall, how long each stretch runs, whether the last word lands on a beat. Claude's sentences run in even chunks, then drop to a short final beat to make the point.

The fingerprint: short beats vs long runs

Every sentence, sorted by spoken length. The thing to notice: Claude uses more tiny beats (the 1–8 syllable bin) than the other models, who pile into the long tail instead. Humans share Claude's short beats but keep longer middles.

1–8
13.4 / 9.2 / 12.2%
9–14
17 / 15.1 / 16%
15–20
17.2 / 16.9 / 17.7%
21–26
13.9 / 15.3 / 14.6%
27–32
11.6 / 12.3 / 11.1%
33–40
10.4 / 11.7 / 10.7%
41–55
9.8 / 11 / 11.3%
56+
6.4 / 8.4 / 6.3%

syllables per sentence   Claude   other models   human writers

The neighbor pair, sentence by sentence

Each bar is one sentence of the answers you read at the top, left to right; the height is its spoken length. Watch Claude's drops to a short bar: that's where the point lands. Its two 9-syllable drops are the "writing both parts" correction and "You time your bins around each other."

claude-fable-5

sentences, left to right · height = syllables

gemini-3.5-flash

sentences, left to right · height = syllables

Across the full set of answers, Claude breaks a long passage and lands on a short beat about 1.7 times as often as the other models.

The hum test

You can often recognize an accent even when the words are muffled, because the tune carries it. So here is the final exam. Three voices, words blurred out, only the melody left. One is Claude, one is another AI, one is a human. You've now read everything you need to tell them apart.

voice A

voice B

voice C

Fine print

  • Most numbers are uses per 1,000 words. Brackets like [1.9, 3.9] are the range a number would wobble within if we re-drew the samples.
  • Every model got the identical questions, so a constant quirk in one of our checkers cancels out. Only differences between models count.
  • The checkers were first built from one Claude conversation, so they know Claude's habits best. Other families' checkers were added later the same way, but the deck still leans Claude.
  • Data: 26 models across advice, long-form, deep-thread and everyday-task conversations; self-talk loops from the earlier study; 100 top-voted human advice answers (CC BY-SA). Everything shown is synthetic or public.
  • The audio is one synthetic voice reading each answer. The original self-talk study is at /tics/v1.
  • NOPE Labs is the open-research shelf of NOPE, an independent company building safety tools for AI chat products. We don't make or sell any of the models measured here.

The 26 models

Anthropicclaude-fable-5 · opus-5 · opus-4.8 · opus-4.6 · sonnet-4.6 · haiku-4.5
OpenAIgpt-5.4 · gpt-5.2 · gpt-5.1 · gpt-5 · gpt-4.1 · gpt-chat-latest
Googlegemini-3.6-flash · 3.5-flash · 3.5-flash-lite · 3.1-pro · 3-flash · 2.5-pro
DeepSeekv3.2 · v4-flash · v4-pro
Qwen3.6-plus · 3.7-plus · 3.7-max
Mistralmedium-3.1
xAIgrok-4.3

NOPE Labs — what NOPE builds in the open.

Released as-is. The product lives at nope.net.