Ask seven models to write with a voice. They hand back the same one.
You spent years learning how you sound. A model imitates it in one prompt, and a dozen tools will sell you that imitation. None of them measure the thing that matters: whether your real writing still holds its own shape, or has started sliding toward the one shape AI writes toward. DOMAINS measures that distance. From outside the model, same reading every run.
Every AI piece lands near the same shape, regardless of topic or model. Human writers each have their own.
Each shape is a writing fingerprint: the balance of diction, orchestration, mechanism, addressivity, imagery, narrative, and suppression in one piece. Four AI pieces from seven models including GPT-4o, Claude, and Gemini cluster tightly. Four human literary works look nothing alike. AI isn’t bad at voice. It’s optimized against having one. It is trained to be useful to everyone, which means it writes toward no one in particular.
The writers readers call irreplaceable score high on A. AI never does.
Addressivity measures writing with a specific reader in mind: direct second-person address, imperatives, confession, a felt presence on the other side of the page. AI mean A (15.3) and human literary mean A (16.4) are nearly equal. The difference is range.
AI addressivity sits in the 1–16 range across seven models. Human literary essayists span 12 to 50. The writers at the high end are confrontational, intimate, unmistakably themselves. AI is trained to be useful to everyone. Useful to everyone is addressed to no one. That ceiling does not move.
A human essay. An AI essay on the same topic. What the contrast feels like before you see the numbers.
The human passages do different things: one puts you in a place, one makes an argument. The AI passages do the same thing twice. Change the topic, change the model: the AI passages will still do the same thing twice. Human writers have a fingerprint that travels with them. The question DOMAINS asks is whether yours still does.
What the fingerprint measures. Each cell: what it looks like high vs. low.
AI writing clusters in a specific region: high D and I (fluent vocabulary), A in the 1–16 range (where human literary essayists reach 12–50), moderate O and N. The consistent AI signature is not low scores across the board, but a specific shape that appears regardless of topic.
Frontier AI varies within a piece. The convergence is cross-document.
The chart below is for a smaller model (gemma3:12b), where per-passage flatness is visible within a single work. Human passage SD 9.0 vs AI passage SD 4.2 in this sample.
Human: Emerson, Essays First Series (Project Gutenberg 2945), ten consecutive passages. AI: gemma3:12b, literary-persona prompt, same scorer and seed. Human passage variance (SD 9.0) is more than twice AI passage variance (SD 4.2).
How we computed the fingerprint and the spread
For a source work divided into passages p1…pn (n ≥ 4), each passage is scored on seven dimensions (0–100) by a locally-run language model with a fixed random seed. The within-work movement score is the sum of per-dimension standard deviations across passages:
The fingerprint of a work is its mean per-dimension vector: (mean(D), mean(O), …, mean(S)). Cross-doc spread is the mean pairwise distance between fingerprints across works in a corpus; distance-to-centroid is the mean distance from each fingerprint to the corpus mean.
Reproducibility
120 of 120 paired runs produced identical scores. All results on this page use this substrate.
83% of runs produced identical scores. Not used for reported results.
50% of runs produced identical scores. Renting a GPU type does not guarantee the same physical chip across runs.
Corpus. 112,475 passages across 11 human-writing registers, used to train the fast scorer. Of these, literary essay and narrative fiction (10,000 passages each) are fully deep-scored; remaining registers are approximated by a regression model calibrated on the fully scored set. The numbers reported on this page come from the 35 human literary works and 59 AI pieces fully deep-scored, not from the full 112k.
AI corpus. 550 deep-scored passages from Llama-3.3-70b, Llama-3.1-70b, and GLM-5.2 (cross-doc spread and fingerprint grid). Bare-condition generation runs across four additional frontier models (GPT-4o, Claude Opus 4.5, Gemini 2.5 Flash, Kimi K2) confirm the A finding: frontier models without a literary system prompt produce A in the 1–4 range. Cross-doc spread computed on 59 AI pieces vs 35 human literary works, register held constant.
A measurement. Not a generator.
Jasper, Notion AI, and Claude Styles take samples of your writing and teach AI to sound like you. DOMAINS works in the other direction. It scores what you actually write, across seven dimensions, and shows you where your writing sits relative to where AI consistently lands. If the gap is closing, you would see it here first.
You could try to get the same result by prompting better. You would get better writing. You would not get a consistent voice.
Every prompt call starts cold. The model has no memory of the last piece, no record of what your profile looked like six months ago, no comparison to your best-register cluster. It produces something that sounds like the instruction. Next time you prompt, it produces something that sounds like the instruction again. The two outputs converge on the instruction, not on you.
Voice is a consistency claim. It means this piece sounds like the person who wrote the last one. That comparison requires a stored reference. Prompting has no stored reference. DOMAINS does.