Why every AI writes the same ad
Each AI answer looks original on its own. Put a thousand side by side and they converge. What the research shows, and why a better prompt won't fix it.
By Ilya Arbabi, 6 October 2026
Language models are individually creative and collectively the same. On standard creativity tests, a model's answer scores about as original as a person's. But put many models' answers side by side and they are far more alike than people's answers are, and switching to another company's model barely helps. So when a thousand businesses ask an AI for "a fun ad for our coffee shop", they get a thousand close cousins of one ad, and each of them thinks theirs is original.
The study
In January 2025, Emily Wenger (Duke University) and Yoed Kenett (Technion) published We're Different, We're the Same: Creative Homogeneity Across LLMs. They gave the same creativity tests to 22 language models and to 102 people recruited online, and compared the answers two ways.
The models came from seven families, chosen so they were not trained by the same company: AI21 Jamba, Google Gemini 1.5, Cohere Command R Plus, Meta Llama 3 70B Instruct, Mistral Large, OpenAI GPT-4o and Microsoft Phi 3.
The tests were three standard measures of divergent thinking:
- Alternative Uses Task: list unusual uses for an everyday object (a book, a fork, a table, a hammer, a pair of trousers).
- Forward Flow: starting from a word (candle, table, bear, snow, toaster), write down whatever comes to mind next, and next.
- Divergent Association Task: name ten nouns that are as unrelated to each other as possible.
Then they measured two different things. Individual originality: how far each answer strays from the obvious, scored automatically. And population-level variability: how different the answers are from each other across all the respondents, measured by turning each answer into a sentence embedding and computing how far apart they sit.
Individually, the models hold their own
On originality, the models and the people were close. The models scored slightly higher on the Alternative Uses Task (0.711 against 0.696) and on the Divergent Association Task (0.801 against 0.753); the people scored slightly higher on Forward Flow (0.637 against 0.603). The differences were real but small.
This is the part everyone experiences. Ask a model for ten uses for a fork and you get a respectable, sometimes clever list. Nothing about any single answer looks derivative.
Together, they collapse
The second measure is where it falls apart. Across respondents, the models' answers were far more similar to each other than the people's were:
| Test | Models | People | Effect size |
|---|---|---|---|
| Alternative Uses Task | 0.459 | 0.738 | 2.2 |
| Forward Flow | 0.534 | 0.835 | 2.0 |
| Divergent Association Task | 0.665 | 0.819 | 1.4 |
Higher numbers mean more variety. By the usual convention an effect size of 0.8 already counts as large; these are two to nearly three times that.
The authors checked the obvious objections. Models write in a more uniform structure than people, so they asked the models to match the length of human answers, and separately compared only single-word answers. The gap stayed. As they put it, it is "the substance, not the structure" of the answers that makes them alike.
They also tried telling the models to be creative. A "very creative" system prompt raised the models' variability on the Alternative Uses Task from 0.459 to 0.576. That's an improvement, and it is still well short of the people's 0.738.
Their conclusion is blunt: LLM users "may find their creative outputs remarkably similar to those of other LLM users regardless of the model used."
Why you can't see it from where you sit
The unsettling part is how invisible this is. You only ever see your own answer. Judged alone, it scores as original as a human's, so it feels fresh. The sameness only shows up at the level of the population: across all the businesses asking the same models the same kind of question at the same time.
In marketing, the population is the feed. Your competitors are asking for "a bold, fun ad for our coffee shop", and the models are handing each of them a close variation of the same bold, fun ad: the same angle, the same joke shape, the same "it's not just coffee" turn. Nobody did anything wrong, and everybody's ad looks like everybody else's.
Switching models doesn't help much
A natural response is to use a different model from everyone else. The study suggests this buys little. The seven families were built by different companies on different data, and their answers still clustered together far more tightly than people's. The authors also looked at five models from Meta's Llama family on their own; related models were, if anything, a little more alike again.
It goes all the way down to the names
The pattern shows up in small choices too. A 2026 study, The Ghost Couple by MichaĆ Brzozowski and Neo Christopher Chung, asked models to invent fictional experts and found each model family reaching for the same names, and the same pairs of names, again and again: Elena Vasquez and Marcus Chen from Claude, Aris Thorne and Lena Petrova from Gemini, Elara Voss from GPT. Those names now appear across hundreds of web pages, books and even fake academic papers. I wrote that study up separately in The ghost couple.
Why turning up the temperature is the wrong lever
The usual advice for more varied output is to raise the temperature, the setting that makes a model pick less likely words. It's the wrong lever for this problem, for a simple reason: temperature adds randomness to each word, not new ideas. The model's sense of what "a creative coffee-shop ad" looks like doesn't change; you get the same ideas, worded less steadily, and past a point, less coherently.
The evidence points the same way. The Ghost Couple study ran at a temperature of 1.0, the default for most model APIs, and still found one Claude version putting Elena Vasquez into two-thirds of its fictional experts. Wenger and Kenett found that even explicitly asking for creativity closes only part of the gap.
What does work: structure
If the model's own choices converge, the fix is to stop letting the model make them. Decide the creative choices outside the model, and ask it to write inside them:
- Decide the angle yourself. Not "write a fun ad", but "write it as a weather report", or "the product never appears until the last second". A model given a specific, unusual constraint can't fall back on the average ad.
- Bring in outside randomness. Draw a random word, place, era or constraint from a long list you wrote, and force the idea to connect to it. The randomness comes from the list, not from the model, so it doesn't converge.
- Feed it what only you know. Your customers' words, your real numbers, the strange thing that happened in the shop last Tuesday. Specifics are the one input no other business has.
- Get several independent drafts, then choose. Several writers who haven't seen each other's work, each pushed a different way, then a person picks.
- Remember what you did last time. Sameness also happens inside one brand, week after week. Steer each new idea away from your recent ones.
What we did about it
This is why Marsell doesn't ask a model to be creative. Before any model writes, Marsell draws the creative choices in code, from large pools, with a random draw that is different for every brand and every week: the angle each writer is forced through, a constraint, an era, a genre, the place, the look. The models write inside that draw. A thousand people typing "give me an ad" get a thousand different starting points. How that works in detail is in The chaos engine.
FAQ
Do all AI models write the same thing?
Not word for word, but far more alike than people do. In Wenger and Kenett's 2025 study, answers from 22 models across seven model families were much more similar to each other than answers from 102 people, on three standard creativity tests.
Is AI less creative than people?
On a single answer, no: the models scored about as original as the people, slightly higher on two tests and slightly lower on one. The difference is in variety across many answers, where the models were far more alike.
Will a better prompt fix it?
Only partly. Telling the models to be "very creative" raised their variety on one test from 0.459 to 0.576, still well below the people's 0.738. Constraints and choices made outside the model help more than asking for creativity.
Does raising the temperature help?
It adds randomness to the wording, not new ideas, and it costs coherence. The study that found models reusing the same fictional names ran at a temperature of 1.0 and still found the same names in up to two-thirds of answers.
How do I make AI-written marketing original?
Make the creative choices yourself, or draw them at random from lists you control, then let the model write inside them. Add specifics only your business has, generate several independent drafts and choose, and steer away from what you did recently.