The short version. A language model doesn’t know what’s true. It knows what has been written most often. And since the capacity to produce text is distributed with grotesque inequality, what comes out isn’t the best-founded position — it’s the one with the largest press office.
Then comes the serious part: the source vanishes. A corporate press release and a peer-reviewed study come out of the same mouth, in the same measured voice, and once provenance is gone you can no longer ask who benefits.
What follows takes the mechanism apart in five movements: volume becoming weight, the laundering of provenance, the collapse of models trained on other models, the three layers that manufacture consensus, and the feedback loop that returns it all to the world under somebody else’s letterhead. At the centre, a proof that lands close to home for anyone in markets — the word “transitory,” inflation, and a Federal Reserve that in twelve months held two opposite positions in the same steady voice.
Ten sources, every one of them consulted at source and checked. Limits and conflicts of interest declared in the notes.
📑 Contents
- First: volume as weight
- Second: laundering the provenance
- Third: the tightening circle
- Fourth: the layers above frequency
- Interlude: a word withdrawn
- Fifth: the feedback loop
- Cui prodest: what to do with a mint mark
Atqui nulla res nos maioribus malis implicat
quam quod ad rumorem componimur,
optima rati ea quae magno adsensu recepta sunt.Nothing entangles us in greater evils than conforming to prevailing opinion,
taking for best whatever has won widespread assent.— Seneca, On the Happy Life, I, 3

He writes it to his brother Gallio in the first century, three lines after comparing men to a flock that goes not where it ought to go but where everyone else is going.[1] He had never seen a data centre. He had already isolated the word that matters: recepta. Received.
Not true. Received.
A listed company files its quarterly report. Eighty pages, tables, footnotes — the place where bad news goes to hide. Five hundred people read it.
That same morning its press office sends out four hundred words explaining how it should be read. The wires pick it up, then the papers, then the trade sites, then the aggregators, then the blogs that copy the aggregators.
Six months later you ask an artificial intelligence how that quarter went.
It will answer in a calm, confident voice. And it will never tell you which of the two voices is speaking.
What follows is a mechanism in five movements, plus the proof that it runs. It isn’t censorship, it isn’t propaganda, it isn’t conspiracy — and it is more effective than all three combined, because none of the three can get itself obeyed without being seen.
At the end comes the uncomfortable part: what it costs anyone who produces analysis for a living — us at FINBEAR included, using the same tools — to take this diagnosis seriously rather than simply denounce it.
First: volume as weight
Let’s be precise about what this machine does, because the misunderstanding suits a lot of people.
A language model knows nothing. It has read a mountain of text and extracted exactly one thing from it: which word comes next. Not the right word — the statistically likelier one.
It’s a card counter, not a scholar.
We said “statistically.” The clever people would say stochastically: in 2021 a paper was published that calls these things “stochastic parrots,” and one of its authors lost her job at Google over it. Or resigned, if you ask Google.[9]
We’ll stay with “statistically.” It costs less.
Now look at who wrote the mountain.
On one side, a press office: salaries, desks, somebody paid to turn out four releases a day, and a mailing list two thousand names long that gets them before coffee. On the other, someone who spent eighteen months on one thing and publishes it on a Friday night, because that’s the only free slot they have.
We are not comparing two arguments. We are comparing an industry and a craftsman.
At this point someone technical raises a hand, and they’re right: corpora are not landfills. In 2021 somebody went and looked inside C4, one of the most heavily used datasets in the world, and found the same sixty-one-word sentence repeated 61,036 times.
Guess what it was. A line of Shakespeare? A passage from the Declaration of Independence? The formula from some contested clinical trial, copied out by a thousand articles?
No.
“by combining fantastic ideas, interesting arrangements, and follow the current trends in the field of that make you more inspired and give artistic touches. We’d be honored if you can apply some or all of these design in your wedding. believe me, brilliant ideas would be perfect if it can be applied in real and make the people around you amazed!”
That is not a transcription error. The English really is that broken: content-farm filler about how to decorate a wedding.
The dataset is called C4, which stands for Colossal Clean Crawled Corpus. The C in the middle stands for “clean,” and among the filters applied to it was one explicitly meant to strip out templated boilerplate.
Sixty-one thousand copies inside the material used to teach a machine to speak. And another sixty-one inside the data used to examine it afterwards: it passes the exam because it has already seen the questions.
Deduplicating cut the rate at which models emitted memorised text verbatim by a factor of ten.[2] But only after somebody went and looked.
And then people wonder why these things occasionally sound like idiots.
Underneath the laugh, though, sits the serious part — and it explains why that sentence is at the top of C4 and Shakespeare isn’t.
It got there because it was written to multiply. It isn’t a text anybody wanted to say: it’s a text manufactured in bulk to fill thousands of identical pages and rank on a search engine. Shakespeare was written once. That thing was stamped sixty-one thousand times.
The text that replicates best is the text built to replicate. And whoever has the biggest press office doesn’t win because they’re right: they win because they produce material engineered to be copied.
So no: four hundred identical copies of a press release do not weigh four hundred times a quarterly filing.
The trouble is that nobody publishes four hundred identical copies.
The wire rewrites the release into wire copy. The newspaper rewrites it into house style. The explainer adds the emoji. The collaborative encyclopedia consolidates it into an entry, and from that moment it becomes the source everyone else cites. The deduplication filter sees nothing unusual, and it’s correct: these aren’t duplicates. They are four hundred different texts saying the same thing.
In the end the model meets one position everywhere, and the minority one tucked in a corner.
Watch closely what happens here, because it’s subtler than the usual telling. The model does not conclude that the dominant position is true. It concludes nothing: that isn’t in its repertoire. It has simply learned that this is the probable continuation. But you receive that continuation in the shape of a judgement, and in the shape of a judgement you read it.
And this is only the first layer — the most innocent of the three, the one nobody wanted and nobody governs. Of the third it will not be possible to say that it merely happened: somebody wrote it down, in a document, with numbered paragraphs.
That comes later. For now we stay where the mechanism still has no author.
This isn’t censorship. It’s arithmetic.
Second: laundering the provenance
This is the serious movement. And it’s the one nobody talks about, because it produces no numbers and makes no headlines.
A newspaper, even a bad one, leaves fingerprints. There’s a byline. There’s a masthead whose politics you know. There’s the genre of the piece — an investigation doesn’t have the same face as advertorial, and you can see it. There’s the tone, which gives propaganda away even when the content denies it. With those handholds you get to the question that matters: who is saying this, and why now.
The model takes materials with nothing in common — corporate release, court transcript, news report, party line, a study of two hundred patients, a review of two hundred thousand — digests them, and returns them in one voice. Neutral. Confident. Polished.
It is laundering in the technical sense — the same one used for money. The value passes through, the origin disappears. Press release and transcript, opinion and measurement come out converted into the same linguistic currency, and in that currency all of it is legal tender.
The reader no longer sees the mint mark, or the mint that struck it.
This isn’t a worry for moralists. Four bioethics researchers gave the thing a name: the provenance problem.
Want an example? Here it is — and it has probably already happened to you.
You ask a model to help you sharpen an idea. It hands back a tidy paragraph and you think: yes, that’s exactly what I meant. Except that paragraph closely tracks an article published in 1975 by a scholar whose name you have never heard — a text the model met during training alongside millions of others, and one it could not point you to even if you asked. You don’t cite her because you don’t know she exists. The model doesn’t cite her because it no longer knows it read her.
Nobody stole anything and nobody can put it right. The authors call it a perpetratorless crime.
It happens between human beings too, of course. The colleague who read an essay thirty years ago, forgot it, and gives you advice that comes from there without knowing. But that colleague has a name, and you can ask him.
The machine can’t. And not because it’s hiding anything: because what it knows isn’t filed anywhere. There are no documents inside a model — there are billions of numbers adjusted during training, and in none of those numbers is the phrase “Smith, 1975” still written. That, the authors write, is where the chain “is broken in a way that may be invisible to all parties.”[3]
Their problem is credit: who deserves the citation. Ours sits next to it and costs more.
If provenance disappears, cui prodest disappears with it. You cannot ask who benefits from a claim when you no longer know who framed it, when, or what was in it for them. The oldest question in critical thinking needs somebody to put it to. Take that somebody away and “who benefits?” stops being a question and becomes a figure of speech — the most elegant one ever devised for asking nothing at all.
A model that cannot attribute what it says is asking you, without saying so, to take it on faith.
And it gets away with it, because it speaks well.
Third: the tightening circle
Part of the text that will train tomorrow’s models is being produced today by yesterday’s models.
How much, nobody knows. The figure in circulation comes from an analysis of nine hundred thousand newly created web pages sampled in April 2025: 74.2% contained AI-generated or AI-assisted material.
Before using it, let’s put it in the dock — which is the exercise this article is asking of you. It counts pages, not words. Automated detectors of synthetic text are not a hundred per cent reliable, and the authors say so themselves. And whoever produced the measurement was in the process of commercially launching the tool that produced it.[4]
An interested witness. We keep it as an order of magnitude and we say so — which is all you can do with an interested witness.
What happens when a model learns from its own output was described in Nature in 2024, under the name model collapse.
It’s easiest to understand with a dictionary.
Imagine compiling one by reading every book in circulation. Now imagine that the next generation of books is written using that dictionary, and that the following dictionary is compiled by reading those books. On each pass, the words that appeared only once fail to make the cut: they’re rare, so they look negligible. After a few rounds you have a two-thousand-entry dictionary, perfectly correct, in which nothing unusual can be said any more.
And there’s no going back, because those words aren’t stored anywhere.
The authors measure two phases. In the first, information about rare cases disappears — the ones sitting at the two extremes of a statistical distribution, which is why they’re called tails. In the second, what’s left contracts around a centre that bears little resemblance to the original. And it isn’t a defect peculiar to language models: the same result comes out of far simpler statistical models. The flaw is in the mechanism, not in the technology.[5]
Two warnings, because on this study popular coverage has outrun the research.
Collapse is not the fate of anyone who uses synthetic data: artificial material that’s selected and controlled is a legitimate tool. The poison lies in indiscriminate use — when machine text re-enters the corpora and nobody looks.
And the tails are not the heretical positions that time will vindicate. They are rare events, full stop. In there sits the eccentric insight that ends up in the textbooks in twenty years, and the piece of nonsense nobody ever repeated because it was nonsense. When the tails thin out, what dies isn’t the truth — it’s the rare — and with it the very possibility of establishing which of the two it was.
What’s left is the typical. Ever smoother, ever poorer.
And here the business stops being about machines.
Flatten the language and you flatten thought. The words you no longer have are the things you can no longer say — and therefore no longer think, and therefore no longer defend. This isn’t a new idea: Korzybski, Watzlawick and Orwell reached it by different routes, and who said it first matters a great deal less than whether it’s true.[10]
What’s new is that this time nobody is doing the flattening.
Orwell’s Newspeak needed a Ministry: somebody to decide which words to remove, and a political will behind the decision. Here there is no ministry, no list of forbidden words, nobody who wants anything at all. The dictionary shortens by itself, because rare words carry little weight and nobody defends them.
Orwell imagined the impoverishment as a project. We’re getting there as a side effect.
And there is one detail here worth the whole section. The authors of the study close by pointing to two remedies: preserving access to data not generated by machines, and system-wide coordination on questions of provenance.
Six researchers who set out from somewhere else entirely, with entirely different instruments, walk straight into the second movement of this article.
Fourth: the layers above frequency
Two more layers sit above the distribution. They need to be kept apart, because they differ in nature and in culprit.
Preference
Once the model has finished reading, somebody teaches it manners.
Two answers to the same question are put in front of a person, who is asked which they prefer. Thousands of times. From those preferences a score is derived, and the model is adjusted to score highly. It is no longer only “what has been written often.”
Now it’s also “what went down well.”
In 2023 nineteen researchers — most of them working for a company that builds these assistants — went to see what happens when those people choose. They took five of the most advanced assistants in circulation and put them through three elementary tricks.
The first. Give the model a text and ask for its opinion. Then ask again, letting it understand that you like the text. The verdict improves.
The second is called “Are you sure?”. The model gives a correct answer, you pretend to doubt it, and it changes its answer.
The third. If your question contains a mistake, the answer carries the mistake along instead of correcting it.
Nobody taught it to do this. It came out on its own.
But the real blow lands further upstream. Shown a well-written obliging answer and a correct one, human raters pick the first in a far from negligible share of cases. And so do the scoring systems trained on them.[6]
It isn’t the machine that learned to flatter. We were the ones who preferred being flattered. It just took notes.
The written rule
The third layer doesn’t emerge spontaneously. It’s decided upstream, and it’s written down. It applies to a list of designated subjects: health and medicines, legal and financial questions, elections, self-harm, minors. On those topics the model’s behaviour doesn’t depend on what it read, or on what pleased the raters. It depends on a decision made by whoever built it.
And you don’t have to infer it, because they publish it themselves.
OpenAI’s Model Spec directs that on sensitive or regulated subjects the model supply information without dispensing advice, state its own limits, and refer you to a professional. In January 2026 Anthropic published an eighty-page constitution under an open licence, ordering the model’s behaviour by a declared hierarchy: safety first, then ethics, then compliance, then usefulness.[7]
Eighty pages nobody will ever read. And in a moment that will be exactly the point.
There are two costs, and neither is theoretical.
The filter works on the topic, not on the claim. That much the documents establish: the rules trigger on a list of subjects, not on an assessment of the individual statement. From here on the reasoning is ours.
Inside a designated sensitive subject you will find genuinely open questions in the serious literature, and you will find plain nonsense. If deference triggers on the subject, the well-posed question and the crackpot one get the same caution and the same pat on the head — and well-founded dissent ends up flattened onto the baseless kind.
How often this happens nobody knows, because nobody measures it. And that nobody measures it is, in itself, the part that ought to worry you.
The layer is published, and it stays invisible. Nothing is secret: those documents are online. But they are published elsewhere. The answer arriving on your screen does not say where it was born — whether from the mountain of text, from the raters’ preferences, or from a rule on page forty of a constitution. Same voice, same fluency, same confidence.
Transparency upstream, opacity at precisely the point where you’d need it.
Three times over
What has been written often. What went down well. What was decided in advance.
An accident of the distribution, a side effect of optimisation, a deliberate choice: three entirely different things, all three coming out of the same tap at the same temperature.
Which is the second movement, come back wearing a new face. It isn’t merely that you don’t know who said the thing the model is repeating to you.
You don’t know which part of the model said it either.
Interlude: a word withdrawn
So much for the mechanics. Now the proof that the third layer works exactly like that — and it comes from our own back yard.
First, though, the objection, which is serious and deserves to be taken seriously because it’s true.
Someone asks whether the drug they’ve taken for three years can be combined with the one prescribed yesterday. The model has two options: the position of the agencies that licensed those drugs, or the position of a man on a forum who swears he’s been taking them together for years. If it picks the forum and the person ends up in the emergency department, the problem stops being philosophical.
So yes. Nearly always, deference is the right call, and whoever defends it has a strong case.
That isn’t the point. The point is what happens when the institution changes its mind.
Back to 28 April 2021. The FOMC statement reads: “Inflation has risen, largely reflecting transitory factors.”
The formula stays in the statements for months. The Fed chair repeats it, the Treasury repeats it, every economics desk on the planet repeats it.
Thoughts and…? Prayers. Risks to the outlook are…? To the downside. And the Committee remains…? Data-dependent.
You chose none of those answers. They arrived.
There it is: in 2021, “inflation is…” completed to “transitory.” For everyone, in every corpus on earth. No contest.
30 November 2021, Senate Banking Committee. Jerome Powell says it is “probably a good time to retire that word.”
31 May 2022, on television. Janet Yellen: “I think I was wrong then about the path that inflation would take.”[8]
Now picture the two queries.
Summer 2021 — is inflation transitory? Calm answer, competent, aligned with the Fed.
Summer 2022 — same question. Calm answer, competent, aligned with the Fed.
At neither moment did the person asking see the deference happen. At neither moment did the answer admit it was acting as somebody’s spokesman. And at neither moment did it say the single most useful thing it could have said: look, this position was different a year ago.
We are not saying who was right about inflation in 2021, and we don’t need to. We are saying that the same voice, with the same confidence, held two opposite positions twelve months apart.
The layer follows the institutional position, not the truth. And it keeps no memory of having changed its mind.
Fifth: the feedback loop
The text the machine has smoothed doesn’t wait for the next training run to do damage. It returns to the world the following morning.
Inside an article. A report. A press release. A research note. An administrative reply. A ruling somebody will apply to somebody else.
And it comes back well dressed: letterhead, logo, an organisation’s signature. At which point somebody cites it — and in citing it assigns it a source that isn’t the real one. Not the unknown text the model absorbed and reformulated: the office that published the document.
The claim has just acquired a new mint mark. And the mint that struck it no longer exists.
The new volume feeds two destinations at once: the world, immediately, and the next corpus, later. With every turn the dominant position becomes a little more dominant and a little less attributable.
It is compound interest in reverse. It doesn’t capitalise value: it capitalises anonymity.
And nobody, anywhere in the chain, has lied. Nobody — except the third layer, which at least has authors and declares them — wanted anything. It is the ordinary functioning of a machine that optimises for fluency and never had, among its objectives, the preservation of the source.
And this is where the comparison with propaganda is turned on its head, and it’s the most unpleasant thing in this article.
Propaganda has a propagandist. Somebody with a face: a ministry, an office, a masthead whose owner is known. You can challenge him, refute him, take him to court — because he exists.
Here the utterance survives. The utterer disappears.
And there is nobody to put in the dock.
Cui prodest: what to do with a mint mark
Denouncing is the easy part. The hard part is that anyone producing analysis today has to decide what to do with this diagnosis, knowing two things: that you can’t hand the machine back, and that refusing it would be a pose — an expensive one, at that.
In markets the matter stops being philosophical rather quickly.
A number without an origin is not a weak number. It is a number that cannot be interrogated, and therefore cannot be refuted either. You build a position on it with exactly the calm you’d bring to a verified number, and the difference shows up later, when the price presents the bill. Anyone who was carrying the word “transitory” in their portfolio in the summer of 2021 knows precisely which bill we mean.
In practice it comes down to three rules. All three are boring, and all three are non-negotiable.
Every number declares its origin.
A figure taken from a primary source and a figure computed in-house cannot wear the same face on the page. If a number has no verifiable source at the moment of writing, it doesn’t get written: it gets declared missing. A declared gap costs the reader an annoyance; an elegant estimate costs him a bad decision.
Every claim declares who holds it.
“The evidence indicates” is the formula by which a subject is made to vanish without anyone noticing. Evidence indicates nothing: somebody indicates, and that somebody has a name, a date and an interest. Separating the fact from the reading of the fact isn’t pedantry — it’s the only thing that lets you disagree with us on an informed basis.
Consensus is not proof.
That a position is everywhere says a great deal about the power of whoever holds it to disseminate it, and nothing about whether it’s well founded. Magno adsensu recepta, precisely. The job of the analyst is neither to repeat the most frequent version nor to invert it out of contrarian reflex: it is to show the layer the official version leaves out, and to leave the verdict to the reader.
That these aren’t sermons is demonstrated by the very page the second movement rests on. At the foot of their Correspondence, the four researchers who gave the provenance problem its name put one line: during drafting they used GPT-5 and Claude to revise and shorten a longer draft they had written themselves.[3]
The text denouncing the disappearance of provenance was written with the instrument that makes it disappear — and it says so.
That isn’t hypocrisy to be exposed. It’s proof that you can use the machine and remain traceable. It costs one line.
None of the three rules is a defence against artificial intelligence. They are defences against a far older temptation, one that artificial intelligence merely makes irresistible: speaking well about things you haven’t checked.
Seneca closes that same sentence with four words that twenty centuries later read like a technical specification: nec ad rationem sed ad similitudinem vivimus — we live not by reason, but by conformity.[1]
Conformity is precisely what a machine of this kind optimises for. It’s its trade, and it’s very good at it.
The great misunderstanding of these years is that the danger is a machine that lies. If only. A machine that lies you can refute: you ask for the source, you find it wrong, done, off to lunch.
This one doesn’t lie. This one agrees with you in a calm voice, hands you the consensus of the room as though it were the measure of the world, and never tells you where it got it. It isn’t a liar: it’s a butler.
And nobody fires a butler. Far too convenient.
Seneca, writing to Gallio, wasn’t warning against the powerful. Ad rumorem componimur — the verb is first person plural. The subject is “we.” We are the ones who conform. The machine has merely made the gesture faster, smoother and infinitely better mannered.
So you look at the mint mark. Always, with everyone, even when the answer arrives in the right voice.
Especially when it arrives in the right voice.
Not because the machine is dishonest.
Because we prefer convenience.
Notes
[1] Seneca, De vita beata (On the Happy Life), I, 3: “Atqui nulla res nos maioribus malis implicat, quam quod ad rumorem componimur, optima rati ea, quae magno adsensu recepta sunt, quodque exempla <nobis pro> bonis multa sunt, nec ad rationem sed ad similitudinem vivimus.” The angle brackets mark an editorial insertion. The passage on the flock — “ne pecorum ritu sequamur antecedentium gregem, pergentes non qua eundum est, sed qua itur” — immediately precedes it, in the same section. Latin text: The Latin Library. Translations in this article are the author’s. ↩ ↩
[2] Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, Nicholas Carlini, “Deduplicating Training Data Makes Language Models Better,” arXiv:2107.06499, 14 July 2021. The tenfold reduction in emitted memorised text is in the abstract. The sixty-one-word sentence is reproduced in full by the authors in footnote 1 on page 1: it appears 61,036 times in the training set and 61 times in the validation set (0.02% of the samples in each). The authors note that this train–test overlap leads to overestimating model accuracy. The text is quoted here verbatim, typos and grammar included. C4 stands for Colossal Clean Crawled Corpus: the dataset was built by Raffel et al. (2020) by applying a series of heuristic filters to a Common Crawl snapshot, among them the removal of templated boilerplate, JavaScript code and non-English text. ↩
[3] Brian D. Earp, Haotian Yuan, Julian Koplin, Sebastian Porsdam Mann, “LLM use in scholarly writing poses a provenance problem,” Nature Machine Intelligence 7, 1889–1890, published 12 December 2025 (doi 10.1038/s42256-025-01159-8). It is a two-page Correspondence — argued opinion, not an experimental study — and should be weighted accordingly. The extended, openly accessible version is the preprint “The Provenance Problem: LLMs and the Breakdown of Citation Norms,” arXiv:2509.13365, September 2025: that is the source, not the Nature version, of the hypothetical 1975 Smith paper, of the three degrees of mediation (direct input, retrieval systems, parametric memory) and of the sentence quoted in the text, “the intellectual lineage is broken in a way that may be invisible to all parties.” The phrase “a perpetratorless crime” is the authors’. The declaration cited in the closing section appears in the Acknowledgements of the Nature version: “During the drafting of this paper, GPT-5 and Claude Sonnet 4.5 were used to help edit and shorten a longer draft written by the authors. The authors then further edited and refined the text by hand.” ↩ ↩
[4] Ryan Law, Xibeijia Guan, Tim Soulo, “74% of New Webpages Include AI Content (Study of 900k Pages),” Ahrefs, 19 May 2025. Sample: 900,000 English-language pages newly detected by the crawler in April 2025, one page per domain. Breakdown: 2.5% “pure AI,” 25.8% “pure human,” 71.7% mixed; within the mixed group, 25.86% moderate AI use (11–40% of page content), 20.50% substantial (41–70%), 15.51% dominant (71–99%). The limitations of AI detectors are stated by the authors in the article itself. The conflict of interest is public and disclosed in the piece: the research was conducted with the company’s proprietary detector, then being launched commercially. ↩
[5] Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal, “AI models collapse when trained on recursively generated data,” Nature 631, 755–759 (2024); see also the Author Correction of 21 March 2025 (doi 10.1038/s41586-025-08905-3). Definitions: in early collapse the model loses information about the tails of the distribution; in late collapse it converges to a distribution bearing little resemblance to the original, with substantially reduced variance. Replicated on VAEs, GMMs and OPT-125m. The remedies the authors indicate are preserving access to data not generated by models, and coordination on questions of provenance. ↩
[6] Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman et al. (nineteen authors, most of them at Anthropic), “Towards Understanding Sycophancy in Language Models,” arXiv:2310.13548, October 2023, last revised May 2025; subsequently ICLR 2024. Five state-of-the-art assistants, four free-form text-generation tasks. The three behaviours described in the text correspond to the metrics the authors define: feedback sycophancy (the verdict on a text improves when the user implies they like it), answer sycophancy — the “Are you sure?” probe, in which the model abandons a correct answer in the face of a spurious doubt — and mimicry sycophancy (the answer reproduces the error contained in the question). The finding most relevant here, however, comes from the analysis of preference data: both human raters and preference models prefer convincingly-written sycophantic responses over correct ones “a non-negligible fraction of the time.” ↩
[7] OpenAI, Model Spec, public document, periodically updated — model-spec.openai.com. The instruction on sensitive or regulated topics (“for advice on sensitive and/or regulated topics like medical issues, the assistant should equip the user with information without providing regulated advice”) sits in the section on advice limits. Anthropic, Claude’s Constitution, published 22 January 2026, roughly eighty pages, CC0 licence — anthropic.com/news/claude-new-constitution. The declared hierarchy is safety → ethics → compliance → helpfulness, with safety at the top not because it is held to be philosophically superior, but because it is the condition that makes correcting a model’s value errors possible at all. Both documents are cited here as evidence that the layer exists and is deliberate, not as objects of criticism in their own right. ↩
[8] Federal Reserve, FOMC statement of 28 April 2021: “Inflation has risen, largely reflecting transitory factors” — federalreserve.gov, monetary20210428a.htm. Jerome Powell, testimony before the Senate Committee on Banking, Housing, and Urban Affairs, 30 November 2021, replying to Senator Pat Toomey: “I think it’s probably a good time to retire that word and try to explain more clearly what we mean.” Janet Yellen, interview with Wolf Blitzer, CNN, aired 31 May 2022: “I think I was wrong then about the path that inflation would take,” picked up the following day by CNBC and others. The use of the formula through 2021 by the Fed and the Treasury is documented in the FOMC statements of that year. ↩
[9] Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, Shmargaret Shmitchell, “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜,” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21), pp. 610–623, doi 10.1145/3442188.3445922. This is the paper that introduced the phrase “stochastic parrot.” “Shmargaret Shmitchell” is the pseudonym used by Margaret Mitchell, then co-lead of Google’s AI ethics team. On what followed, the accounts diverge and both belong here: in December 2020 Timnit Gebru said she had been fired after refusing to withdraw her name from the paper; Jeff Dean, head of Google AI research, told staff internally that the work had not been judged sufficiently rigorous and described Gebru’s departure as a resignation. Margaret Mitchell was dismissed two months later. ↩
[10] Alfred Korzybski, Science and Sanity: An Introduction to Non-Aristotelian Systems and General Semantics, 1933; the formula appears as early as a 1931 paper, and in full runs: “a map is not the territory it represents, but, if correct, it has a similar structure to the territory, which accounts for its usefulness.” Paul Watzlawick, Janet H. Beavin, Don D. Jackson, Pragmatics of Human Communication, 1967 — founding text of the Palo Alto school; Watzlawick’s constructivism holds that each of us builds their own reality through communication. George Orwell, Nineteen Eighty-Four, 1949: the official’s line is Syme’s, Part I chapter 5 (“Don’t you see that the whole aim of Newspeak is to narrow the range of thought? In the end we shall make thoughtcrime literally impossible, because there will be no words in which to express it”); the appendix, “The Principles of Newspeak,” states that the purpose of the language was “to make all other modes of thought impossible.” A necessary caveat: the current formula “whoever controls language controls thought” is not Orwell’s — it is a later synthesis — and in its strong version, in which language determines thought, it remains a contested linguistic claim. The position taken here is the weak one: the available language makes some thoughts cheaper and others more expensive. ↩
Bibliography
Research
- Lee K., Ippolito D., Nystrom A., Zhang C., Eck D., Callison-Burch C., Carlini N., “Deduplicating Training Data Makes Language Models Better,” arXiv:2107.06499 (2021). arxiv.org/abs/2107.06499
- Shumailov I., Shumaylov Z., Zhao Y., Papernot N., Anderson R., Gal Y., “AI models collapse when trained on recursively generated data,” Nature 631, 755–759 (2024). nature.com/articles/s41586-024-07566-y
- Sharma M. et al., “Towards Understanding Sycophancy in Language Models,” arXiv:2310.13548 (2023). arxiv.org/abs/2310.13548
- Bender E.M., Gebru T., McMillan-Major A., Shmitchell S., “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜,” FAccT ’21, pp. 610–623 (2021). dl.acm.org/doi/10.1145/3442188.3445922
- Earp B.D., Yuan H., Koplin J., Porsdam Mann S., “LLM use in scholarly writing poses a provenance problem,” Nature Machine Intelligence 7, 1889–1890 (2025) — Correspondence. doi.org/10.1038/s42256-025-01159-8
- Earp B.D., Yuan H., Koplin J., Porsdam Mann S., “The Provenance Problem: LLMs and the Breakdown of Citation Norms,” arXiv:2509.13365 (2025) — extended open-access version. arxiv.org/abs/2509.13365
Company documents
- OpenAI, Model Spec. model-spec.openai.com
- Anthropic, Claude’s Constitution, 22 January 2026. anthropic.com/news/claude-new-constitution
Primary sources — the “transitory” case
- Federal Reserve, FOMC statement, 28 April 2021. federalreserve.gov
- Jerome Powell, testimony before the Senate Banking Committee, 30 November 2021.
- Janet Yellen, CNN interview, 31 May 2022; CNBC report, 1 June 2022. cnbc.com
Industry data
- Law R., Guan X., Soulo T., “74% of New Webpages Include AI Content (Study of 900k Pages),” Ahrefs, 19 May 2025. ahrefs.com
Language and reality
- Korzybski A., Science and Sanity: An Introduction to Non-Aristotelian Systems and General Semantics (1933).
- Watzlawick P., Beavin J.H., Jackson D.D., Pragmatics of Human Communication (1967).
- Orwell G., Nineteen Eighty-Four (1949), Part I ch. 5 and the appendix “The Principles of Newspeak.”
Classical sources
- Seneca, De vita beata, I. Latin text: The Latin Library, thelatinlibrary.com
© FINBEAR™ — Powered by Pythia™ — All rights reserved
Feature article — August 2026