You as Training Data: What Language Models Repeat About Your Project
Your public docs are training data. A model reads them, compresses them into a lossy summary weighted by whatever gets said most, and hands that summary out as your first impression, before anyone reaches your site.
Someone is about to look you up, and increasingly they do it by asking an assistant “who is X” or “is X any good” and reading the paragraph that comes back, rather than scanning a page of blue links and weighing the sources themselves. That paragraph is the whole encounter. No competing pages sit behind it to check against, no ranked sources to weigh, just one confident summary that reads like a verdict.
The paragraph was not looked up. It was reconstructed from a compressed statistical model of everything written near your name, and that reconstruction has a specific, describable shape. So you can look at your own public record the way the model does, and see roughly what it will say about you before it says it.
Two ways a model learns your name
A large language model, the kind behind ChatGPT, Claude, Gemini, or Copilot, knows about an entity through two channels worth keeping apart.
The first is parametric memory: the patterns baked into the model’s weights during training, from a corpus that scraped a large slice of the public web. If your name appeared there often enough, in consistent enough contexts, some compressed impression of you sits in the weights. Nothing is stored as a fact you could look up. What is stored is a statistical tendency, its sense of which words tend to follow your name.
The second is retrieval-augmented generation, RAG for short: the answer engine runs a live search when you ask, pulls back a handful of documents, and writes its answer from those. Retrieval was introduced to fix the memory problem, to ground an answer in real fetched text and cite it instead of confabulating from the weights (Lewis et al., 2020). An “answer engine” is any of these search-plus-summary systems, Perplexity and Google’s AI Overviews being the clearest examples, where the product is the synthesized answer rather than a list of results.
Both channels compress. The weights average your public presence into a tendency; retrieval picks a few passages out of everything and lets the model stitch them into prose. Neither hands back your record. Both hand back a lossy version of it, and the losses are not random.
The compression has a direction
Ted Chiang gave the memory version of this its lasting image. “Think of ChatGPT as a blurry jpeg of all the text on the Web,” he wrote in early 2023. It keeps much of the information “but, if you’re looking for an exact sequence of bits, you won’t find it; all you will ever get is an approximation” (Chiang, 2023). The mechanism he names is interpolation: with no exact record, the model fills the gap by averaging its nearest neighbors, the way a compression algorithm reconstructs a patch of sky it did not store. His essay opens with a 2013 Xerox bug where photocopiers silently altered the numbers on a floor plan, changing room sizes with no flag that anything had been approximated. That is the failure mode exactly: a confident output with the compression damage invisible inside it.
The direction of the loss is what matters for reputation. Compression keeps the frequent and drops the rare. Whatever framing appears most often near your name, in the most authoritative-sounding sources, survives the squeeze. The minority characterization, the contested detail, the dissenting read, sits in the thin tail of the distribution, and the tail is the first thing compression discards.
OpenAI’s own researchers pinned the frequency effect to something concrete in September 2025. Their paper treats hallucination, a model producing a fluent statement that is simply false, as a structural byproduct of how these systems are graded. Nearly every benchmark scores a wrong answer and an “I don’t know” identically, so a model that guesses beats a model that abstains, the way a student guesses on an unpenalized exam. Hallucinations, they write, “persist due to the way most evaluations are graded, language models are optimized to be good test-takers, and guessing when uncertain improves test performance” (Kalai et al., 2025). The rate is highest exactly where the training data is thinnest. Their own example is a birthday the model has seen too rarely to store, so it guesses one, fluently and wrong.
For a well-documented entity, a public company or a famous person, the compression is roughly accurate, because the frequent framing and the true framing are the same thing. When your public record is thin or conflicting, a small studio, a new project, a person with a common name, retrieval comes back sparse, the model leans on a parametric guess, and it emits the highest-probability adjacent pattern with the fluency it would use for a fact. Dahl et al. found exactly this in the legal domain, where hallucination climbs as the cases get more obscure (Dahl et al., 2024). You get a smooth, confident, wrong answer, and nothing in its tone tells the reader it is the model filling in sky.
Grounding helps, and it is nowhere near enough
The objection is that answer engines already fetch live sources, so this is solved for anyone whose site is up. The measured record from 2024 through 2026 says otherwise.
The Tow Center for Digital Journalism tested eight generative search tools against 1,600 queries where the correct source was known. Collectively they returned wrong source information on more than sixty percent of them. Perplexity did best and still erred on 37 percent; Grok-3 was wrong 94 percent of the time. The paid tiers were more confidently wrong than the free ones, and several tools cited copied or syndicated versions over the original source (Jaźwińska & Chandrasekar, 2025). Grounding was switched on for all of it.
The legal-tooling numbers are sharper, because there the vendors sell accuracy as the product. Stanford’s RegLab and HAI tested purpose-built legal research tools with retrieval, marketed with claims like “hallucination-free” citations. In their words, the tools “each hallucinate between 17% and 33% of the time” (Magesh et al., 2025). Between one in six and one in three answers carried a false claim, in a domain where a fabricated citation ends a career, from teams with every resource to get it right. The researchers’ verdict on the marketing was blunt: “We demonstrate that the providers’ claims are overstated.”
Grounding fails for reasons that stack. Retrieved passages can be on-topic without containing the answer. Fetched documents can contradict each other, and the model picks or blends. Long context can make it worse: models attend to the start and end of what they are given and neglect the middle, an effect Liu et al. named “lost in the middle,” accuracy dropping by double digits as the relevant passage moves toward the center of a long input, so a model can hold the correct source in its window and still write past it (Liu et al., 2024). When the retrieved evidence about you is thin, the whole condition of a long-tail entity, the model falls back on the parametric guess retrieval was supposed to replace, and it fails quietly, with a citation attached that may not support the sentence it is hung on.
The labs are working the problem: AI Overviews link the pages they draw from (Reid, 2024), Anthropic ties Claude’s claims back to source passages (Anthropic, 2025), and OpenAI’s paper proposes rewarding a calibrated “I don’t know.” That is real progress on the mechanism. The measured gap above already reflects those efforts; it is where the shipping tools stand today.
The framing you wrote gets laundered into fact
This part touches reputation most directly, so I want to hold the demonstrated apart from what I am inferring.
What is demonstrated: models compress toward the dominant framing, and the dominant framing near your name is very often the one you wrote yourself. Your marketing copy, your bio, your about page are usually the densest, most repeated, most authoritative-sounding text about you on the open web, because you published them and others quote them. The model compresses that self-description and re-emits it with the promotional context stripped off. The startup’s “leading platform,” the founder’s own bio adjectives, come back in a flat, sourceless register reading as neutral fact. Your marketing acquires the authority of an impartial summary by passing through a machine the reader trusts to be one.
The same mechanism runs in the other direction, and this is where it stops being a branding curiosity. When the frequent framing near your name is wrong, the compression hands the wrong one out with the same confidence. Martin Bernklau, a German court reporter, had his name appear for years alongside the crimes he covered. In August 2024, Microsoft Copilot compressed that co-occurrence into the obvious wrong pattern and described him as a convicted child abuser and fraudster, printing his real address and phone number beside the accusation (Goodwins, 2024). Jonathan Turley, a law professor, was named by ChatGPT as the accused in a sexual-harassment case that never happened, complete with a citation to a Washington Post article that was never written (Verma & Oremus, 2023). Neither man touched his own public record. The compression assembled a false version from what sat near their names and delivered it in the register of settled fact.
Where I am inferring rather than reporting: I am calling this “framing inheritance,” the dominant characterization getting absorbed and re-emitted as neutral, and treating it as one named phenomenon. No single paper establishes it under that name. I am assembling it from the compression argument, the frequency-and-hallucination results, and the fabrication cases above. Take it as the essay’s own reading of the evidence, a wager about a real pattern rather than a settled finding with its own citation.
Your public surface on the left (site copy, bios, press, forum mentions, name collisions), sized by how often each framing recurs. The model's compression in the middle, keeping the dense center and dropping the thin tail. On the right, the single paragraph it emits, the first impression a reader gets, colored by evidence state per the color-as-evidence method: which claims are verified, which are the model's interpolation.
The loop that hardens it
There is a second-order version, and the direction is clear even where the severity is argued. Once a compressed description of an entity is published, an AI-written summary landing on a blog, a content farm, a scraped forum, it becomes text like any other, and the next model crawls it. Shumailov and colleagues showed in Nature that models trained on their own recursively generated output degrade, with the tails of the distribution vanishing first, an effect they called model collapse (Shumailov et al., 2024). Applied to reputation: the flattened framing of you gets written down, re-crawled, and fed to the next model, so the compression compounds and the minority-but-true details in the tail erode a little further each pass. The caveat is one the research carries itself. Shumailov’s fully-synthetic setup is a worst case, and later work found that mixing fresh real data with synthetic avoids the catastrophic version. The direction holds; the magnitude at web scale is not settled. Watch it as a bending trajectory, the wall still well out of sight.
Can you shape the mirror
Some, at the margins, and most of the market selling this is overpromising.
There is a real mechanism underneath. Answer-engine optimization (AEO, or GEO for generative engine optimization) means shaping your content so a model is more likely to cite or reproduce it. The founding paper tested such tactics against a benchmark of generative-engine queries: adding citations, direct quotations, and statistics, plus cleaner writing, raised a source’s visibility in generated answers by up to forty percent. The sharp finding is what did not work. Old-style keyword stuffing performed about ten percent worse than doing nothing, because the optimization that helps operates at the level of the passage a model wants to quote, not the keywords a crawler once counted (Aggarwal et al., 2024).
Two things get sold under that one name, and the line between them is clean. Writing that is clear, factual, quotable, and easy to ground, an about page with concrete dated facts, real numbers, structure a model can lift a sentence from, is worth doing, and it is mostly good fundamental writing pointed at a new reader. The rest, the “share of model” monitoring dashboards and the once-hyped llms.txt file, is where the overpromising lives. On llms.txt, a proposed standard file that hands a model a clean index of your site (Howard, 2024): its own proposer never pitched it as a ranking lever, and by 2026 no major AI crawler consumes it for retrieval (Spriestersbach, 2025). Ship it as tidy plumbing if you like, but do not pay a premium for it.
Your room to act is capped by law as much as by mechanism. The clearest legal test so far is Walters v. OpenAI: a radio host sued after ChatGPT told a journalist he had been accused of embezzlement, which was false. In May 2025 the court granted OpenAI summary judgment, partly on the ground that no reasonable reader in the journalist’s position would have taken the raw output as an assertion of fact (Walters v. OpenAI, 2025). That is a narrow win, tied to a public-figure plaintiff who never believed the output, and it is not a general shield. It does mark the current posture: the courts give the labs wide latitude, so your practical recourse is a documented, dated correction request to the provider rather than a lawsuit you will win. The problem has case law now, going back to Mata v. Avianca, where lawyers were sanctioned for filing a brief citing six cases ChatGPT invented outright, fake quotes and all (Mata v. Avianca, 2023).
Read your surface the way the model does
The move this desk makes with a contested crypto project applies here directly, which is why the topic sits in our lane. When we audit a launch, we do not ask what the team meant to say. We read the public surface the way a hostile newcomer reads it, and we find the specific places a misreading forms. An answer engine is that newcomer, industrialized, and it reaches your reader first.
So the audit is literal. Query the major engines with “who is X” and the adversarial variants a skeptic or a competitor would type, and read what comes back as data about your compressed record: the specific false or flattened claims, and which sources it cites for each. That output is the reputation the machine is already handing out, drawn from a surface you can change. You cannot rewrite the compression. You can change what sits dense enough near your name to survive it, which means making the true framing the frequent one before the model decides for you.