Ask ChatGPT to describe your company on a Monday and again three weeks later, and you can get two different companies back. The facts drift. The sentiment shifts. Sometimes a competitor gets named first; sometimes a detail is simply wrong. This is not a glitch, and it is not something a single screenshot will ever reveal. It is how the system works, and it is why the brands paying attention have stopped treating AI answers as a thing you check and started treating them as a thing you monitor.
The short version: because the answer changes, a one-time look tells you almost nothing. Build a fixed set of prompts, run them on a schedule, and log four things every time: mention frequency, sentiment, accuracy, and the sources the model cites. That running record is what turns a vague sense that “the AI says something weird about us” into a dated, diagnosable trend you can actually act on.
AI reputation monitoring is the ongoing practice of tracking how assistants like ChatGPT describe a brand, measured across mention rate, sentiment, accuracy, and cited sources over time rather than in a single snapshot. The distinction between over time and in a snapshot is the entire discipline.
Why the answer keeps changing
ChatGPT pulls from two channels that update at completely different speeds. The first is its trained knowledge, frozen at a cutoff date and refreshed only when the underlying model is updated. The second is live web search. According to OpenAI’s help docs, ChatGPT decides on a per-query basis whether a question benefits from the web, then rewrites it into targeted searches and links to what it finds.
Which channel answers depends on the question, and that alone makes outputs unstable. Layer on two more variables and the drift compounds. Models sample their responses probabilistically, so the same prompt asked twice can produce different wording, different emphasis, and occasionally different facts. And personalization means the version of you that one user sees may not match what another sees. A company can be described accurately one week from stale training data, then inaccurately the next week after the model runs a search and lands on a bad source.
The scale is what makes this worth the effort. OpenAI reported more than 900 million weekly active users in early 2026, and by mid-2026 the ChatGPT app had crossed roughly a billion monthly users. When that many people ask an assistant about a category before they ever reach a website, the assistant’s answer is the first impression, and it is one that most companies never see.
Accuracy is the signal that matters most
Presence is easy to celebrate and easy to misread. The harder question is whether the model is right. The evidence says it often is not.
Columbia Journalism Review’s Tow Center for Digital Journalism tested eight generative search tools on their ability to correctly identify the source of real published material. The Tow Center study found the tools answered more than 60 percent of those queries incorrectly. ChatGPT’s search feature was wrong roughly two-thirds of the time. The best performer still missed better than a third. Worse for anyone trying to gauge risk, the tools rarely signaled doubt. They returned confident, wrong answers instead of declining, and the paid tiers were often more confidently wrong than the free ones. Several fabricated links outright or cited copied versions of the original article.
A separate BBC investigation reached the same neighborhood from a different angle. When BBC research asked four leading assistants to summarize 100 of its news articles, journalists judged 51 percent of the answers to contain significant issues. Nineteen percent of answers citing the BBC introduced factual errors in statements, numbers, or dates. Thirteen percent of the quotes attributed to the BBC were altered from the original or did not appear in the cited article at all.
Read those two findings together, and the takeaway for a brand is direct. A model can name your company and still get the story wrong, and it will deliver that error with the same confident tone it uses for the truth. Hallucinated facts are the highest-priority thing a monitoring program exists to catch, because being described inaccurately at scale is worse than not being described at all.
What to measure on every run
The brands that manage this well track a consistent set of signals, so that any change stands out against the baseline. Measure the same fields every time:
- Mention rate: how often ChatGPT names the company across your target prompts.
- Sentiment: whether the description reads positive, neutral, or negative.
- Accuracy: whether the specific claims are correct, scored against known facts.
- Prominence: whether the brand is named first or buried beneath competitors.
- Cited sources: which URLs the model pulls when it searches, since those pages are the leverage points.
The last one deserves emphasis. When ChatGPT searches, the pages it cites are the raw material for its answer. Knowing which sources it trusts for your category tells you exactly where to focus if you want to change what it says.
Setting up monitoring, step by step
Start with a baseline, then repeat it on a cadence. The baseline is the control group every future run gets measured against, so it needs to be built deliberately and written down in full.
- Build a prompt set that mirrors how real buyers ask: a direct brand query, a category or discovery prompt, a competitor comparison, and a problem-led question a customer would actually type.
- Run each prompt several times in fresh sessions, both with and without web search enabled, and record the exact responses with the date attached.
- Log structured fields for every run: the prompt, mention yes or no, sentiment, an accuracy score, position, and any sources cited.
- Set a cadence that fits the stakes: weekly while you are actively improving, monthly once results stabilize, daily for high-stakes or fast-moving categories.
- Watch for drift and diagnose its cause, because a model update, a new piece of web content, and lost media coverage each leave a different fingerprint in the data.
One conversation is not data. Because ChatGPT samples its answers and may or may not search the web on any given run, a single result can mislead in either direction, flattering one day and alarming the next. Run each prompt multiple times before trusting a trend, and never try to force the numbers with thin or spammy content, which both the model and its underlying sources increasingly filter out.
What to do when ChatGPT gets it wrong
You cannot log into ChatGPT and edit what it says. AI answers sit downstream of your web and earned presence, so the fix is to change what the model learns from. That means strengthening the authoritative, consistent, citable material that AI systems retrieve and trust.
The work is unglamorous, and it compounds. Update owned pages so the accurate facts are easy to find and easy to quote. Earn coverage in publications the model already trusts, since a correction that lives only on your own domain carries less weight than the same fact confirmed by an independent outlet. Keep entity data uniform across the web: the company name, the descriptions, the key figures, so the model resolves everything to one coherent identity instead of stitching together conflicting fragments.
This is the discipline Status Labs has refined since AI search began. A Status Labs breakdown of the monitoring method lays out the same protocol the firm runs for clients, and its reputation management whitepaper puts the practice in the wider context of where AI and reputation are heading. For teams that would rather watch the mechanics explained than read them, the firm’s Status Labs videos walk through how AI systems form and repeat brand narratives. While much of the market is still taking one screenshot and guessing, the firms measuring on a schedule are the ones catching problems early enough to fix them.
A working framework
- Build a fixed prompt set that mirrors how real buyers ask.
- Baseline it with multiple runs, fresh sessions, and both search modes.
- Track five signals every time: mention rate, sentiment, accuracy, prominence, and sources.
- Re-run on a set cadence and compare each result against the baseline.
- Diagnose any drift, then fix it at the source with authoritative, consistent content.
