← The Clair Blog

    How to tell when ChatGPT is making it up (and when to check)

    ChatGPT does not lie, it guesses confidently, which is worse. Here is how to spot when it is making things up, and a simple ladder for how hard to check.

    A note before the steps. The tools and their habits keep shifting, so treat any specific behaviour here as a snapshot. The durable bit is the judgment underneath, match how hard you check to where the answer is going. That part does not age, and it works the same whether you are using ChatGPT, Claude, Gemini or Perplexity.

    You asked ChatGPT a straight question and it gave you a confident, tidy, completely reasonable-sounding answer. A figure, maybe, or a source, or a quote. It felt solid.

    And somewhere at the back of your mind was the quiet worry that you were about to repeat it in a meeting, put it in a deck, send it to someone who would believe you, and you had no real way of knowing if any of it was true.

    That worry is the correct instinct. The problem is that nobody has given you a calm way to act on it, so you either check nothing and hope, or you check everything and lose the time the tool was supposed to save.

    The short answer: ChatGPT does not lie, it guesses, which is worse, because a liar knows the truth and a guesser does not. It predicts the most plausible next words rather than looking anything up, so it cannot tell you when it is wrong and will deliver a fabricated fact with exactly the same confidence as a real one. The skill is not catching every error, it is matching how hard you check to where the answer is going. That is the Trust Ladder, and it is below.

    Does ChatGPT lie, or is it something else?

    It is something else, and the distinction matters. ChatGPT is not deceiving you, because deception needs an intention it does not have. It generates answers by predicting what text should plausibly come next, based on patterns it learned, rather than by checking facts against anything.

    So when it does not know something, it does not stop. It fills the gap with the most likely-sounding version and hands it over without a flicker of doubt.

    This has a name, AI hallucination, and the important thing to understand is that it is not a glitch the next update will fix. It is a feature of how these tools work, baked in, and the researchers building them are fairly open that it will probably never be fully solved. Claude does it. Gemini does it. Perplexity does it.

    The cruel part is that the tool cannot audit itself. It has no internal sense of which of its answers are solid and which it invented, so it will never be the thing that warns you. That job stays yours.

    How to tell when it is making it up

    You cannot catch everything, but the bluffs have tells, and once you know them you spot them fast.

    • It sounds certain and offers nothing to stand on. Genuine answers tend to come with a source, a caveat, or a visible bit of working. A confident assertion with no hedging and nothing to check is the first flag, especially on anything specific.
    • The details are suspiciously precise. A made-up answer often arrives dressed in exact dates, names, page numbers and citations, because specificity is what makes it feel real. Lawyers have been sanctioned for filing AI-written briefs full of cases that sounded perfectly plausible and did not exist, and a newspaper once printed a summer reading list of real authors paired with books they never wrote.
    • The answer changes when you ask again. Because there is no fact database underneath, a tool that is guessing will often contradict itself if you ask the same thing a second time, or push back on the answer. If the story shifts, it never had solid ground to begin with.
    • The errors are wrapped in things that are right. This is the dangerous one. It will get a whole paragraph correct and quietly flip a statistic the wrong way, increase instead of decrease, or swap one name for a similar one.
    • It is worse on the obscure and the specific. The thinner the public information on a topic, the more it guesses, so niche subjects, recent events, and anything about your own company or a small organisation are where it invents most freely.

    Two prompts catch most of it. Before you rely on something, try: "list your sources for this, with links, and flag anything you are not certain about." Then check the links actually exist and say what it claimed. Or simply ask the same question again in a fresh chat and see whether the answer holds.

    The Trust Ladder: how hard to check, based on where it is going

    Here is the part that turns the worry into a habit. You do not verify everything to the same degree, you verify in proportion to where the output is going and what happens if it is wrong. Picture four rungs.

    RungWhat sits hereWhat to do
    Use itBrainstorms, rough drafts, outlines, explaining a concept to yourself, getting unstuck. Stays with you, easy to undoUse it and move on, the cost of an error is almost nothing
    Skim itTone rewrites, reformatting, a summary of something you have already read. Goes out internally, low stakes, but your name is on itRead it once so it sounds like you and nothing has drifted, then send
    Check itA stat, date, name, quote, or a summary of something you have not read yourself. You are going to repeat it or act on itVerify the load-bearing parts against the original source, this is exactly where it is confidently wrong
    Don't send it uncheckedAnything client-facing, board-facing, legal, financial, or public. High stakes, hard to walk backAI drafts, a human approves, no exceptions

    Most of what people get wrong is putting a Rung 4 task on Rung 1 trust, taking an invented figure straight from the chat window into a board pack because it sounded right. The ladder is just the habit of asking, before you trust anything, which rung this actually sits on.

    Where each tool sits

    They are not equally trustworthy on facts, and knowing the rough shape helps you choose. ChatGPT is the confident guesser, brilliant for momentum, happy to invent a statistic to keep you moving. Claude tends to hedge more and reads more carefully, though it will still fabricate a citation if you let it.

    Perplexity shows its sources, which is its whole point, but the sources are sometimes only loosely related to the claim, so you still open them. Gemini is steady on data and structure and weaker on nuance. None of them remove the need to check, they just change how much.

    If you have not picked a main tool yet, our take on ChatGPT vs Claude walks through how to choose, and whether ChatGPT Plus is worth paying for is the next question after that.

    The habit that makes this automatic

    You do not need to memorise any of this. You need one question, asked before you trust an answer: where is this going, and what happens if it is wrong. That single question places the output on the ladder and tells you whether to use it, skim it, check it, or never send it unchecked.

    For how to run the checks themselves, see how to fact-check AI output. For the jobs where the answer is "do not hand it to ChatGPT at all," we wrote a companion piece: the tasks you should never give ChatGPT.

    For the related problem of getting honest pushback instead of empty agreement, see how to get ChatGPT to push back instead of just agreeing. And for the everyday version of all this, the thirty-second scan before anything leaves your screen, see the pre-send check.

    That is the actual skill with AI, and it is not technical. It is knowing the difference between a tool that is right most of the time and a tool you can trust, and never confusing the two when it matters.

    Know when to trust your AI

    The tools won't tell you when they're wrong. Your judgment has to.

    The Trust Ladder is the start of it. The AI Starter Kit builds the rest, setting ChatGPT and Claude up so the output is worth trusting in the first place, and teaching the judgment to know what to lean on and what to check.

    Get the AI Starter Kit

    AI that actually works for you. ChatGPT and Claude.

    Clair helps non-technical professionals know when to trust their AI, when to check it, and when to skip it.

    Common questions

    Does ChatGPT lie?

    No, it guesses. ChatGPT predicts the most plausible next words rather than checking facts, so it can be confidently wrong without knowing it. There is no intention to deceive, which is exactly why it delivers an invented answer with the same certainty as a real one.

    How can I tell if ChatGPT is wrong?

    Watch for confidence with nothing to check, oddly specific details like exact dates or citations, sources you cannot actually find, and answers that change when you ask the same question again. Errors wrapped inside otherwise correct paragraphs are the hardest to catch, so treat anything factual as a first draft.

    What does ChatGPT get wrong most often?

    Statistics, dates, names, quotes, and citations, plus anything niche or about a specific company or small organisation. The thinner the reliable public information on a topic, the more it fills the gaps by guessing.

    Can you trust ChatGPT?

    For low-stakes work like brainstorming, drafting, and reformatting, yes. For anything you will repeat, act on, or send to people who matter, treat it as a confident first draft and check the parts that carry weight. How hard you check should match where the answer is going.

    How do I fact-check AI output?

    Ask it to list its sources with links and verify they exist and say what it claimed, search the key fact independently in a normal search engine, or ask the same question again in a fresh chat and see whether the answer holds. If it shifts, it was never on solid ground.

    The AI Starter Kit

    Set up ChatGPT and Claude properly. Once.

    The workflows, rules and habits that hand the boring recurring work to your AI and give you your evenings back.

    Get the Kit →