A note before the steps. The tools and their habits keep shifting, so treat any specific behaviour here as a snapshot. The durable bit is the judgment underneath, match how hard you check to where the answer is going. That part does not age, and it works the same whether you are using ChatGPT, Claude, Gemini or Perplexity.
You asked ChatGPT a straight question and it gave you a confident, tidy, completely reasonable-sounding answer. A figure, maybe, or a source, or a quote. It felt solid.
And somewhere at the back of your mind was the quiet worry that you were about to repeat it in a meeting, put it in a deck, send it to someone who would believe you, and you had no real way of knowing if any of it was true.
That worry is the correct instinct. The problem is that nobody has given you a calm way to act on it, so you either check nothing and hope, or you check everything and lose the time the tool was supposed to save.
The short answer: ChatGPT does not lie, it guesses, which is worse, because a liar knows the truth and a guesser does not. It predicts the most plausible next words rather than looking anything up, so it cannot tell you when it is wrong and will deliver a fabricated fact with exactly the same confidence as a real one. The skill is not catching every error, it is matching how hard you check to where the answer is going. That is the Trust Ladder, and it is below.
Does ChatGPT lie, or is it something else?
It is something else, and the distinction matters. ChatGPT is not deceiving you, because deception needs an intention it does not have. It generates answers by predicting what text should plausibly come next, based on patterns it learned, rather than by checking facts against anything.
So when it does not know something, it does not stop. It fills the gap with the most likely-sounding version and hands it over without a flicker of doubt.
This has a name, AI hallucination, and the important thing to understand is that it is not a glitch the next update will fix. It is a feature of how these tools work, baked in, and the researchers building them are fairly open that it will probably never be fully solved. Claude does it. Gemini does it. Perplexity does it.
The cruel part is that the tool cannot audit itself. It has no internal sense of which of its answers are solid and which it invented, so it will never be the thing that warns you. That job stays yours.
How to tell when it is making it up
You cannot catch everything, but the bluffs have tells, and once you know them you spot them fast.
- It sounds certain and offers nothing to stand on. Genuine answers tend to come with a source, a caveat, or a visible bit of working. A confident assertion with no hedging and nothing to check is the first flag, especially on anything specific.
- The details are suspiciously precise. A made-up answer often arrives dressed in exact dates, names, page numbers and citations, because specificity is what makes it feel real. Lawyers have been sanctioned for filing AI-written briefs full of cases that sounded perfectly plausible and did not exist, and a newspaper once printed a summer reading list of real authors paired with books they never wrote.
- The answer changes when you ask again. Because there is no fact database underneath, a tool that is guessing will often contradict itself if you ask the same thing a second time, or push back on the answer. If the story shifts, it never had solid ground to begin with.
- The errors are wrapped in things that are right. This is the dangerous one. It will get a whole paragraph correct and quietly flip a statistic the wrong way, increase instead of decrease, or swap one name for a similar one.
- It is worse on the obscure and the specific. The thinner the public information on a topic, the more it guesses, so niche subjects, recent events, and anything about your own company or a small organisation are where it invents most freely.
Two prompts catch most of it. Before you rely on something, try: "list your sources for this, with links, and flag anything you are not certain about." Then check the links actually exist and say what it claimed. Or simply ask the same question again in a fresh chat and see whether the answer holds.
The Trust Ladder: how hard to check, based on where it is going
Here is the part that turns the worry into a habit. You do not verify everything to the same degree, you verify in proportion to where the output is going and what happens if it is wrong. Picture four rungs.
| Rung | What sits here | What to do |
|---|---|---|
| Use it | Brainstorms, rough drafts, outlines, explaining a concept to yourself, getting unstuck. Stays with you, easy to undo | Use it and move on, the cost of an error is almost nothing |
| Skim it | Tone rewrites, reformatting, a summary of something you have already read. Goes out internally, low stakes, but your name is on it | Read it once so it sounds like you and nothing has drifted, then send |
| Check it | A stat, date, name, quote, or a summary of something you have not read yourself. You are going to repeat it or act on it | Verify the load-bearing parts against the original source, this is exactly where it is confidently wrong |
| Don't send it unchecked | Anything client-facing, board-facing, legal, financial, or public. High stakes, hard to walk back | AI drafts, a human approves, no exceptions |
Most of what people get wrong is putting a Rung 4 task on Rung 1 trust, taking an invented figure straight from the chat window into a board pack because it sounded right. The ladder is just the habit of asking, before you trust anything, which rung this actually sits on.
Where each tool sits
They are not equally trustworthy on facts, and knowing the rough shape helps you choose. ChatGPT is the confident guesser, brilliant for momentum, happy to invent a statistic to keep you moving. Claude tends to hedge more and reads more carefully, though it will still fabricate a citation if you let it.
Perplexity shows its sources, which is its whole point, but the sources are sometimes only loosely related to the claim, so you still open them. Gemini is steady on data and structure and weaker on nuance. None of them remove the need to check, they just change how much.
If you have not picked a main tool yet, our take on ChatGPT vs Claude walks through how to choose, and whether ChatGPT Plus is worth paying for is the next question after that.
The habit that makes this automatic
You do not need to memorise any of this. You need one question, asked before you trust an answer: where is this going, and what happens if it is wrong. That single question places the output on the ladder and tells you whether to use it, skim it, check it, or never send it unchecked.
For how to run the checks themselves, see how to fact-check AI output. For the jobs where the answer is "do not hand it to ChatGPT at all," we wrote a companion piece: the tasks you should never give ChatGPT.
For the related problem of getting honest pushback instead of empty agreement, see how to get ChatGPT to push back instead of just agreeing. And for the everyday version of all this, the thirty-second scan before anything leaves your screen, see the pre-send check.
That is the actual skill with AI, and it is not technical. It is knowing the difference between a tool that is right most of the time and a tool you can trust, and never confusing the two when it matters.
Know when to trust your AI
The tools won't tell you when they're wrong. Your judgment has to.
The Trust Ladder is the start of it. The AI Starter Kit builds the rest, setting ChatGPT and Claude up so the output is worth trusting in the first place, and teaching the judgment to know what to lean on and what to check.
Get the AI Starter KitAI that actually works for you. ChatGPT and Claude.
Clair helps non-technical professionals know when to trust their AI, when to check it, and when to skip it.