AI

Why I Started Fact-Checking Every AI Answer, Even the Ones I Was Sure About

By Mr. Edmilson · September 22, 2026 · 9 min read

Laptop showing a blurred chat interface next to a notebook, pen, and magnifying glass on a desk

The mistake that changed how I use AI wasn’t a big, obvious one. It was small, confident, and almost invisible, which is exactly why it stuck with me. I’d asked a chatbot for the founding year of a local business I was writing about, a detail I genuinely didn’t know and had no reason to doubt. It gave me a year, stated plainly, no hedging. I wrote it into a draft. A friend who actually knew the owner mentioned, almost as an aside, that the year was wrong by nearly a decade. Not a controversial fact. Not something buried or ambiguous. Just wrong, delivered with the same tone of confidence as everything else in that conversation.

The Specific Thing That Bothered Me

What got under my skin wasn’t that the AI made a mistake. Every source makes mistakes, including people, including reference books, including my own memory. What bothered me was that I had no way to tell, from the answer itself, that this particular claim was any less reliable than the ones sitting right next to it in the same response, which were accurate. There was no verbal hedge, no “I believe” or “I’m not fully certain,” nothing in the phrasing that distinguished the wrong fact from the right ones. It read as uniformly confident, and I’d absorbed that confidence uncritically because nothing in the delivery gave me a reason not to.

That’s a different problem than “AI sometimes gets things wrong,” which I already knew and had already made peace with. The actual problem is that the wrongness doesn’t announce itself. A source that’s unreliable in a consistent, visible way is easy to work around. A source that’s right ninety percent of the time with no signal for which ten percent to doubt is much harder to use safely, because it trains you, gradually and without your noticing, to stop checking.

Several open reference books and a handwritten notebook spread across a wooden desk

What I Actually Changed About How I Use It

I didn’t stop using AI for research, which was my first instinct and, on reflection, an overcorrection. What I changed was more specific: I now treat any AI-generated claim with a concrete, checkable detail, a date, a number, a name, a statistic, as a lead to verify rather than a fact to use. Ideas, structure, phrasing suggestions, brainstorming, none of that needs the same scrutiny, because being wrong about a phrasing suggestion doesn’t put an inaccurate claim into the world. But anything that could end up stated as fact somewhere else, in something I write, something I tell someone, a decision I make, gets checked against an actual source before it leaves the draft stage.

In practice this means I’ve started treating AI chat the way I’d treat a smart, well-read friend who’s occasionally and unpredictably wrong about specifics. I’d still ask that friend for a starting point, a general shape of the answer, a pointer toward where to look. I just wouldn’t quote them in an article without checking first, and I wouldn’t have thought twice about that standard for a person. I’m not sure why it took a wrong founding year to apply the same standard to a chatbot.

The Verification Habit That Actually Sticks

The habit that’s worked best for me is asking a follow-up question in the same conversation: where does that information come from, and how confident is the model in it. This doesn’t always produce something citable, sometimes it just restates the claim with slightly more hedging, but often enough it surfaces something useful, a related fact that contradicts the first one, or an explicit acknowledgment that the detail might be outdated or approximate. That signal, when I get it, is worth more than the original answer, because it tells me where to spend my actual verification time instead of treating every claim as equally solid or equally suspect.

For anything genuinely important, I’ve settled on a simple rule: two independent sources, at least one of which isn’t itself an AI summarizing the same underlying material. That last clause matters more than it sounds like it should. I’ve caught myself “verifying” an AI’s claim by asking a second AI, which tells me almost nothing, since both are plausibly drawing on similar patterns and can be confidently wrong in the same direction. An actual primary source, a original document, a person who was there, a dataset I can look at myself, is the only check that really counts.

Two smartphones side by side on a desk showing different abstract chat conversations

Where This Gets Genuinely Hard

The honest complication is that full verification isn’t always practical, and pretending otherwise would make this whole habit collapse under its own weight. I don’t have time to independently verify every claim in every casual answer I get, and for a lot of low-stakes questions, that level of rigor would be a waste of the very time AI is supposed to save me. The rule I’ve landed on is proportional: the more a claim is going to travel, the more it’s going to end up in something public, something someone else will rely on, the more scrutiny it gets. A fact I’m using to satisfy my own private curiosity gets a lighter check than a fact that’s going into a piece of writing other people will read and trust because my name is on it.

I’ve also had to accept that this makes AI genuinely slower to use for research than it initially appears, and that the speed I thought I was gaining was partly illusory, borrowed against verification work I was skipping. That’s not a reason to abandon it. It’s a reason to be honest about what it’s actually good for, which turns out to be narrower and more specific than “answering questions,” and closer to “generating a fast first draft of an answer that still needs the same scrutiny any first draft deserves.”

What I’d Tell Someone Starting to Use AI for Research

Notice the tone, but don’t trust it. Confident phrasing and accurate phrasing look and sound identical, and that’s not a bug that’s going to get fixed by a better prompt, it’s closer to a structural feature of how these systems generate language. Treat every declarative sentence with a specific fact in it as a claim to verify, not a fact already established, regardless of how certain it sounds.

Build the verification step into your process from the start rather than trying to retrofit it later, because retrofitting it is exactly what happened to me, and it took an actual public mistake to force the change. It’s much easier to build the habit of checking before something matters than to build it after something’s already gone out wrong. And keep a mental note of the specific moment that changed your relationship to these tools, if you have one, because mine, a small wrong number in a low-stakes blog post, is the thing that’s kept the habit sticky months later, long after the initial embarrassment faded.

Where I’ve Landed

I still ask AI tools things constantly, probably more than I did before this happened, because once I stopped treating every answer as equally reliable, I got more comfortable using them for the parts they’re actually good at. The founding-year mistake didn’t make me trust these tools less in some vague, general way. It made me trust them more precisely, in the specific places where that trust is actually earned, and more skeptically everywhere else. That’s a better relationship with the tool than either blind trust or blanket suspicion, and it took getting one small fact wrong, publicly, to actually build it.

The Categories of Mistake I’ve Learned to Watch For

Not all AI errors are the same shape, and learning to tell them apart has made my verification time much more efficient. The first category is genuinely outdated information stated as current, prices, staff, policies, anything that changes over time and was accurate at some point in the training data but isn’t anymore. These are the easiest to catch once you know to look, because the fix is just checking the date on anything time-sensitive.

The second category is more subtle: plausible-sounding specifics that were never true at all, a statistic that sounds reasonable, a quote attributed to the wrong person, a detail that fits the pattern of what should be true without actually being true. This is the category that got me with the founding year, and it’s the hardest to catch, because there’s no obvious tell, no outdated timestamp to notice. The only real defense is treating every specific number or name as unverified until I’ve checked it somewhere else, which sounds exhausting written out like that but becomes fairly quick once it’s routine.

The third category is confident synthesis of genuinely conflicting sources, where the model picks one version of a disputed fact and states it as settled when actual expert sources disagree. This one I’ve only caught by accident, when I happened to already know a topic was contested and noticed the answer flattened that complexity into false certainty. I don’t have a great systematic defense against this one yet, beyond generally distrusting any answer on a topic I already know to be genuinely debated.

How This Changed the Way I Write, Not Just Research

The unexpected side effect has been in my own writing, independent of AI entirely. Once I started noticing how confident, well-constructed language can carry inaccurate information without any tonal warning sign, I started paying more attention to my own certainty in what I write. Am I stating something because I’ve actually verified it, or because it sounds right and fits the sentence I was already building? That’s an uncomfortable question to ask of your own writing, and I don’t think I asked it enough before this happened. The AI mistake was, in a strange way, a mirror for a habit I already had and hadn’t examined closely.

The One Exception I’ve Made

There’s one category of use where I’ve deliberately relaxed this rule: early brainstorming, where the entire point is generating raw material to react to, not finished claims. When I’m trying to think through possible angles on a topic, possible explanations for something, possible structures for a piece of writing, I let the AI say things I know might be wrong, because I’m not going to use any of it unverified anyway, it’s just fuel for my own thinking. The rule only applies once something moves from “raw material I’m reacting to” into “a specific claim I’m about to rely on or repeat.” Keeping that line clear in my own head has made the whole system workable instead of exhausting.

Mr. Edmilson

Leave a Reply

Your email address will not be published. Required fields are marked *