Skip to content
All posts
8 min read

What should happen when the AI gets it wrong

Every demo is the happy path. What matters is what a tool does at the edge of what it knows — because a confident wrong answer to a customer is worse than no answer at all. Here's the question to ask, and what a good answer sounds like.

aicustomer serviceescalationlocal businesstrust
Cover illustration for “What should happen when the AI gets it wrong”

It's 8:50 on a Saturday morning and there's a woman at the counter with a confirmation on her phone. She's here for a service the shop stopped offering back in the spring. Nobody who works there sent her that message. Something answered her question at eleven the night before — quickly, fluently, in complete sentences — and it was wrong. Now there's a person standing in the shop who is not wrong at all. She rearranged her morning around this. She's holding proof.

the failure people brace for is the wrong one

When owners worry about letting software talk to their customers, they picture it not understanding. Someone asks something strange, the thing gets confused, there's an awkward moment, everyone moves on. That's the failure they rehearse in their heads, and it's mild. A confused tool is a nuisance.

The failure that actually costs you is the opposite shape. The tool understands the question perfectly well, has no grounds whatsoever for an answer, and produces one anyway — in the same calm, competent register it uses for the things it genuinely knows. There's no tremor. It doesn't hedge. Language models have no "not sure" face, and nothing about the way the text arrives tells you whether it came from your actual policy or from the space where your policy should have been.

That's the part worth sitting with. A wrong answer delivered uncertainly is a small problem, because the customer discounts it. A wrong answer delivered with total composure is a much bigger one, because the customer believes it, acts on it, and shows up with a receipt.

two that made it into the public record

Most of these never leave the shop. Two didn't.

In February 2024, the British Columbia Civil Resolution Tribunal ruled against Air Canada in a claim brought by a passenger named Jake Moffatt. He'd asked the airline's chatbot about bereavement fares while booking travel after his grandmother died. The chatbot told him he could book at the regular rate and apply for the reduced fare afterward, within a set window. That was not the airline's policy — the airline's own web page contradicted it. Air Canada argued, among other things, that the chatbot was a separate entity responsible for its own statements. The tribunal was unmoved, and found the airline liable for negligent misrepresentation: a customer shouldn't have to cross-check one part of a company's website against another to work out which part is telling the truth.

In October 2023, New York City launched a chatbot to answer small-business questions about city rules. In March 2024, The Markup tested it and published what came back: the tool told users an employer could take a cut of workers' tips, that a landlord could turn away a tenant using a housing voucher, and that no rule required a business to accept cash. Each was wrong, and wrong in a direction that would have put an owner in real trouble for following it. In follow-up testing days later, asked point-blank whether it could be relied on for professional advice, it answered yes.

Neither tool was broken in an exotic way. Both were doing what this kind of software does at the edge of what it knows: producing the most plausible-sounding continuation and handing it over with a straight face. The lesson isn't that these were bad products. It's that fluency and accuracy are separate properties, and only one of them is visible.

the question to ask in a demo

Every demo is the happy path. That isn't a scandal — it's what a demo is for. The scripted question gets the scripted answer, the room nods, and you learn nothing about the only thing that matters.

So ask this instead, out loud, in the room:

Show me what it does when it doesn't know.

Not tell me. Show me. Hand them a question the tool has no business answering — something specific to an operation like yours: a policy that changed, a service you discontinued, an exception you make for one kind of customer and nobody else. Then watch what comes out. You'll learn more in that minute than in the rest of the hour.

It's the question the demo is built to avoid, which is precisely why it's worth asking. The reaction tells you nearly as much as the output. A vendor who has thought hard about this has an answer ready and is a little pleased you asked. A vendor who hasn't will steer you back to the script.

what a good answer looks like

Four things, roughly. None of them impressive, which is the point.

Scope with an edge on it. A tool that answers three kinds of question well and declines the rest is worth more than one that'll have a go at anything. Ask where the boundary is and who drew it. If there isn't really a boundary — if the pitch is that it copes with whatever comes up — that's not a strength, it's an unowned risk sitting on your side of the table. Breadth is cheap to demo and expensive to be wrong about.

An honest stall instead of a guess. The behavior you want at the edge is boring: the tool acknowledges the question, doesn't answer it, states plainly that a person will confirm, and then actually causes that to happen. Something like I'm not certain about that one — let me get you a real answer today is a perfectly good customer experience. It isn't a failure. People are entirely used to a human at a counter needing to go check on something. Nobody has ever been offended by it. The urge to make a tool sound omniscient is a vendor instinct, not a customer one.

A real route to a person, with the context attached. Every tool on earth claims to escalate. What's worth checking is what escalation means in practice. Does it reach someone with authority to decide, or does it drop into a queue? Does that person arrive already holding the conversation, or do they open by asking the customer to repeat everything they just typed? A handoff that dumps a frustrated stranger into a fresh line with no history isn't a handoff. It's a second failure wearing the first one's clothes.

A log you'd actually read. Ask where the answers live afterward. Not a metrics dashboard — the literal text of what the thing told your customers, in a list you can skim over coffee. This is what gets waved away most often and matters most. Both public failures above surfaced because a human read transcripts. Neither would have shown up on a chart of resolution rates. If you can't see what the tool told people, you have no way to know it was right.

this isn't really an AI question

The useful thing about show me what it does when it doesn't know is that it has almost nothing to do with AI. You'd ask it of an answering service before signing. You'd want it from a new hire in their first week on the front desk. Someone who guesses confidently is a bigger problem than someone who asks, and every owner who has trained anyone knows this already.

The model transfers cleanly: bounded scope plus honest escalation beats broad confident coverage. True of people, vendors, and software alike. What's different about software is only that it's far more fluent than a nervous new hire, and therefore much better at concealing the moment it stopped knowing things. That model also makes a bad slide, which is roughly why the other one keeps getting built.

For what it's worth, I have a stake here: I work on Sarah, the front-desk assistant on the Catalyst side of things. So read this as an interested party's view, and go make somebody's tool fail in front of you rather than taking my word for it. Being straight about ours — Sarah works over text today, not voice. Voice is on the roadmap and isn't live. Live texting switches on for a business as carrier approval for business messaging clears, and until then a person on our team runs the front desk by hand. Slower than the pitch. Honest about where the edge is.

the point

Back to the counter. The right move is almost certainly to honor the confirmation, because she did nothing wrong and the shop's software told her something the shop now has to own — which, per that tribunal in British Columbia, is not merely a moral position but increasingly a legal one.

What's worth noticing is that the morning was avoidable, and not by finding a smarter tool. It was avoidable by using one that had been told what it wasn't allowed to answer, and what to do when it hit that line.

A tool that stops and admits it doesn't know is not failing. It's working. The one to be wary of is the one that never stops — that has something for every question, delivered at identical confidence whether it's reading your real policy or filling a hole with something shaped like your policy. You can't hear the difference. Your customer certainly can't. The only defense is a thing built to know where its knowledge ends, and to say so early, to a person who can sort it out — before someone is standing at your counter on a Saturday morning, holding a confirmation you never sent.

Building something? Let's talk.

Tell us what you're working on and a real person will reply.