📚 Part of a Learning Journey
This post is part of the Working Effectively With AI pathway.
The first time I ran an AI-assisted triage system against our ERP support desk, roughly six weeks ago, it burned through 5.2 million tokens and produced almost nothing I could use.
Not because it lacked the tools. I'd connected it to the MCP for our ERP system and given it real access: read the ticket, use that connection to investigate, come back with what was actually causing the problem. Not every ticket came back useless. But a lot of them did, and the pattern across the bad ones was the same. The answers were generic: a lot of the ticket's own wording handed back with a diagnosis wrapped around it, and the checking behind that diagnosis was thin, no more than a plausible-looking guess dressed up as a finding. Every ticket got treated like the first ticket anyone had ever raised. No memory of the hundred similar ones before it. No sense of what usually turns out to be true. The tools were good. The instructions I'd given it were too imprecise to make good use of them. Which mattered, because the whole reason I was building this was that we were struggling to answer the volume of questions coming in.
Around the same time, I noticed the same shape of problem somewhere else entirely. I'd started using AI to help draft development tickets: descriptions of bugs and change requests to hand off to a developer. The output read fine. It was also, consistently, more verbose and more ambiguous than what I'd have written myself. Not wrong, exactly. Just soft at the edges, in a way that meant a developer picking it up still had to go and find out what was actually meant.
Two different workflows. Same underlying failure.
The problem isn't the wrong answer. It's that you can't tell.
If an AI system is confidently wrong, that's bad. But it's a solvable kind of bad: you find out, you correct it, you move on.
The harder problem is that confident-and-wrong looks identical to confident-and-right. The tone doesn't change. The fluency doesn't drop. A model that has genuinely worked something out and a model that's pattern-matched its way to something plausible will describe both with exactly the same certainty. There's no tell.
I'd already run into a milder version of this without naming it: complex issues where I was driving Claude Code myself, manually correcting it mid-investigation because I could see it heading somewhere wrong before it got there. That's fine when I'm the one steering. It doesn't scale to a system meant to help other people, less experienced in the underlying platform than I am, get to a right answer faster. If the system is wrong with the same confidence it's right, it isn't actually helping them. It's just moving the guessing downstream to someone worse-placed to catch it.
That's the real stakes, for me. The point of building this wasn't to save myself time. It was to take a genuinely painful, repetitive part of the job (triage) and make it something more people could do well, faster, freeing up time for the higher-value work underneath it. Misinformation doesn't help that. A system that sounds authoritative and is sometimes just wrong is worse than a slower system that admits what it doesn't know, because the first one erodes trust in a way that's very hard to see happening in real time.
Two sides of the same coin
Once I could name the problem, the fix turned out to have two halves, and they're really the same discipline pointed in two directions.
On the triage side: don't let the expensive path run by default. Most tickets don't need deep investigation: a cheap, cheap-first pass (does this ticket even have enough information to diagnose? does a known pattern already match?) handles a large share of the queue on its own, and only escalates to genuinely expensive investigation when the cheap pass can't resolve it. This isn't just about cost, although the cost difference is real. It's that reserving the expensive, careful work for the cases that actually need it is what makes the careful work possible at all. Run everything at maximum effort and maximum effort becomes the average, which in practice means it becomes sloppy.
On the ticket-writing side: the discipline is almost the opposite instinct, applied to the same problem. Instead of holding back effort, you hold back assumption. Rather than letting a model fill in ambiguity with plausible-sounding domain knowledge (the thing it's extremely good at, and exactly the behaviour that makes wrong answers indistinguishable from right ones), the system has to be willing to say this can't be verified from here, here's what would confirm it rather than quietly papering over the gap. A precise, honest "I don't know yet" is worth more downstream than a fluent guess, because someone can act on the first and gets misled by the second.
Both halves come down to the same instinct: add less, not more. Not more checking, not more prose, not more confident-sounding hedged coverage of every possibility: less. Fewer expensive investigations, reserved for where they're earned. Fewer assumptions dressed up as facts. The instinct in both cases runs against the grain of what feels like diligence (surely more checking is safer, surely more detail is more helpful), and in both cases it wasn't. It was noise wearing the costume of thoroughness.
The discipline doesn't stay fixed once you've named it
Six weeks after that first run, while I was writing this piece, I ran the triage system again on a larger-than-normal batch. The diagnosis quality was genuinely better this time: the epistemic half of the discipline had held. The cost hadn't. That run burned nearly as many tokens as the first one, for a completely different reason.
Two things had gone wrong, separately, and neither was a knowledge gap. I'd run the update path that checks our source diffs and refreshes the system's skills, and started the triage run straight after without clearing the context in between. Separately, I'd been editing the playbook and reintroduced a pattern I'd deliberately removed a while earlier specifically to control cost: agents fanning out to work sections of a ticket in parallel. The edit was small, but it meant every one of the 168 agents dispatched that run loaded the full playbook, all six hundred-odd lines of it, into its own context before doing anything else. A hundred and sixty-eight agents at roughly 35,000 tokens each comes to close to 5.9 million tokens: in the same range as what the very first, worst run had cost, for a batch that didn't need anywhere near it.
I knew both of those practices. I'd built the cost control that got quietly undone. What let it back in wasn't ignorance, it was an edit made while multitasking, on a larger-than-usual run, having just come back from two weeks away. The point isn't the excuse. It's that the discipline this whole piece is about is not a fact you learn once and then hold. It's a state you have to keep checking for, the same way the fanning-agent pattern quietly came back after I'd already removed it once: getting the epistemic half right that run didn't protect the cost half at all. They're separate disciplines that happen to point the same direction, and either one can slip while the other is working fine.
I came to this from the practical end, ticket by ticket, cost report by cost report. It's also, it turns out, a documented pattern well outside software: Leidy Klotz's Subtract is built entirely around the observation that people systematically reach for addition as the default fix, even when subtraction would solve the problem better and cheaper. It's a strange thing to find your own AI tooling problem sitting inside a book about architecture and bureaucracy and household clutter, but the shape is exactly the same: the fix usually isn't doing more. It's being disciplined about doing less, in the right places, and staying disciplined about it after the first time you get it right.
Both halves, every run
The system isn't finished and I'm not pretending it is. It's still running as an experiment, not on a schedule, and I'm watching each run for both halves of this: is the diagnosis honest about what it doesn't know, and did the cost stay where it should. The plan is to move it onto a scheduled cadence once both of those hold consistently, not before.
The models keep getting faster and more capable, and that trend isn't slowing down. Which makes this more urgent, not less: a more capable model is a more convincing one, in both directions. The gap between confident-and-right and confident-and-wrong doesn't close as the tools improve. If anything, it gets harder to see.
Source:
- Subtract: The Untold Science of Less by Leidy Klotz (Flatiron Books, 2021)
📚 This post is part of a learning pathway
Working Effectively With AILearning where AI works, where it doesn't, and how to build production-ready tools at team scale
View the full pathway →