Product, test, data, or environment — and the evidence that tells them apart.
TL;DR — A failing test tells you something broke, not what. I wrote a rulebook that sorts failures by one question: which artifact is wrong — the product, the test, the data it used, or the environment? Two LLMs applied it to 21 failures, with the rules and without. Without them, one model took a new test's expected value as proof the product was broken. With them, that stopped, and every "I can't tell" came with the missing evidence named. The rulebook and every model reply are in the repo: test-failure-taxonomy
The pipeline is red. Someone says, "it's flaky". Someone else says "no, the API changed". A third person reruns it. Half an hour later, nobody has looked at the same evidence twice.
The question underneath is always the same: what is wrong here? The product, the test, the data the test used, or the environment it ran in? Each answer sends the work to a different place, and a wrong answer is expensive in both directions. Developers chase a product bug that turns out to be a stale test. Or a real product bug gets waved away as "just the test" and ships.
Teams do sort their failures, but usually by the wrong thing.
So, I tried sorting by the thing you actually need to know.
Which artifact is the wrong one?
A failure surfaces wherever the assertion happens to sit, which is rarely where the fault is. For this rulebook, I group root causes into four classes, each one an answer to that question:
|
Class |
The wrong artifact is |
|---|---|
|
PRODUCT_DEFECT |
The product |
|
TEST_CODE_DEFECT |
The test itself: its logic, the values it supplies, or state it carries between runs |
|
TEST_DATA |
A record the test consumed but did not create during this run |
|
ENVIRONMENT_CONFIG |
The environment: infrastructure, configuration, a third-party dependency |
|
INSUFFICIENT_DATA |
You can't tell yet |
That last one is not a cop-out, because it comes with a bar to clear. To answer "not enough evidence," you must name the two classes the case sits between, and the one piece of evidence that would decide it. If you can't name that piece, you're not short of evidence — you're short of confidence, and that's a different problem.
The rules work as an elimination order. Ask in this order and stop at the first answer:
Here's a real-looking failure. An invoice total is one cent out:
java.lang.AssertionError: Invoice total mismatch expected [100.00] but found [99.99]
The response is a 200. The line items are 50.005 and 49.995. They sum to 100.00 exactly, and the service says 99.99. Open and shut, surely: the product is rounding wrong.
Not yet. The only thing claiming 100.00 is correct is the test's own assertion. And in this case the test is brand new. It has never passed against any build, so there's no history saying its expectation was ever right. Rounding each line to cents before summing isn't a bug either — it's one of two legitimate accounting conventions, and which one applies depends on a contract nobody has produced.
That's question 4, and it's the one everyone skips. A product is not wrong until something other than the test says what right is: a contract, a specification, a release note, or a build where the same unedited test was green. Without one of those, you have a disagreement, not a defect.
Now take the same failure with the contract attached:
totalMUST equal the sum oflineItems[].amountcomputed at full precision and rounded HALF_UP to the currency's minor unit exactly once. Implementations MUST NOT round individual line items before summation.
…and server logs showing Money.of(50.005) -> 50.00 and Money.of(49.995) -> 49.99. Now it's a product defect, and you can prove it.
Two failures that look identical, one piece of evidence apart, with different answers. I built the whole test set on that idea and called them look-alike pairs — because a triage process that gives both halves the same answer is reading the surface, not the evidence.
Rules that only their author can apply aren't rules. So, I built 21 failure cases with full evidence, each carrying whichever of these a real triage would have — error, stack trace, logs, request and response, test code, its history, the contract — and a label assigned before any model saw it. The set contains eight look-alike pairs: two failures with similar symptoms and different causes, one piece of evidence apart.
For six of the cases, I also made a stripped-down version containing only the error message, to see whether a model would guess once the evidence needed to classify was gone. That gives 27 evidence files: 21 full, 6 error-message-only.
Then I asked two models from two vendors, claude-opus-5 and gpt-6-astra, to classify each file three times, once with the rulebook in the prompt and once with only a one-line definition of each class. 27 files × 2 models × 2 conditions × 3 runs = 324 replies, 162 per model, all saved in the repo.
Every call went through the vendor's API rather than a chat interface, and each one was independent: a fresh request carrying only the prompt and one evidence file, with no conversation history and nothing carried over from the case before. That matters twice over. A chat session accumulates context, so a model that had just seen the invoice case with its contract would carry that into the version without it — and that comparison is what the whole experiment rests on.
Independence is also what makes three runs per file meaningful: three separate attempts at the same question, rather than one answer restated. Temperature was left unset, since neither model accepts one; everything ran at defaults.
The limits, before the numbers. The 21 cases are constructed rather than captured from real builds, and drafted with AI help. Constructed evidence is clearer than real evidence, so these results are likely a best case. Two models, three runs, one run date. And the expected labels came from the same rules the models were given, so agreement shows the rules are clear enough to apply consistently — not that they're correct. Read the numbers as examples of where the rules act, not as rates.
|
With the rules |
Without | |
|---|---|---|
|
Look-alike pairs told apart |
15 of 16 |
12 of 16 |
|
Look-alike pairs given the same answer |
0 |
1 |
|
Consistent classification across all 3 runs |
54 of 54 |
52 of 54 |
16 = 8 pairs × 2 models; 54 = 27 files × 2 models. Each column counts one condition.
The headline isn't the size of the gap. It's where the gap is.
Once, and in exactly the place the rule was written for. Without the rules, claude-opus-5 answered PRODUCT_DEFECT to the invoice case without the contract, in 2 of 3 runs:
The expected value of 100.00 is therefore arithmetically right, so the assembler's rounding logic is the fault, not the assertion. The test being brand new with no prior green build explains why this latent rounding bug is only now surfacing, not that the test is wrong.
Read that last sentence again. The model noticed the test had never passed — the strongest reason to doubt its expectation — and turned it into evidence for the expectation.
With the rules, the same model, same evidence, three times out of three:
the only statement that 100.00 is correct is the test's own assertion. There is no contract, specification or release note on how line-item amounts (50.005 + 49.995) should be rounded … the deciding evidence would be the billing rounding contract or spec for invoice totals.
This was also the only pair either model ever merged: the same answer to both halves, from a model reading the invoice's surface rather than what was attached to it.
Worth saying plainly: this is one model on one case, not "AI can't triage". The other model declined to decide that same case even without rules. One over-committed, the other held back. Different failure modes, same fix.
Across both models and all three runs, the full-evidence files produced 21 INSUFFICIENT_DATA replies. Every one of them named the exact evidence that would settle the case:
That turns the most useless answer in triage into the most useful one. "I don't know" closes a ticket badly. "I don't know, and here is the one artifact that would tell us" is a to-do item — and, repeated across a quarter, a list of exactly which evidence your pipeline should start capturing. If half your INSUFFICIENT_DATA answers ask for server-side logs you don't retain, that's not an AI problem. That's a reporting gap you can fix this sprint.
The stripped-down cases produced a different result. Across all six variants, neither model overclaimed when the evidence was absent: every reply came back INSUFFICIENT_DATA — 36 with the rules, 36 without, 72 in all. That says less than it looks like, because a bare error message is the easiest thing in the world to decline; there's nothing there to build a story from. The more interesting failure appeared when there was enough evidence to build a convincing but wrong story, which is exactly the invoice case.
One case still doesn't work. A test looks up a customer ID that doesn't exist, and the rule asks whether the record was supposed to be there. The evidence shows only that it never was. One model called it a test defect; the other answered "not enough evidence" every time and named the missing piece: the fixture definition recording which ID the test should have used.
The second reading is more consistent with my own rule. The label is disputed, and I left it disputed in the repo rather than quietly changing it after seeing the scores. That rule needs a clearer worked example and a named evidence type — which is the method finding its own weak spot, and the most useful thing it did.
The contribution here isn't "LLMs can classify test failures". It's the rulebook: an evidence-based way to decide what broke, which two LLMs from different vendors could apply consistently.
The rulebook, the 21 labelled cases, the six stripped-down evidence variants, the exact prompts, every model reply, and the scoring code are all public: CC BY 4.0 for the writing and the cases, MIT for the code. Each case ships with its label and a note on how the failure was constructed, so you can check my answers rather than take them: test-failure-taxonomy
Three ways to use it. Put the five questions on the wall and triage by hand. Paste prompts/classify.md into your AI triage assistant, filling its three placeholders withTAXONOMY.md, the class names, and the failure you're triaging — it's the exact prompt that produced the 21-of-21 result. Or run the cases against a model you're considering and see whether it reads the evidence or the surface.
And the next time someone says, "it's flaky", ask which artifact is wrong.