The Joining Tax: A Year of Governing Claude Code
In January 2026 Anthropic published a retrospective on multi-agent systems that described what I'd b 2026-10-8 05:58:1 Author: hackernoon.com(查看原文) 阅读量:2 收藏

In January 2026 Anthropic published a retrospective on multi-agent systems that described what I'd been running for a year, and described it as the mistake. Their wording, near enough: teams build elaborate multi-agent systems with separate agents for planning, execution, review and iteration, only to discover that they suffered from lost context at each handoff and spent more tokens coordinating than executing. This is the same company whose earlier post reported a multi-agent setup beating a single agent by 90%, so the retrospective reads as a walk-back. And an April 2026 preprint goes further still: hold the reasoning-token budget constant and a single agent matches or beats the multi-agent setups, and it argues the reported multi-agent wins were mostly extra computation nobody was counting.

I have a lead, a business analyst, an architect, a developer, QA, two reviewers who argue opposite positions, and a doc updater. They've been running against my own production product since May 2025, so I've paid the tax that retrospective describes, every month. And I kept the thing anyway. The criticism is right about the cost, but I think it's wrong about what the structure is for.

What I built

A task goes to the lead. The lead talks to the business analyst, and the two of them work up questions and discussion items from the business side until the plan is sound. Then it goes to the architect, where the same loop runs again on the technical side. Each step has a review companion whose only job is holding a second opinion, and the standing instruction for anything in a review seat is: always doubt, never believe, always fact-check.

Then a TDD phase. QA or the lead writes the definition of done as test behaviour, from the use cases agreed with the BA, and the developer works against those unit, integration and end-to-end tests until they're green. If the developer needs to touch a test to make it green that's a red flag, so the lead steps in, sorts out what kind of problem it actually is, and consults the BA or the architect depending on the answer.

Once the developer signs off on green it goes to review, with a bit of back and forth until that's green too. I ran two reviewers there for a while, one arguing refactor everything and one arguing don't touch a single extra line, with the architect judging between them. Then it comes to me, and after my review, testing. If I have feedback the lead logs it and it goes back to the developer, or the lead handles it if it's small, or I do it if neither of them can cope. The lead's main job by then is owning the issue list, where an issue means a code review finding or a bug or anything that isn't clearly reaching the definition of done.

At the end a doc updater reads the paper trail the lead kept and hardens the docs and the scripts, so the next session starts from a better place. I basically had spec-driven development running before it was mainstream, and it's a heavy process, not worth pretending otherwise. The runner that does it is on GitHub as agentweft: a flow is a spec, a role is a markdown file, and every step's output is written down so a run that dies can be picked up where it fell over.

Where the criticism is right

It's expensive, first of all. Two hours of this drains a full Max 5x quota, on a heavy day just in 90 minutes, and coordination is a real line item inside that. It's also slow in the places you don't want slow. A feature used to cost me 3 or 4 days of planning, then 4 or 5 days of review that was mostly rescuing slop, then testing. Tuning the rules and the examples over the year is what halved the review: it's 1 or 2 days of deep dive and verification now, because there's less to rescue. Planning is still about 2 days, faster mostly because the examples, the common questions and the common approaches are written down and I don't spell them out again. And testing plus bug-fixing is another 1 or 2 days that was always there and stays there, because I'm the final sign-off and that part doesn't get delegated.

There were stretches where the paper trail itself was the problem. I had decision-track files running 2k, 3k, 5k lines, a few that hit 10k, and I'd have to split them and iterate over them. Some bad features left behind a dozen files of 2k or 3k lines each, all notes about what to fix. I knew that was wrong while I was doing it; the tool was misbehaving and I was responding by writing more. That stretch eventually taught me to treat the tool as a tool and not a human being.

And the honest one: this workflow is pretty opinionated. I'm heavy on planning and heavy on review, and it works the way I already worked without AI, so it's my companion first of all, and whatever transfers to your setup is probably not the exact configuration.

Why I kept it

I kept it because the roles do two mechanical jobs, and a single long-running agent can't do either of them for itself.

The first is a bounded context per unit of work. Instruction compliance decays inside a session as the agent generates code, on the order of a few percent of odds per function in the one controlled study I've found that measures it. So a small unit of work with a fresh context has less room to drift. That's roughly what a role boundary buys. The architect thinks like anything else would, but it starts clean and it's done before it drifts. There's also a peer-reviewed paper from ACL in 2025 that made me a bit more confident about this than my own experience did. The interesting bit is a test the authors ran on their own system, switching off one part of it at a time: the gain came from hiding the earlier steps from the agent, not from splitting the work up.

The second is write scope. Reads stay wide open and I push for that, I want agents grepping and investigating anything. But writes are partitioned by role, which in agentweft is a tool scope declared on the step, and QA is the reason: the architecture docs and the definition of done are closed to its writes, since that's the same failure as a developer editing a test until it goes green. They were trying to find shortest paths and they were trying to cheat the request.

A single agent can't give you that, because there's no boundary for it to enforce against itself. Separation of duties is a hundred years older than any of this, and it exists for the same reason. So when I read that multi-agent systems lose context at every handoff and spend more on coordination than execution, I don't disagree; the handoff isn't a bug in the design, it's the price of the boundary, and the boundary is the whole reason those roles exist.

The tax, and how you pay it

Splitting work across specialists is not a new problem. When frontend and backend live separately on a team you get cleaner ownership on each side and you pay for it at the join. That joining tax is real enough that I spent years trying to avoid it in teams. Split your agents by role and it's the same tax for the same reason: the original context-narrowing problem gets solved, but the cost moves elsewhere and the final goal isn't much closer. So the work is in making the seam cheap, and five things did most of that for me.

The first was capping the debate. Each reviewer gets 2 rounds of feedback, the original review plus clarification, and the architect can extend to 3 cycles if it genuinely wants more argument; the developer gets 2 or 3 iterations before anything escalates to me, with a hard cap of 5. Without the caps the reviewers will argue for as long as the tokens last, which is why agentweft carries them in the flow spec: the reviewer sends work back twice at most, and then it has to look at what came back.

The second was giving every role a priority order, not only what it does, but what it should prioritise, where it's allowed to decide on its own, and in what order. Most of the arbitration I used to do by hand was me holding a priority list in my head that nobody else had a copy of, and writing it down per role let a lot of the pipeline run without me.

The third was making the handoff structured, so a result and a verdict cross the boundary instead of a wall of pasted text the next role has to re-read and reinterpret. And the fourth was keeping one owner for the issue list: 10 review iterations and easily 100 findings across a feature is fairly normal for me. Without a single role owning that list you get the 10k-line decision track, because that's where an unowned issue list ends up.

The fifth was closing the loop at the end. The doc updater exists so the lessons from this run turn into docs and scripts before the next one starts, and that's the only reason the rules corpus shrank over the year instead of growing; it sits at about a fifth of its old size now. The issues got repetitive over time, and a repetitive issue is one a script can take over; in agentweft those are gates, a regex or a command that either passed or didn't, and the flow hands the same promises to the role doing the work and the gate checking it.

When one agent wins

If I already know which files change and how, the whole apparatus is negative value. A single agent with a single task does it better without all the back and forth. The crossover sits roughly where a task stops fitting in one head: if a single agent would have to compact, summarise its own context and carry on, 20 or 50 times to finish something, it arrives at the end having quietly forgotten the beginning. At that point paying the joining tax gets cheaper than paying for what the agent forgets. And the bigger the scope, the better the trade gets, which is kind of the opposite of what you'd guess from the coordination overhead.

If I rebuilt it today I'd keep the write partitioning, the tests as definition of done, the caps and the doc-updater loop. But I'd think harder about how many roles need to exist, because some of mine are there for symmetry with how human teams are organised, and symmetry with a human team is not a technical argument. That part of the criticism lands on me.

I still read every line that comes out of it, and I expect to keep doing that. But the reason I'll defend the shape, after a year and against people with better data than mine, is that the alternative isn't a cheaper system. It's the same cost, and it sits somewhere I can't see it.


文章来源: https://hackernoon.com/the-joining-tax-a-year-of-governing-claude-code?source=rss
如有侵权请联系:admin#unsafe.sh