I noticed something interesting. Whenever an AI agent starts doing strange things, the immediate response of any team is to add more of something. More context. More tools. More logging. Anything so that the AI "has all it needs." This feels right, I guess. It turns out that's usually exactly what's causing the problem.
A few months ago, my team went down that road with a project that was supposed to be simple, and turned into anything but. We were building an internal agent for our support engineers, something to help them triage production incidents faster. Pull logs, check what got deployed recently, look at open tickets, take a guess at the root cause. Standard stuff on paper. The kind of task that sounds like a weekend build until you hand it to people who are on call at 2 am and have zero patience for a tool that wastes their time.
So version one was, well, everything at once. We bolted on log search, ticket lookup, deploy history, a Slack search tool, runbook lookup, a metrics query tool. Twenty-two tools total, all sitting in front of one agent with a system prompt that had crept up to several pages. Worked great in the demo. It always does.
Then it went to three engineers who were actually on call that week, and their notes were blunt. Slow. Guessed at things it had no business guessing at. One of them told me flat out that it took 28 seconds just to answer "what changed in the last hour," and the answer it gave mixed up two completely different services. I won't pretend that didn't sting a little.
Instead of going back and rewriting the prompt for the fourth time, we actually sat down and traced where all that time and context was going.
Most of it wasn't the actual problem we were solving. It was bookkeeping. Every time a tool ran, its full raw output got dropped straight into the conversation, and none of it ever got cleaned up, so the agent kept carrying it forward turn after turn. A single log query alone could dump a few thousand tokens.
Four or five steps into a real debugging session and we were sitting around 70 to 80k tokens of context, and honestly most of that was just stale junk from three steps ago that nobody needed anymore.
The tool count made it worse in a way I hadn't fully appreciated until we measured it. With 22 tools registered, the model had to weigh all 22 of them on every single turn, even for a question where 18 of them were completely irrelevant. That's not free. It adds latency by itself, and it gives the model that many more chances to pick the wrong one.
Here's roughly what that looked like next to what we changed it into.

No bigger model. If anything, we went the opposite direction and shrank the thing down.
We cut the 22 tools to 6 by grouping them into skills. A log-investigation skill, a deploy-history skill, a ticket-search skill, that kind of split. The agent only loads whichever one applies to the current step instead of holding all of them in view the whole time. Fewer choices in front of the model, fewer wrong picks.
This took us a couple of tries before it clicked, so it's probably worth spelling out. Every skill is its own small manifest file, nothing more than a name, a short line describing when it applies, and the one or two tools it owns. None of that sits in the agent's context by default. When a query comes in, a small, fast classifier pass looks at it and decides which skill fits, and it does this without ever seeing the full tool catalog at all. It's a bit like asking a receptionist which room you need instead of handing you the floor plan for the entire building.
Once that classifier picks a skill, only that skill's tools get bound into context for the step. So "what changed in the last hour" pulls in the deploy-history skill and nothing else — logs, tickets, metrics never show up in the conversation at all. In token terms, that's the difference between reasoning over something like 3,500 tokens of combined tool schemas versus roughly 350. Then the skill unloads once the step is done, and the next turn starts fresh.
Honestly, this single change accounted for more of our savings than anything else we tried. Loading every tool's schema up front is a fixed tax you pay on every turn no matter how simple the question is, and it scales with how many tools exist, not with how hard the question actually is. Narrowing that down to only what a given step could plausibly need turned a fixed cost into something closer to a variable one.

Second big change: we stopped shoving raw tool output straight into the conversation. Results get written to a scratch file now, and the agent reads back only the slice it actually needs at that moment. The filesystem does the remembering instead of the context window, so a 3,000-token log dump doesn't get dragged along into every future turn just because it happened once.
We also added a compaction step. Every few turns, whatever's been checked and ruled out gets boiled down into a short note, the raw history gets dropped, and the agent keeps working from that summary going forward.
And for anything genuinely parallel, checking three services at once, say, we hand each one off to its own subagent instead of making the main agent track three separate threads in its head. Each one just reports back a short result when it's done.
On our set of past incidents, average context per turn went from around 75k tokens down to about 18k. Median response time dropped from 28 seconds to roughly 7. And accuracy went up too, which honestly caught me off guard more than the speed did. Our internal eval score moved from the low 60s into the mid 80s out of 100, and near as I can tell that's mostly because the agent isn't wading through irrelevant context anymore right when it needs to make a call.
Same model the whole time. Same weights. The only thing that changed was how much we handed it at once, and whether it had to ask for more before it got it.
I get why the instinct is to add. More tools feel generous, more context feels thorough, more instructions feel like covering your bases. In practice, all it does is give the model a bigger pile to dig through before it can do anything useful.
What actually worked for us didn't feel like an AI trick at all. It felt like ordinary engineering hygiene. Smaller surface area. Clear ownership over what each piece is responsible for. Memory that gets swept instead of piling up in a corner forever. None of it was clever, if I'm honest. It was just tidy.
We spent months teaching the agent to do more. The fix was teaching it to forget.