One of the principles that has stuck with me ever since I read The Systems Bible by John Gall is:
The problem isn't the problem, coping is the problem.
This principle encourages designing systems that continue functioning even when a component of the system fails. By way of example, let's consider HTTP requests. When an HTTP request fails, our first instinct is to add a retry. But adding retries doesn't actually fix the problem; the second request can fail. And the third. By adding retries, we've hopefully made the problem occur less frequently. But there is absolutely nothing we can do to ensure an HTTP request will always succeed. So, how do we cope?
To decide how to cope, we have to take a broader look at the system. Let's extend the example. The HTTP request is for push notifications between two servers: server A is the sender, B the recipient. There are all sorts of reasons A might not be able to deliver to B. In order to cope in this situation, A could implement a new endpoint allowing B to pull notifications.
Now, A gets to decide how much effort to spend mitigating delivery failures, or whether further mitigation is worthwhile at all. Addressing the delivery problem becomes an optimisation rather than a necessity.
In general, when designing systems, you should plan how the system will cope with a problem before planning how the system will mitigate the problem. For the purpose of this article, I will define mitigation as reducing the frequency, but not eliminating the problem. Other examples of mitigation could be increasing timeouts because processing is taking too long. Or increasing cache sizes because responses are too slow.
This isn't to say that mitigation is pointless. Adding mitigation for an existing behavior will almost always be quicker than adding a fallback to cope. In the middle of an incident, adding retries might be the difference between 100 angry emails and 10. Coping provides the fallback, while mitigation tries to reduce how often we need it. By designing the fallback first, we can decide how important the mitigation is.
This principle was very important when I was recently building a personal agent to replace a spreadsheet. Natural language can be a fickle beast, with a lack of context or precision leading to misunderstanding. Without adding the ability to cope, I'd still be playing whack-a-mole, tweaking prompts trying to get 100% of the evals to pass.
The purpose of the spreadsheet the agent replaced was to keep an account of expenses between two households. We will often pick things up for each other from the shops. Then send a message in Slack saying how much the expense was. To which the response almost always was, "Have you added that to the spreadsheet?" Periodically, we would check the spreadsheet and settle the difference.
I kept the design for the agent simple; it receives single messages in Slack and decides which tool (if any) to call. No back and forth. I wanted to keep the non-deterministic part to a minimum (deterministic code is predictable, faster, and cheaper). In this system, the main tool I expect to be called is add_expense. And the main problem I ran into was classification between shared and individual expenses, where a shared expense would be split between the households.
I was tempted to solve this problem by crafting a system prompt and tool descriptions so beautiful the LLM gods would weep. Alas, I am only mortal, so I added the find_and_edit_expense tool. To be clear, there is no issue with adding a new tool, but an increase in complexity has to be earned. When an expense is mentioned in Slack, it gets recorded, and the Slack channel is notified. If the user notices that the details of the notification were wrong, they can send a follow-up message correcting the mistake. I will admit, I did burn a few hours attempting to create that beautiful prompt. I knew I'd still need the edit tool because no prompt was going to overcome the inherent non-determinism, but it was still an interesting exercise.
Couldn't resist mate
The few hours I did burn on tweaking the prompt were spent trying to add context while keeping the prompt short and general. Keeping it short is important for keeping cost down. And, making sure the prompt is kept general attempts to not overfit to the evals. I even removed a few evals which I decided were reasonable for the LLM to be getting wrong. After all, I was going to add editing to cope.
Once the edit tool was in place, it took a lot of pressure off the add tool. The add tool no longer had to be perfect; it was now good enough. Tuning the prompt became an optimisation, rather than a requirement. I can also add further context to improve error rate if we notice a set of expenses being classified incorrectly.
Coping by editing was fine for our low-stakes expenses ledger. The main cost of a bad classification is the friction of having to correct it. In a higher-stakes environment where incorrect tool calls are more expensive, we would want a different way to cope. For example, we could send the notification, but delay the execution, giving the user time to cancel. Or, at still higher stakes, require explicit approval through deterministic code before executing risky tool calls.
Before closing, I also want to call out that sometimes logging the failure, or even doing nothing and moving on, is the right fallback. As long as that's a deliberate decision. Everything is a trade-off, and in some cases the added complexity of doing more is not worth it.
Design the fallback first. Take the pressure off mitigation and let it become an optimization problem. Some problems can't be prevented. All we can do is decide how the system will cope.