There's a specific moment when an AI system stops being an experiment and becomes a real engineering problem: when the agent calls your API and doesn't understand what you send back. Not because the model is poor. But because your API was never designed for it.
For years, backend teams have designed contracts — endpoints, error schemas, status codes, response messages — with a single type of client in mind: the human who opens a terminal, reads the 429, and decides what to do. That human has judgment. They interpret. They infer. An AI agent does none of those three things in the same way. It follows instructions. It retries. And if the contract is ambiguous, it improvises in ways that corrupt data or collapse queues.
The result is predictable: agents in production that work in demos and degrade silently in the real world. This isn't a prompt problem. It's an architecture problem. And the teams solving it aren't rewriting models — they're rewriting the backend.
Current APIs talk to humans; agents need machine-readable contracts with explicit semantics and no ambiguity.
The most dangerous errors aren't the ones that break the agent — they're the ones the agent misinterprets and keeps executing.
Designing for agents isn't adding a layer on top; it's rethinking what information a non-human client needs to make correct decisions.
The Problem: Contracts Designed for Eyes, Not Machines
When a developer designs an endpoint, they think about who will consume it: another developer, a web interface, a mobile app. All of them share something: there's a human in the loop who can interpret an unexpected response. If the server returns an HTTP 422 with a cryptic validation message, the developer reads it, understands it, and fixes it. If it returns a 200 with an empty body because the operation was partially successful, the developer catches it in the logs and reports it.
An AI agent does none of that. It receives the response, interprets it according to the schema it has in context, and decides its next action. If the contract is ambiguous — if the same 200 can mean full success, partial success, or "we'll process it later" — the agent picks an interpretation. And that interpretation can be catastrophic.
The challenge of designing rate limits for machine callers is a good illustration: a 429 has two readers — the human debugging in a terminal and the retry loop parsing it in code. Most APIs are only written for the first. This isn't a minor issue when the second is an autonomous agent deciding how many times to retry, with what backoff, and whether the failed operation is important enough to escalate up the decision chain.
The parallel with critical software failures in high-stakes environments is apt: expert users are left "unable to do anything" because a system controlling critical functions behaved unexpectedly. There's no visible catastrophic failure. Just silent degradation. That's exactly what happens when an agent consumes an API not designed for it: it doesn't explode, it simply stops doing what was expected — and nobody sees it until the damage is done.
Semantics: What the Status Code Can't Say Alone
The underlying problem is semantic. HTTP status codes were designed to describe the state of the network transaction, not the business intent. A 200 says "the request arrived and the server responded." It doesn't say whether the operation had the effect the caller expected. A 400 says "something in your request is wrong." It doesn't say whether that "something" is recoverable, whether the agent should retry with different data, or whether it should halt the entire chain and escalate.
For a human, that ambiguity is tolerable: read the message, infer the context, act. For an agent, ambiguity is an error vector. And the more autonomous the agent — the longer its planning horizon, the more actions it chains before waiting for human confirmation — the more that error amplifies.
An agent doesn't read between the lines. If the contract isn't explicit, the agent invents the missing line. And that invented line can trigger ten more actions before anyone notices.
The solution isn't adding more text to the message field of an error response. It's structuring response semantics to be machine-interpretable without ambiguity. This involves several concrete design decisions:
Domain-specific error codes, not just generic HTTP. Instead of a catch-all 400, an error schema that distinguishes between "invalid parameter," "operation not permitted in this state," "quota limit temporarily reached," and "quota limit permanently reached."
Explicit action instructions in the response: should the agent retry? When? With the same data or modified data? Should it escalate to a human?
Operation state separated from request state: knowing the request was received is not the same as knowing the operation had the expected effect.
This isn't science fiction. It's what the best-designed APIs on the market already do: Stripe distinguishes between card errors, network errors, fraud errors, and configuration errors — each with a code, a description, and a recommended action. That level of granularity was designed for human developers, but it turns out to work perfectly for agents too. The lesson is that designing with precision for humans is a good proxy for designing for machines. The problem is when an API was designed lazily for humans: the agent suffers that laziness in amplified form.
Idempotency: The Life Insurance Your API Probably Lacks
If there's one concept that separates APIs designed for agents from those that aren't, it's idempotency. An agent that retries isn't a bug — it's the expected behaviour under uncertainty. The question is what happens when that retry reaches your backend.
If your "create order" endpoint isn't idempotent, an agent retrying on a timeout can create the same order three times. If your "send notification" endpoint isn't idempotent, the user receives three emails. If your "process payment" endpoint isn't idempotent... the scenario is obvious.
The Idempotency Key Pattern
The standard solution is the idempotency key pattern: the caller sends a unique identifier with each operation, and the server guarantees that operation — with that identifier — executes exactly once, regardless of how many times the request arrives. Stripe has been using it for over a decade. It's not a new idea. But it is an idea that most SME backends haven't implemented because, when the client was always a web interface with a human behind it, the probability of double submission was low and the cost of implementing idempotency seemed higher than the cost of rare edge cases.
With autonomous agents, that equation changes completely. An agent operating under uncertain network conditions, with variable latencies and no visual confirmation of success, will retry. Always. The question isn't whether your backend will receive duplicate requests. It's when. This connects directly to what we've explored about why agents in production fail in ways that test suites never catch — many of those silent failures root here.
Beyond Idempotency: State as Contract
There's an extension of the problem that goes beyond simple idempotency: long-running operations. When an agent triggers a process that doesn't complete synchronously — document generation, an intensive calculation, a third-party integration — it needs to know what state that operation is in at any given moment. A "202 Accepted" isn't enough. It needs a mechanism to query the state, receive an update when it changes, or retrieve the result when it's available.
Designing that mechanism for agents — with explicit states, clear transitions, and signals the agent can interpret without ambiguity — is an architectural decision that must be made before writing the first endpoint, not when the agent is already in production and the team is investigating why half the operations are stuck in "processing" indefinitely.
It's worth reading alongside this what we wrote about how tool orchestration decisions compound when nobody defined why each tool is in the chain: the same design gaps surface, just at a different layer of the stack.
Architecture: Designing the Backend for the Client That's Coming
The practical conclusion isn't "rewrite all your APIs." That's the kind of advice that sounds good at a conference and paralyses teams in reality. The practical conclusion is more surgical: identify which operations will be consumed by agents and design those operations with an explicit contract for non-human clients. The rest can stay as is.
That identification isn't trivial. It requires understanding which part of the workflow will be automated, what level of autonomy the agent will have, and which operations are critical — irreversible, with external effects, with economic cost. Exactly the same questions you should be asking before deciding what to automate and why.
In practice, the work typically organises into three layers:
Contract layer: versioned response schemas, error codes with domain semantics, explicit action instructions. What the agent needs to interpret each response without ambiguity.
Guarantees layer: idempotency on operations with external effects, explicit state management for async operations, rate limits with actionable information on when and how to retry.
Observability layer: traceability of which agent ran which operation, when, with what result. Without this layer, debugging anomalous agent behaviour is like searching for a failure's cause in logs that don't record what you need to see.
What we see in projects where the agent arrives late to the architecture conversation is always the same sequence: the prototype works, the demo convinces, the team ships to production, and two weeks later there's a behaviour nobody understands. The AI team says the model is working fine. The backend team says the APIs are returning the right thing. Both are correct. The problem lives in the space between the two: the contract that nobody defined precisely because nobody expected the client to be non-human.
The most expensive technical debt isn't the code written badly. It's the contract assumed implicit that nobody writes down.
Changing that doesn't require months or big migrations. It requires asking the right questions before the agent goes to production: what does this client need to know to act correctly? What happens if this operation arrives twice? How does the agent know a long-running operation has finished, and with what result? What does it do if it doesn't know?
These are engineering questions, not AI questions. And their answers determine whether the agent in production is an operational advantage or a systematic source of incidents.
If you're at that inflection point — the agent already exists, the backend wasn't ready, and the team is patching instead of designing — our work in custom product development for internal teams includes exactly this kind of architecture review: no unnecessary rewrites, surgical decisions about which contracts need to evolve and which can stay as they are. If you want to start there, the first conversation is free.






