Every time an AI agent starts a new execution, it begins from zero. The context window is empty. What happened yesterday, what it learned last week, the preferences the user set a month ago — none of that exists, unless someone deliberately put it there. This seemingly trivial fact is the root of one of the most common — and most costly — architectural failures we see in AI systems that are already running in production.
The problem isn't fundamentally technical. It's conceptual. The teams building these systems confuse two fundamentally different things: context and memory. They treat them as synonyms, mix them into the same layer, and then wonder why the system behaves erratically, why it doesn't "remember" what it should, or why it burns through tokens at a rate that makes the CFO uncomfortable.
This confusion has consequences that reach well beyond technical performance. It directly affects the end user's experience: the agent that doesn't know who you are, that repeats questions you've already answered, that loses the thread between sessions, that "learns" something in one conversation and forgets it in the next. These are failures the user feels even if they've never heard of context windows or embeddings.
Context and memory are distinct architectural layers, with radically different purposes, lifecycles, and costs.
Stuffing everything into the context is the equivalent of reading a client's full file every time you pick up the phone — possible, but absurd at scale.
Without a deliberate memory strategy, your agents are no smarter than a 2019 chatbot with better prose.
Architecture: What Context Is and Why It's a Scarce Resource
Context is what the model sees right now. It's the window — limited, ephemeral, expensive — within which the agent processes information and generates responses. Every token you put into that window has a cost: in latency, in money, in the model's reasoning capacity. Context doesn't persist. When the execution ends, it's gone.
This has an implication many teams ignore until it's too late: context is not a storage system, it's a workspace. Using it like a database — stuffing in the full history, all preferences, all relevant documents — is the equivalent of photocopying a client's complete file and reading it end-to-end before every support call. It works with ten clients. With ten thousand, the system collapses under its own weight.
The context window has a hard limit measured in tokens, which varies by model. GPT-4o handles up to 128,000 tokens; some specialised models much less. But even within that limit, there's a documented phenomenon practitioners call "lost in the middle": the model's attention degrades for information placed in the centre of the context. More input does not mean the model uses it all well.
The classic mistake is building an automation agent — one that processes orders or handles incidents — and feeding its context with weeks of interaction history "just in case." The result: saturated windows, degraded reasoning, and inference costs that spike with no proportional improvement in user experience. In projects where we've audited this pattern, reducing redundant context has cut the cost per execution by 40–60% with no perceptible drop in response quality.
Memory: The Layer Your Agents Probably Don't Have
Memory is everything the agent should know — or be able to retrieve — beyond the current execution. It's persistent, structured, and must be designed, not improvised. And here's the trap: most teams don't design it. They simulate it by stuffing things into context, or they simply don't have it and expect the user to repeat everything each session.
An AI system without a memory architecture isn't an intelligent system. It's a system that appears intelligent for exactly as long as one conversation lasts.
Memory in agent systems has at least three layers with radically different characteristics:
Short-term memory: the thread of the session
This is what the agent needs to remember within an extended session: the steps it has already taken, the tools it has already invoked, the intermediate decisions it has made. This layer is typically managed with a structured buffer injected into the context selectively — not completely. The mistake is injecting the raw history. The good practice is injecting a compressed summary or serialised state that the model can interpret efficiently.
Episodic memory: what happened in previous sessions
This is the layer most frequently missing. The user already mentioned what industry they work in. They already set their format preferences. They already rejected a similar proposal two weeks ago. If the agent doesn't have access to that information — or has it but doesn't know when to retrieve it — the experience fragments. The user feels like they're talking to someone who never takes notes.
Implementing episodic memory requires non-trivial design decisions: what gets stored? For how long? At what granularity? Who can delete or correct it? These are product questions before they are engineering questions, which is exactly why they tend to go unanswered until the user complains.
Semantic memory: domain knowledge
This is the structured knowledge base the agent operates on: product catalogues, company policies, technical documentation. This layer doesn't change every session, but it needs efficient update and retrieval mechanisms. This is where RAG (Retrieval-Augmented Generation) enters the picture — not as magic, but as an engineering solution with its own maintenance and quality requirements. The agent doesn't "know" the catalogue: it retrieves it when needed, and the quality of that retrieval depends on how well you've indexed and curated your data.
If you're building or auditing an agent system, it's worth revisiting how agent observability intersects with this problem — context and memory are precisely the kind of things that traditional logs don't capture well.
The Cost of Confusion: When User Experience Pays the Bill
The misunderstanding between context and memory isn't academic. It materialises in behavioural patterns that the end user experiences directly, even if they lack technical words to describe them.
The first is session amnesia: the agent that remembers nothing between conversations. The user has to re-explain their situation every time. It's the equivalent of calling customer support and the agent has never heard of you, even though you've been a client for five years. The solution is not to dump the full history into context — that's too expensive and degrades reasoning. It's to design an episodic memory layer that retrieves what's relevant, surgically.
The second is context noise: the agent that receives too much information and reasons worse. This one is less obvious to the user, but manifests as less precise, more generic responses, or errors that make no sense given what the system "should know." The engineering team blames model limitations. The real cause is usually poorly managed context that saturates the model's attention.
The third — and perhaps the most dangerous — is cross-execution inconsistency. Two users ask the same question in similar contexts and receive different answers, not because the model is probabilistic, but because the state injected into each context was different. In regulated environments — finance, healthcare, legal — this inconsistency isn't just a UX problem: it can be a compliance problem. This connects directly to what we've analysed before about why test suites pass while users suffer: context and memory failures are exactly the kind of issues that don't show up in unit tests but hit the user on first contact.
Deliberate Design: How to Separate the Layers in Practice
The good news is that separating context and memory doesn't require rewriting everything from scratch. It requires, above all, design clarity before writing a single line of code. These are the decisions that must be made explicitly:
Define the lifecycle of each piece of data. Is this datum relevant only for this execution, for this session, or indefinitely? The answer determines which layer it belongs to. An API error message that just occurred: context. The user's language preferences: persistent memory. The last three exchanges: a session buffer managed outside the raw context.
Design the retrieval policy, not just the storage. Many teams focus on how to store memory and not on how and when to retrieve it. An agent that retrieves everything it has stored about a user on every execution makes the same mistake as an agent with no memory at all: it saturates the context with information irrelevant to the specific task. Retrieval must be semantically relevant and temporally bounded.
Implement active history compression. For long sessions, instead of accumulating the full message history, maintain a structured summary that updates as the conversation progresses. Frameworks like LangGraph or LlamaIndex have primitives for this, but the compression strategy — what to preserve, what to discard — is a product decision that can't be left to the framework's default.
An agent's memory architecture is as important as its prompt. Probably more so, because everyone reviews the prompt and nobody reviews the memory.
Audit context cost regularly. Token usage per execution is a product metric, not just an infrastructure one. If it grows steadily without a corresponding improvement in response quality, context is being used as a warehouse. This metric belongs in the same dashboard as user satisfaction metrics. If you want a framework for instrumenting this, the work we've done on consistency in AI systems is a useful starting point.
Decide who can correct memory. This is the question nobody asks until a user calls furious because the agent has "learned" something incorrect about them and keeps applying it weeks later. Persistent memory needs accessible correction and deletion mechanisms — both for the user and for the operations team. This isn't an implementation detail: it's a trust requirement.
Ultimately, designing context and memory layers well means designing the user experience of your agents. Every technical decision — what gets injected, what gets retrieved, what gets compressed, what gets discarded — has a direct counterpart in how the user perceives the system's intelligence, coherence, and usefulness. An agent that "remembers" what matters isn't just more efficient: it's more worthy of trust.
If you're building agent systems or evaluating why the ones you have aren't performing as expected, at Room 714 we work on digital product architecture from layer design through to production deployment. A diagnostic conversation can save months of guesswork.






