The hardest engineering problem in AI products is not the model; it is memory. What does the system remember between sessions, between users, between tenants? Get it wrong and you ship a product that leaks customer data, hallucinates context, or forgets what it was supposed to be doing.
The failure is rarely dramatic. It looks like an agent casually mentioning a fact from a different customer's workspace, or confidently applying last quarter's pricing to this quarter's proposal, or greeting a returning user with context from someone else's session. Each incident is small; the trust damage is total, because the user can no longer predict what the system knows.
Three layers, kept separate
Session memory: what the model needs in this conversation. Short, transient, scoped to one interaction. User memory: long-term, preferences and past tasks for a single user. Persisted, encrypted, never crossed with other users. Domain memory: knowledge about the customer's data and workflows. Tenant-isolated by default; never shared across customers.
The layers need different write policies, not just different stores. Session memory is written freely and discarded whole. User memory is written selectively: distilled facts and stated preferences, never raw transcripts, because raw transcripts accumulate stale and contradictory context that the model will faithfully resurface at the worst moment. Domain memory is written only through the same permission checks a human would face; the retrieval layer must carry the tenant's ACLs, so the model can never read what the logged-in user could not.
The isolation guarantee
The system has to be designed so that domain memory cannot leak across tenants under any prompt. This is the single biggest reason enterprise buyers reject AI products: they cannot verify the isolation. Make it architecturally impossible, not just policy-prohibited. Defaults matter more than promises.
Architecturally impossible has a concrete meaning: the tenant ID is bound at the connection or index level, below the layer any prompt can reach. Per-tenant collections or row-level security enforced by the database, credentials scoped so the retrieval service for tenant A cannot form a query against tenant B, and no code path where the model composes the tenant filter itself. If a prompt injection would have to break Postgres row-level security to leak data, you can put that sentence in a security review and survive the follow-up questions. If isolation is a WHERE clause the application remembers to add, you cannot.
Forgetting is a feature
Most teams obsess over what the system remembers and ignore the harder half: what it must forget, and when. Memory needs expiry by default (session context dies with the session, distilled user facts carry a last-confirmed date and decay when stale), correction ("that is no longer true" must actually delete, not append a contradiction), and deletion that propagates through every store, including embeddings and caches, within a bounded time you are willing to write into a DPA. A memory system without deletion mechanics is a compliance incident on a timer.
What the user controls
Users should be able to see what the system remembers about them, edit it, and delete it. This is a product feature, not a privacy footnote. Memory that the user does not control is memory the user will not trust.
The test worth shipping against: a user should be able to answer "why did it say that?" in one click, tracing any personalised behaviour back to the remembered fact that produced it, and remove that fact on the spot. Products that pass this test convert memory from a liability conversation into their stickiest feature, because a memory the user has curated is a memory they will not want to rebuild elsewhere.