Cloudflare Agent Memory Beta 2026: How to Build AI Agents That Remember Across Sessions [Code Tutorial]
What You'll Learn
- What Cloudflare Agent Memory stores and how its private-beta status changes the deployment decision.
- How namespaces, profiles, sessions, and four memory types define the storage model.
- How ingest, remember, recall, list, and deletion operations fit into a Worker-based agent.
- Which limits, privacy controls, and verification steps belong in a production design.
What Cloudflare Agent Memory Is and Is Not
Cloudflare Agent Memory is a managed service for keeping selected agent context outside the active model prompt. Cloudflare’s April 17, 2026 announcement describes it as a private beta that extracts useful information from conversations, stores that information in named profiles, and returns relevant context when an application asks for it.
The service is not a second model that remembers everything by itself. An application still decides which profile to use, when to ingest a conversation, when to expose recall to a model, and when a person must confirm a consequential action. The official Agent Memory documentation also describes the product as persistent and scoped rather than universal or shared by default.
The distinction matters for the word memory. A language model receives tokens for a call. Agent Memory stores records outside that call and later retrieves selected records. That makes memory an application capability with an API, a storage scope, and a deletion path. It does not give the model human experience, intention, or a guarantee that every recalled item is current.
Cloudflare’s product announcement presents several possible uses, including user preferences, company rules, support history, project state, and shared knowledge for teams. Those are design patterns, not proof that every workload will benefit. A short, stable conversation may need no durable memory at all. A long-running coding or support workflow may need it because repeated context reconstruction costs time and can lose important decisions.
For a broader explanation of the agent layer around memory, see Current Affair’s guide to planning, memory, and tool use. The practical question is where state should live and which part of the system may read or change it.
Why Durable Memory Sits Outside the Prompt
Keeping every conversation turn in the prompt seems simple until the conversation becomes long. More history increases input size and can make important details harder to find. Pruning solves the size problem but can remove a preference, a decision, or a prior failure that the agent needs later. Cloudflare positions Agent Memory between those two choices by extracting selected information and retrieving it on demand.
The service is designed around the agent context lifecycle. A harness drives model calls and tool calls. The model proposes text or actions. State includes the current conversation plus information stored outside it. When the harness compacts the conversation, it can send the relevant messages to Agent Memory for ingestion instead of discarding every older detail.
This does not mean that every token becomes a durable memory. The service extracts discrete items and classifies them. A useful memory should be specific enough to retrieve, have a clear source, and fit the application’s scope. “The user prefers concise answers” is more useful as a preference than a full transcript copied into every future prompt.
Cloudflare’s Agent Memory announcement explains that the service uses an opinionated extraction and retrieval design. The company’s stated goal is to keep the primary agent from spending its context on storage strategy. That can simplify an agent harness, but it also makes the service’s extraction decisions part of the application’s reliability surface.
Memory should be treated as evidence, not as an instruction that automatically overrides current policy. If a stored note says a user prefers one workflow, the application can use it to tailor a response. If the note would trigger a refund, access change, deployment, or payment, the current request and a confirmation rule should still control the action.
Current Affair’s agent operating-system analysis discusses why state and permissions matter when agents call external tools. Agent Memory covers one state layer. It does not replace authentication, authorization, audit logging, or business rules.
Namespaces, Profiles, and Sessions
Agent Memory uses a two-level isolation model. A namespace defines a memory domain for an application, environment, tenant, or memory layer. A profile is an isolated memory store for one entity, such as a user, team, agent, organization, tenant, or application object. A session groups memories that came from one interaction or conversation inside a profile.
The boundaries should be chosen before code is deployed. A support product may use one namespace for production and another for testing. It may then create one profile per customer or organization. A team knowledge profile can be shared by approved agents and people, while a personal preference profile should not be exposed to another account.
Cloudflare’s namespace and profile documentation describes the conceptual scope as namespace, then profile, then memory. A recall operation on one profile does not return records from another profile. Sessions add a further label for inspection and deletion, but they do not replace profile isolation.
| Scope | Typical use | Design question |
|---|---|---|
| Namespace | Application, environment, tenant group, or memory layer | Which systems are allowed to share this domain? |
| Profile | User, team, agent, organization, or object | Whose durable context belongs in this store? |
| Session | One conversation or interaction | Which messages and memories came from this run? |
| Memory | Fact, event, instruction, or task | What can be recalled, corrected, or deleted? |
Profile naming is an access decision, not just a string-format choice. A predictable profile name can make retrieval easy, but a weak mapping between an authenticated user and a profile can expose private information. Derive the profile from a trusted identity or tenant record, validate the mapping on the server, and do not let a model choose an arbitrary profile name.
Current Affair’s vector database comparison explains why logical isolation and retrieval indexes are separate concerns. Agent Memory provides its own profile boundary, while the application remains responsible for identity and access control.
The Four Memory Types
Cloudflare documents four memory categories. Facts represent stable knowledge about a person, project, or tool. Events represent completed actions anchored to a time. Instructions represent reusable procedures or conventions. Tasks represent short-lived work that is active within a session.
The categories create different update behavior. If a user changes a preference, the newer fact can supersede the older one on the same topic while the history remains available. If a deployment happens, it is an event that adds to the timeline rather than replacing every earlier deployment. If a team changes a runbook, the new instruction should supersede the older version. A task can become less relevant when the session ends.
Cloudflare’s memory-model documentation says ingestion extracts, classifies, deduplicates, and stores these types. The service also keeps raw conversation messages alongside extracted memories according to the documentation. That is useful for provenance, but it means retention, deletion, and sensitive-data rules need to be part of the design.
| Type | Example | Update behavior |
|---|---|---|
| Fact | A project uses a specified API or a user prefers a response style | Newer information can supersede the older topic value |
| Event | A deployment, decision, milestone, or observed incident | Events accumulate along a timeline |
| Instruction | A runbook, workflow, coding convention, or procedure | Newer instructions can supersede older versions |
| Task | An open investigation, follow-up, or current session item | Short-lived and less important after the session |
A category is not a truth score. A fact can still be wrong. An event can be recorded with the wrong date. An instruction can become unsafe after a system changes. Retrieval results should carry enough source and time context for the application or a reviewer to challenge them.
For applications that also use retrieval-augmented generation, Current Affair’s embeddings API comparison can help separate an embedding model decision from a memory policy decision. Similarity search is only one part of deciding what deserves durable storage.
Ingest, Remember, Recall, and List
The public announcement describes a small operation set that maps well to an agent loop. Ingest processes conversation messages and asks the service to extract durable items. Remember stores one item when the application already knows exactly what matters. Recall searches the stored profile and returns a synthesized answer grounded in matched content. List returns a manageable view of stored records for inspection. Get and delete operations are also available in the Workers API.
Use ingest for a batch, not as a reflex after every model turn. The official Cloudflare getting-started guide recommends ingesting after the user goes idle, when a conversation is compacted, or at another natural checkpoint. Repeated ingestion of the same conversation is idempotent according to the Workers API documentation.
Use remember when the application or an approved agent has a precise memory to store. A direct write tool should have a narrow contract and a policy for what can be saved. Do not give the model an unrestricted memory writer and assume that fluent output will produce good records.
Use recall when the current task depends on durable context that is not in the active conversation. The result is a synthesized answer, so the application should preserve the query, profile, session context, and any decision that follows. Use list, get, and summary operations for administration, inspection, and support rather than putting an entire profile into every model prompt.
| Operation | Best use | Control to add |
|---|---|---|
| Ingest | Extract durable items from a message batch | Batch at idle or compaction and preserve session ID |
| Remember | Store one known item immediately | Limit who or what may write and define confirmation rules |
| Recall | Find relevant context for a current task | Scope the profile and treat the answer as helpful context |
| List or get | Inspect stored memory and metadata | Restrict administrative access and paginate results |
| Delete | Remove a memory, session, profile, or namespace | Log the request and confirm the intended scope |
Deletion is part of the feature, not a cleanup task for later. A user may ask to remove one memory, a conversation, or every record in a profile. The application should know which operation matches that request and should avoid silently deleting a wider scope than the user intended.
How a Worker Connects to Agent Memory
Cloudflare exposes Agent Memory to Workers through an agent_memory binding. The binding points to a namespace. Worker code then gets a profile by name and calls the profile methods. The Workers API documentation provides the binding shape and method contracts.
The binding does not make memory available to the model automatically. Your Worker or agent framework must decide when to call recall, how to pass the result into the model context, and how to handle an empty answer. The application must also decide which authenticated request maps to which profile.
A minimal configuration pattern looks like this:
agent_memory binding = MEMORY namespace = support-prodThe exact Wrangler format depends on whether the project uses JSON or TOML. The important ideas are the binding name and the namespace. Keep the namespace in deployment configuration, not in user input. Use separate settings for development, staging, and production when those environments must not share memory.
Cloudflare says a profile is created when the application first accesses or writes to it. The first access to a new profile may take longer while the service creates the profile. A request handler should therefore handle startup latency and return a useful error if the memory service is unavailable.
Current Affair’s coding-assistant comparison is relevant when memory is added to a developer tool. A coding agent can store project conventions or prior decisions, but it should not treat a note as a substitute for the repository, tests, or the current source tree.
A Minimal Worker Integration Pattern
The most useful integration pattern has three paths. The first path reads durable context before a model call. The second path writes a conversation batch after a natural checkpoint. The third path handles explicit memory or deletion requests under application policy.
const profile = await env.MEMORY.getProfile(profileName)
const recalled = await profile.recall('project conventions for this task')
const context = recalled.answer || 'No relevant memory found'
await profile.ingest(messages, { sessionId: sessionId })This code is illustrative. It shows the order of operations, not a complete application. A real handler needs authentication, input validation, error handling, rate limits, logging, and a model call that receives the recalled context. It should also avoid ingesting the same batch repeatedly unless the idempotent behavior is intentional.
The profile name should be derived from the application’s data model. For a personal assistant, it could map to an authenticated user. For a support system, it might map to an organization or customer account. For a coding tool, it could map to a project or team. The correct choice depends on who should share the memory.
Do not expose every method to every model. A read-only recall tool may be appropriate for normal conversations. An explicit remember tool can be limited to user-approved commands. Delete should usually remain an application or support operation with a clear confirmation path.
Current Affair’s AI coding-agent guide gives context on tool-connected development systems. The same principle applies here: the model can propose a memory action, but the application should validate the scope before execution.
Batch Ingestion at Idle or Compaction
Calling ingest after every model turn is a poor default. It increases work, stores transient text that may not deserve persistence, and can turn an active conversation into a stream of duplicate or low-value records. Batch ingestion after the user pauses or after the harness compacts context gives the service a larger but more meaningful unit to analyze.
The getting-started guide presents a sample that schedules ingestion after user activity stops. It uses a cursor so the next batch contains messages that have not already been processed. The sample also associates the batch with a session ID. Those are implementation choices that developers can adapt to their own runtime.
A checkpoint should include more than “the user stopped typing.” Consider whether the conversation contains a stable preference, a completed event, a reusable instruction, or an active task. Do not store secrets, access tokens, payment details, or unrelated personal information merely because the text appeared in a message. Redaction before ingestion can reduce risk.
Ingestion is not the same as summarization. A summary compresses history for a model call. Ingestion extracts records that may be recalled later. A good pipeline can use both: summarize the current run for continuity, then ingest only durable items with the proper profile and session.
Cloudflare’s announcement describes raw messages being retained with extracted memories for provenance and search. That makes an idle checkpoint a data-governance event. Decide what retention period applies, who can inspect the raw transcript, and how a deletion request propagates to linked messages and memories.
Recall as a Model-Callable Tool
A memory binding in the Worker is not enough. The model needs a tool or context provider that describes when memory should be searched and how recalled information should be treated. Cloudflare’s getting-started guide uses the Agents SDK and Session API to expose memory recall as a model-callable search tool.
The instruction should be specific. Search memory when the request depends on prior preferences, project state, conventions, decisions, or long-running tasks. Do not search memory to repeat a fact the user just provided. When recall returns a result, use it as helpful context. When the memory would drive an irreversible action, confirm the important detail with the user.
That policy prevents two common errors. The first is over-searching, where every easy question triggers a memory call and adds latency. The second is over-trusting, where a stored note becomes an unreviewed command. Memory is a source of context. Current user intent, permissions, and business rules still have priority.
Cloudflare documents `recall()` as returning a synthesized answer grounded in stored content. If no memories match, it returns an empty answer. Your tool wrapper should preserve that distinction instead of filling the empty result with a guess.
Current Affair’s case-study article on agentic workflows shows why durable context should support a workflow rather than hide accountability. A service agent can recall a customer preference, but the application must still check the live account and policy before changing anything.
HTTP API Access and Authentication
Applications that do not run inside Workers can use Cloudflare’s HTTP API. The API uses the same namespace and profile model as the Workers interface, but requests are made to Cloudflare’s account API with an API token that has the appropriate Agent Memory permissions.
The HTTP API documentation shows the standard Bearer token pattern and endpoints for namespace management, profiles, sessions, ingest, remember, recall, list, get, delete, and summaries. Profiles are created automatically when first written.
Keep the API token on the server. Never place it in a browser bundle, a user prompt, a log message, or a memory record. Use the narrowest permission set available and separate tokens for different environments. If an external service calls the API on behalf of a user, authenticate the user before selecting the namespace and profile.
The HTTP interface also makes deletion and support workflows explicit. A profile delete marks the profile and its memories and messages for deletion. A session delete affects records tagged with that session. A memory delete removes the memory and linked source messages. Your user interface should explain which scope a request affects.
Cloudflare documents common HTTP errors including 400 for invalid namespace format, 401 for authentication failure, 404 for a missing namespace or profile, and 409 when a namespace already exists. Handle these as visible operational states. Do not ask a model to retry an authentication error with a different token.
Limits, Deletion, and Data Governance
Private-beta services should be integrated with their documented constraints visible in the design. Cloudflare’s limits page documents messages per ingest call, message size, recall query size, session and profile naming limits, namespace naming limits, and list pagination.
The documented limits are 500 messages per ingest call, 32 KB or 32,768 UTF-8 bytes for message content, 1 KB or 1,024 UTF-8 bytes for a recall query, 64 characters for a session ID, 100 characters for a profile name, 32 characters for a namespace name, and a list page size from 1 to 1,000 with a default of 20.
| Constraint | Documented value | Implementation response |
|---|---|---|
| Messages per ingest call | 500 | Chunk larger histories and track session progress |
| Message content | 32 KB or 32,768 UTF-8 bytes | Validate and split content before submission |
| Recall query | 1 KB or 1,024 UTF-8 bytes | Use a concise search topic and reject oversized input |
| Session and profile names | 64 and 100 characters | Generate stable bounded identifiers |
| Namespace name and list page | 32 characters and 1 to 1,000 records | Validate names and paginate administrative views |
Limits are not the only governance issue. The memory model can retain raw messages, extracted content, source references, and timestamps. Apply data minimization before ingest. Set retention and deletion procedures. Restrict profile access. Record which user or process requested a write. Provide a way to correct a memory when a user says it is wrong.
Deletion should be tested as a workflow. Verify that deleting one memory does not remove unrelated records. Verify that deleting one session leaves other sessions intact. Verify that profile deletion is reserved for the correct tenant or user. A successful API response is not enough if the application selected the wrong profile.
Current Affair’s embeddings comparison is useful when planning semantic retrieval, but an embedding index should not be allowed to define retention by itself. The product policy must decide what can remain searchable.
When the Private Beta Fits Your Architecture
Agent Memory fits best when an agent needs durable, scoped context across conversations and the team prefers a managed extraction and retrieval layer over building every memory component itself. It can be a reasonable candidate for user preferences, support history, project state, team conventions, and session continuity.
It may not be the right first step for a short-lived assistant, a workflow that already has a well-governed system of record, or a workload that requires a fully portable storage layer before the service’s export and API behavior meet the organization’s needs. Private beta also means availability, behavior, and documentation can change as Cloudflare refines the product.
Before adoption, answer five practical questions. Which namespace and profile should hold the data? Which events trigger ingest? Which model-callable actions are read-only? What happens when recall is empty or memory conflicts with live data? How will users inspect, correct, export, and delete their records?
Run a small evaluation with representative conversations. Measure extraction quality, recall relevance, stale-memory handling, latency, cost, deletion behavior, and failure recovery. Include prompt injection inside a conversation, a changed user preference, a cross-tenant access attempt, and an irreversible action that requires confirmation.
Cloudflare’s documentation gives the building blocks, not a finished data policy. The safest implementation treats the service as a managed state layer inside a larger agent system. The Worker owns identity and orchestration. Agent Memory handles the documented profile operations. The model receives only the context it needs. Verification and human approval remain available when the stakes require them.
That is the useful reading of “memory” in this private beta. It is not a promise that an agent remembers like a person. It is a structured way to extract, scope, retrieve, inspect, and delete selected context across sessions.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles