How Memory Works in DeerFlow?

How Memory Works in DeerFlow?

Livia Ellen

April 1, 2026

The #1 repo on GitHub this week is a superagent harness, DeerFlow by ByteDance. It stores memory in JSON.

Most frameworks simulate memory by replaying chat history, stuffing everything into the context window and hoping the model picks what matters.

DeerFlow does something different.

It doesn’t store conversations. It extracts facts about the user, scores them by confidence, and injects what fits within a 2,000-token budget into each prompt, async, without bloating the context.

In this article, I’ll break down how it works based on reverse-engineering a self-hosted DeerFlow instance I run in my docker and inspecting its JSON memory.

What Is DeerFlow

DeerFlow (Deep Exploration and Efficient Research Flow) is an open-source super agent harness that orchestrates sub-agents, memory, and sandboxes to do almost anything, powered by extensible skills.

They are currently trending #1 in Github with 49.2K stars. Check out their repository here.

The DeerFlow UI looks like a ChatGPT or Claude interface where you have chat interface + agent options on your left sidebar.

I tested it by having conversations for about 3 hours, running it locally on Docker Container in my Macbook.

What I tested:

The part that doesn’t get enough attention is the memory layer.

The Memory Interface

Memory Panel UI in DeerFlow (Screenshot from Personal Memory)

The Memory panel organizes everything into:

You can see everything the system learned about you, when it learned it, and how confident it is.

The source field on each fact is the actual thread UUID, so the system knows exactly which conversation each piece of knowledge came from.

Where the memory sits

Memory in DeerFlow lives as a file on disk: backend/.deer-flow/memory.json.

memory.json file in a self hosted deerflow

It’s local and it persists across every session. A structured JSON file.

Within one session, DeerFlow builds a structured profile. Each fact has similar content like we have seen in the Memory panel UI. Facts below 0.7 confidence don’t get included. The store caps at 100 facts total, evicting the lowest-confidence ones first when it overflows.

How It Actually Works

DeerFlow’s Lead Agent runs on LangGraph, a graph-based orchestration framework where each agent turn is a node in a stateful execution graph.

Memory isn’t a plugin or sidecar. It’s baked directly into the middleware chain as position #8, meaning it runs on every agent turn automatically.

The key component is MemoryMiddleware.

Deerflow Middleware Chain (generated by Claude)

MemoryMiddleware sits right after TitleMiddleware and before the vision and loop-detection layers.

This is intentional, memory updates should happen after the title is generated (so the LLM knows what the conversation was about) but before any loop or clarification checks cut the session short. MemoryMiddleware doesn't update memory synchronously. It queues the conversation for async processing.

DeerFlow Memory System Write Flow

The flow works like this:

Before appending a new fact, the system checks for exact content duplication (normalized by stripping whitespace). Note: different phrasing of the same semantic fact will still get added, the dedup is text-based, not semantic.

You can point model_name at a cheaper model for the extraction step, the memory LLM doesn't need to be your best model, it just needs to be good at structured extraction.

memory config yaml

What I found interesting on step 3 is if the same thread_id already has a pending update in the queue, the new entry replaces it rather than appending. You never process stale mid-conversation snapshots, only the final state of each thread gets processed.

Thread replacement Queue

Memory ingestion doesn’t affect the user experience on getting response from the agent from latency perspective. If a user sends a message, they will get a response right away. This is example of a good harness in agentic system.

Debounce Timeline

Where the Memory Actually Shows Up

When you start a new conversation, every fact that fits within a 2,000-token budget, sorted by confidence, plus all the user, history, and topOfMind summaries get injected into the system prompt inside tags. There's no hardcoded "top N" limit; it's a token budget, not a count. Tiktoken counts the tokens precisely as facts are added one by one until the budget runs out.

DeerFlow memory.json (Image generated with Claude)

The agent sees something like this at the start of every session, from my actual session, formatted exactly as format_memory_for_injection().

The agent never has to ask "what are you working on?" It already knows based on my Top of Mind History.

What It Looks Like In Practice

After a 3 hours session, this is what DeerFlow showed in its Fact section in Memory panel. This is what's being compacted into 2000 token.

My Facts section in my memory after 3 hours

Memory Retrieval Test

I tried testing more to understand if it’s able to retrieve my information based on the conversation I had with them.

Retrieval worked as expected.

Simple Retrieval Check

I tried deleting the memories through the chat. It only adds more preference in the fact table that "I don't want to talk about this topic despite the LLM Extractor steps showing it has factsToRemove diff recorded.

LLM Extractor failed to delete facts

There is definitely a workaround, which is modifying the json within the codebase. This shows limitations on this memory system.

Another limitation I encountered was when I talked about memory benchmark in one of my sessions, despite my user profile showing English preference, it shows questions recommendation in Chinese.

Recommended Question in Chinese

Conclusion

Highlight from Deerflow Memory

What It Doesn't Do

Why It Matters

Most agent frameworks treat memory as a retrieval problem: embed everything, store it in a vector database, and retrieve the most similar chunks at query time.

That works, but it comes at a cost. You’re adding an embedding model, a vector store, and a retrieval step to every request. That means more infrastructure, more latency, and more complexity.

DeerFlow takes a different approach: don’t store conversations, store understanding.

It runs an asynchronous LLM extraction pass that distills raw conversations into structured facts, then injects those facts directly into future prompts. The expensive step happens after the response is already delivered, on a debounced 30-second timer, using any model you choose.

Overall, it delivers on its promise as a SuperAgent harness.

The memory system is efficient and production-ready in its current form, with some clear limitations.

It combines an agentic approach to scoring and injecting facts, but still lacks a robust memory intelligence layers on some usecase. Overall, as a SuperAgent Harness deep research and agentic tool. Deerflow team did a great work on implementing memory system on this release.

The result is an agent that knows you, not because it remembers everything you said, but because it builds a persistent model of your preferences, goals, and context. And that model lives as a simple JSON file on your machine.

DeerFlow builds memory directly into its middleware, so it works out of the box, on every conversation, by default.

That’s the part worth paying attention to.