How to make your clients more context-aware with OpenMemory MCP
How to make your clients more context-aware with OpenMemory MCP
Taranjeet Singh
May 13, 2025
In the rapidly evolving landscape of AI, Large Language Models (LLMs) have transformed how we interact with technology. Yet, they face a fundamental limitation: they forget everything between sessions.
What if there could be a way to have a personal, portable LLM “memory layer” that lives locally on your system, with complete control over your data?
Today, we're excited to present OpenMemory MCP - a private, local-first memory layer powered by Mem0 that enables persistent, context-aware AI across MCP-compatible clients such as Cursor, Claude Desktop, Windsurf, Cline, and more. It provides a memory service that runs entirely on your machine, with a built-in UI for memory visibility, auditing, and control. It runs entirely on your local system, keeping your data under your control while enabling truly personalized experiences across any MCP-compatible tool.
This guide explains how to install, configure, and operate the OpenMemory MCP Server. It also covers the internal architecture, available features, and real-world applications.
What You'll Learn
- What OpenMemory MCP Server is and why it matters
- Step-by-step setup guide
- Dashboard features and UI breakdown
- Security, Access Control, and Architecture overview
- Practical use cases with examples
What is OpenMemory MCP Server?
OpenMemory MCP is a private, local memory layer for MCP clients. It provides the infrastructure to store, manage and utilize your AI memories across different platforms, all while keeping your data locally on your system.
In simple terms, it's like a vector-backed memories layer for any LLM client using the standard MCP protocol and works out of the box with tools like Mem0.
Key Capabilities:
- Add, retrieve, list and delete memory objects via MCP server tools (
add_memories,search_memory,list_memories,delete_all_memories). - Store semantically indexed data using
Qdrantas a vector store under the hood. - Runs fully on your local infrastructure (
Docker + Postgres + Qdrant) with no data sent outside. - Pause or revoke any client’s access at the
app or memory level, with audit logs for every read or write. - Built-in UI for observability and manual control
OpenMemory MCP Server serves as the backbone of your memory-aware AI stack, enabling clients to operate with shared, persistent context.
🔁 How it works (the basic flow)
Here’s the core flow in action:
- You can spin up OpenMemory (API, Qdrant, Postgres) with a single
docker-composecommand. - The API process itself hosts an MCP server (using Mem0 under the hood) that speaks the standard MCP protocol over SSE.
- Your MCP client opens an SSE stream to OpenMemory’s
/mcp/...endpoints and calls methods likeadd_memories(),search_memory()orlist_memories(). - Everything else, including vector indexing, audit logs and access controls, is handled by the OpenMemory service.
Step-by-step guide to set up and run OpenMemory.
In this section, we will walk through how to set up OpenMemory and run it locally.
The project has two main components you will need to run:
api/- Backend APIs and MCP serverui/- Frontend React application (dashboard)
Step 1: System Prerequisites
Before getting started, make sure your system has the following installed.
- Docker and Docker Compose
- Python 3.9+ - required for backend development
- Node.js - required for frontend development
- OpenAI API Key - used for LLM interactions
- GNU Make
Please make sure Docker Desktop is running before proceeding to the next step.
Step 2: Clone the repo and set your OpenAI API Key
You need to clone the repo available at github.com/mem0/openmemory by using the following command.
git clone <https://github.com/mem0ai/mem0.git>
cd openmemory
Next, set your OpenAI API key as an environment variable.
export OPENAI_API_KEY=your_api_key_here
This command sets the key only for your current terminal session. It only lasts until you close that terminal window.
Step 3: Setup the backend
The backend runs in Docker containers. To start the backend, run these commands in the root directory:
# Copy the environment file and edit the file to update OPENAI_API_KEY and other secrets
make env
# Build all Docker images
make build
# Start Postgres, Qdrant, FastAPI/MCP server
make up
Once the setup is complete, your API will be live at http://localhost:8000.
Step 4: Setup the frontend
The frontend is a Next.js application. To start it, just run:
# Installs dependencies using pnpm and runs Next.js development server
make ui
After successful installation, you can navigate to http://localhost:3000 to check the OpenMemory dashboard, which will guide you through installing the MCP server in your MCP clients.
3. Features available in the dashboard (and what’s behind the UI)
The OpenMemory dashboard includes three main routes:
/ – dashboard/memories – view and manage stored memories/apps – view connected applications
- Install OpenMemory clients
- Get your unique SSE endpoint or use a one-liner install command
- View memory and app stats
- See how many memories you’ve stored
- See how many apps are connected
Refresh or manually create a Memory
You can open the filters panel to pick:
- Which apps to include
- Which categories to include
- You can inspect and manage individual Memories
Security, Access control and Architecture overview.
When working with the MCP protocol or any AI agent system, security becomes non-negotiable.
🎯 Security
OpenMemory is designed with privacy-first principles. It stores all memory data locally in your infrastructure, using Dockerized components (FastAPI, Postgres, Qdrant).
🎯 Access Control
Fine-grained access control is one of the core things focused on in OpenMemory.
🎯 Architecture
- Backend (
FastAPI + FastMCP over SSE) :
- exposes both a plain-old REST surface (
/api/v1/memories,/api/v1/apps,/api/v1/stats) & - an MCP “tool” interface (
/mcp/messages,/mcp/sse/<client>/<user>) that agents use to call (add_memories,search_memory,list_memories) via Server-Sent Events (SSE).
Vector Store (
Qdrant via the mem0 client) : All memories are semantically indexed in Qdrant, with user and app-specific filters applied at query time.Relational Metadata (
SQLAlchemy + Alembic) :
- track users, apps, memory entries, access logs, categories and access controls.
Together, this gives you a self-hosted LLM memory platform: ⚡ Store & version your chat memory in both a relational DB and a vector index. ⚡ Secure it with per-app ACLs and state transitions (active/paused/archived). ⚡ Search semantically via Qdrant. ⚡ Observe & control via dashboard.
Practical use cases with examples.
✅ Multi-agent research assistant with memory layer
Imagine building a tool where different LLM agents specialize in different research domains.
- Each agent stores what it finds via
add_memories(text)and a master agent later runssearch_memory(query)across all previous results.
✅ Intelligent meeting assistant with persistent cross-session memory
We can build something like a meeting note-taker (Zoom, Google Meet, etc.) that:
- Extracts summaries via LLMs.
- Remembers action items across calls.
✅ Agentic coding assistant that evolves with usage
Your CLI-based coding assistant can learn how you work by storing usage patterns, recurring questions, coding preferences and project-specific tips.
Now your MCP clients have real memory. You can track every access, pause what you want and audit everything in one dashboard. The best part is that everything is locally stored on your system.