How to make your clients more context-aware with OpenMemory MCP

How to make your clients more context-aware with OpenMemory MCP

Taranjeet Singh
May 13, 2025

In the rapidly evolving landscape of AI, Large Language Models (LLMs) have transformed how we interact with technology. Yet, they face a fundamental limitation: they forget everything between sessions.

What if there could be a way to have a personal, portable LLM “memory layer” that lives locally on your system, with complete control over your data?

Today, we're excited to present OpenMemory MCP - a private, local-first memory layer powered by Mem0 that enables persistent, context-aware AI across MCP-compatible clients such as Cursor, Claude Desktop, Windsurf, Cline, and more. It provides a memory service that runs entirely on your machine, with a built-in UI for memory visibility, auditing, and control. It runs entirely on your local system, keeping your data under your control while enabling truly personalized experiences across any MCP-compatible tool.

This guide explains how to install, configure, and operate the OpenMemory MCP Server. It also covers the internal architecture, available features, and real-world applications.

What You'll Learn

What is OpenMemory MCP Server?

OpenMemory MCP is a private, local memory layer for MCP clients. It provides the infrastructure to store, manage and utilize your AI memories across different platforms, all while keeping your data locally on your system.

In simple terms, it's like a vector-backed memories layer for any LLM client using the standard MCP protocol and works out of the box with tools like Mem0.

Key Capabilities:

OpenMemory MCP Server serves as the backbone of your memory-aware AI stack, enabling clients to operate with shared, persistent context.

🔁 How it works (the basic flow)

Here’s the core flow in action:

  1. You can spin up OpenMemory (API, Qdrant, Postgres) with a single docker-compose command.
  2. The API process itself hosts an MCP server (using Mem0 under the hood) that speaks the standard MCP protocol over SSE.
  3. Your MCP client opens an SSE stream to OpenMemory’s /mcp/... endpoints and calls methods like add_memories(), search_memory() or list_memories().
  4. Everything else, including vector indexing, audit logs and access controls, is handled by the OpenMemory service.

Step-by-step guide to set up and run OpenMemory.

In this section, we will walk through how to set up OpenMemory and run it locally.

The project has two main components you will need to run:

Step 1: System Prerequisites

Before getting started, make sure your system has the following installed.

Please make sure Docker Desktop is running before proceeding to the next step.

Step 2: Clone the repo and set your OpenAI API Key

You need to clone the repo available at github.com/mem0/openmemory by using the following command.

git clone <https://github.com/mem0ai/mem0.git>
cd openmemory

Next, set your OpenAI API key as an environment variable.

export OPENAI_API_KEY=your_api_key_here

This command sets the key only for your current terminal session. It only lasts until you close that terminal window.

Step 3: Setup the backend

The backend runs in Docker containers. To start the backend, run these commands in the root directory:

# Copy the environment file and edit the file to update OPENAI_API_KEY and other secrets
make env

# Build all Docker images
make build

# Start Postgres, Qdrant, FastAPI/MCP server
make up

Once the setup is complete, your API will be live at http://localhost:8000.

Step 4: Setup the frontend

The frontend is a Next.js application. To start it, just run:

# Installs dependencies using pnpm and runs Next.js development server
make ui

After successful installation, you can navigate to http://localhost:3000 to check the OpenMemory dashboard, which will guide you through installing the MCP server in your MCP clients.

3. Features available in the dashboard (and what’s behind the UI)

The OpenMemory dashboard includes three main routes:

/ – dashboard
/memories – view and manage stored memories
/apps – view connected applications

  1. Install OpenMemory clients
  1. View memory and app stats
  1. Refresh or manually create a Memory

  2. You can open the filters panel to pick:

  1. You can inspect and manage individual Memories

Security, Access control and Architecture overview.

When working with the MCP protocol or any AI agent system, security becomes non-negotiable.

🎯 Security

OpenMemory is designed with privacy-first principles. It stores all memory data locally in your infrastructure, using Dockerized components (FastAPI, Postgres, Qdrant).

🎯 Access Control

Fine-grained access control is one of the core things focused on in OpenMemory.

🎯 Architecture

  1. Backend (FastAPI + FastMCP over SSE) :
  1. Vector Store (Qdrant via the mem0 client) : All memories are semantically indexed in Qdrant, with user and app-specific filters applied at query time.

  2. Relational Metadata (SQLAlchemy + Alembic) :

Together, this gives you a self-hosted LLM memory platform: ⚡ Store & version your chat memory in both a relational DB and a vector index. ⚡ Secure it with per-app ACLs and state transitions (active/paused/archived). ⚡ Search semantically via Qdrant. ⚡ Observe & control via dashboard.

Practical use cases with examples.

✅ Multi-agent research assistant with memory layer

Imagine building a tool where different LLM agents specialize in different research domains.

✅ Intelligent meeting assistant with persistent cross-session memory

We can build something like a meeting note-taker (Zoom, Google Meet, etc.) that:

✅ Agentic coding assistant that evolves with usage

Your CLI-based coding assistant can learn how you work by storing usage patterns, recurring questions, coding preferences and project-specific tips.

Now your MCP clients have real memory. You can track every access, pause what you want and audit everything in one dashboard. The best part is that everything is locally stored on your system.