Reranker-Enhanced Search - Mem0

Reranker-enhanced Search

Reranker-enhanced search adds a second scoring pass after vector retrieval so Mem0 can return the most relevant memories first. Enable it when keyword similarity alone misses nuance or when you need the highest-confidence context for an agent decision.

You’ll use this when…

Reranking raises latency and, for hosted models, API spend. Benchmark with production traffic and define a fallback path for latency-sensitive requests.

The Configure it and See it in action snippets below use the Python SDK. The self-hosted TypeScript SDK supports the Cohere, Zero Entropy, Sentence Transformer, Hugging Face, and LLM rerankers; see TypeScript SDK.

TypeScript SDK

The self-hosted TypeScript SDK (mem0ai/oss) ships five rerankers: Cohere, Zero Entropy, Sentence Transformer, Hugging Face, and the LLM reranker. Configure one under reranker, then opt in per search with rerank: true. Keys are camelCase (apiKey, not api_key).Provider SDKs are peer dependencies. Install the one your reranker needs:

pnpm add cohere-ai          # cohere
pnpm add zeroentropy        # zero_entropy
pnpm add @huggingface/transformers   # sentence_transformer, huggingface
# llm_reranker defaults to openai (already a core dependency); install another
# provider's SDK only if you nest a different one under config.llm

Hosted Rerankers (Cohere, Zero Entropy)

Both call a hosted API and read their key from config or the provider’s environment variable (COHERE_API_KEY, ZERO_ENTROPY_API_KEY).

import { Memory } from "mem0ai/oss";

// Cohere reranker (defaults to the rerank-v3.5 model)
const memory = new Memory({
  reranker: {
    provider: "cohere",
    config: { apiKey: process.env.COHERE_API_KEY },
  },
});

const results = await memory.search("What are my food preferences?", {
  filters: { userId: "alice" },
  rerank: true,
});
// Zero Entropy reranker (defaults to the zerank-1 model)
const memory = new Memory({
  reranker: {
    provider: "zero_entropy",
    config: { apiKey: process.env.ZERO_ENTROPY_API_KEY },
  },
});

Local Cross-Encoders (Sentence Transformer, Hugging Face)

Both run a cross-encoder locally with Transformers.js: no API key, no network at inference time. Because Transformers.js runs ONNX weights, the default models are the ONNX mirrors of the Python SDK’s defaults (sentence_transformerXenova/ms-marco-MiniLM-L-6-v2, huggingfaceXenova/bge-reranker-base). Point model at any ONNX-exported cross-encoder on the Hub to override.

const memory = new Memory({
  reranker: {
    provider: "sentence_transformer", // or "huggingface"
    config: {
      // model: "Xenova/bge-reranker-base",  // override the default
      device: "cpu", // Transformers.js device: "cpu" | "wasm" | "webgpu"
      maxLength: 512, // max tokens per query-document pair
      normalize: true, // sigmoid-normalize logits to [0, 1] (default)
    },
  },
});

const results = await memory.search("What movies do I like?", {
  filters: { userId: "alice" },
  rerank: true,
});

batchSize and showProgressBar are accepted for config parity with the Python SDK but are no-ops in this runtime, because a memory search reranks a small candidate set in a single in-process forward pass. The model is downloaded once and cached in-process on first use.

LLM Reranker

To score with an LLM instead of a dedicated reranker, use the llm_reranker provider. It builds its own LLM from the reranker’s config (defaulting to openai / gpt-4o-mini) rather than reusing the Memory’s main llm:

const memory = new Memory({
  reranker: {
    provider: "llm_reranker",
    config: { apiKey: process.env.OPENAI_API_KEY },
  },
});

const results = await memory.search("What movies do I like?", {
  filters: { userId: "alice" },
  rerank: true,
});

Nest a different provider under config.llm to override the default:

const memory = new Memory({
  reranker: {
    provider: "llm_reranker",
    config: {
      llm: {
        provider: "anthropic",
        config: { apiKey: process.env.ANTHROPIC_API_KEY },
      },
    },
  },
});

Config Reference

Provider Default model Key config fields
cohere rerank-v3.5 apiKey, model, topK
zero_entropy zerank-1 apiKey, model, topK
sentence_transformer Xenova/ms-marco-MiniLM-L-6-v2 model, device, maxLength, normalize, topK
huggingface Xenova/bge-reranker-base model, device, maxLength, normalize, topK
llm_reranker openai / gpt-4o-mini provider, model, apiKey, llm (nested override), topK

rerank is opt-in per search and a no-op when no reranker is configured. If the reranker call fails, Mem0 logs a warning and returns the original vector-ranked results.

Feature Anatomy

Supported Providers

Provider Comparison

Provider Latency Quality Cost Local deploy
Cohere Medium High API cost
Sentence Transformer Low Good Free
Hugging Face Low–Medium Variable Free
LLM Reranker High Very high API cost Depends

Basic Setup

from mem0 import Memory

config = {
    "reranker": {
        "provider": "cohere",
        "config": {
            "model": "rerank-v3.5",
            "api_key": "your-cohere-api-key"
        }
    }
}

m = Memory.from_config(config)

Confirm results["results"][0]["score"] reflects the reranker output: if the field is missing, the reranker was not applied. Set top_k to the smallest candidate pool that still captures relevant hits. Smaller pools keep reranking costs down.

Provider-Specific Options

# Cohere reranker
config = {
    "reranker": {
        "provider": "cohere",
        "config": {
            "model": "rerank-v3.5",
            "api_key": "your-cohere-api-key",
            "top_k": 10,
            "return_documents": True
        }
    }
}

# Sentence Transformer reranker
config = {
    "reranker": {
        "provider": "sentence_transformer",
        "config": {
            "model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
            "device": "cuda",
            "max_length": 512
        }
    }
}

# Hugging Face reranker
config = {
    "reranker": {
        "provider": "huggingface",
        "config": {
            "model": "BAAI/bge-reranker-base",
            "device": "cuda",
            "batch_size": 32
        }
    }
}

# LLM-based reranker
config = {
    "reranker": {
        "provider": "llm_reranker",
        "config": {
            "provider": "openai",
            "model": "gpt-5-mini",
            "api_key": "your-openai-api-key",
            "top_k": 5
        }
    }
}

Keep authentication keys in environment variables when you plug these configs into production projects.

Full Stack Example

config = {
    "vector_store": {
        "provider": "qdrant",
        "config": {
            "host": "localhost",
            "port": 6333
        }
    },
    "llm": {
        "provider": "openai",
        "config": {
            "model": "gpt-5-mini",
            "api_key": "your-openai-api-key"
        }
    },
    "embedder": {
        "provider": "openai",
        "config": {
            "model": "text-embedding-3-small",
            "api_key": "your-openai-api-key"
        }
    },
    "reranker": {
        "provider": "cohere",
        "config": {
            "model": "rerank-v3.5",
            "api_key": "your-cohere-api-key",
            "top_k": 15,
            "return_documents": True
        }
    }
}

m = Memory.from_config(config)

A quick search should now return results with both vector and reranker scores, letting you compare improvements immediately.

Async Support

from mem0 import AsyncMemory

async_memory = AsyncMemory.from_config(config)

async def search_with_rerank():
    return await async_memory.search(
        "What are my preferences?",
        filters={"user_id": "alice"},
        rerank=True
    )

import asyncio
results = asyncio.run(search_with_rerank())

Inspect the async response to confirm reranking still applies; the scores should match the synchronous implementation.

Tune Performance and Cost

# GPU-friendly local reranker configuration
config = {
    "reranker": {
        "provider": "sentence_transformer",
        "config": {
            "model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
            "device": "cuda",
            "batch_size": 32,
            "top_k": 10,
            "max_length": 256
        }
    }
}

# Smart toggle for hosted rerankers
def smart_search(query, user_id, use_rerank=None):
    if use_rerank is None:
        use_rerank = len(query.split()) > 3
    return m.search(query, filters={"user_id": user_id}, rerank=use_rerank)

Use heuristics (query length, user tier) to decide when to rerank so high-signal queries benefit without taxing every request.

Handle Failures Gracefully

try:
    results = m.search("test query", filters={"user_id": "alice"}, rerank=True)
except Exception as exc:
    print(f"Reranking failed: {exc}")
    results = m.search("test query", filters={"user_id": "alice"}, rerank=False)

Always fall back to vector-only search: dropped queries introduce bigger accuracy issues than slightly less relevant ordering.

Migrate from v0.x

# Before: basic vector search
results = m.search("query", filters={"user_id": "alice"})

# After: same API with reranking enabled via config
config = {
    "reranker": {
        "provider": "sentence_transformer",
        "config": {
            "model": "cross-encoder/ms-marco-MiniLM-L-6-v2"
        }
    }
}

m = Memory.from_config(config)
results = m.search("query", filters={"user_id": "alice"})

See It in Action

Basic Reranked Search

results = m.search(
    "What are my food preferences?",
    filters={"user_id": "alice"}
)

for result in results["results"]:
    print(f"Memory: {result['memory']}")
    print(f"Score: {result['score']}")

Expect each result to list the reranker-adjusted score so you can compare ordering against baseline vector results.

Toggle Reranking Per Request

results_with_rerank = m.search(
    "What movies do I like?",
    filters={"user_id": "alice"},
    rerank=True
)

results_without_rerank = m.search(
    "What movies do I like?",
    filters={"user_id": "alice"},
    rerank=False
)

Log the reranked vs. non-reranked lists during rollout so stakeholders can see the improvement before enforcing it everywhere.

Combine With Metadata Filters

results = m.search(
    "important work tasks",
    filters={
        "AND": [\
            {"user_id": "alice"},\
            {"category": "work"},\
            {"priority": {"gte": 7}}\
        ]
    },
    rerank=True,
    top_k=20
)

Verify filtered reranked searches still respect every metadata clause: reranking only reorders candidates; it never bypasses filters.

Real-World Playbooks

Customer Support

config = {
    "reranker": {
        "provider": "cohere",
        "config": {
            "model": "rerank-v3.5",
            "api_key": "your-cohere-api-key"
        }
    }
}

m = Memory.from_config(config)

results = m.search(
    "customer having login issues with mobile app",
    filters={"agent_id": "support_bot", "category": "technical_support"},
    rerank=True
)

Top results should highlight tickets matching the login issue context so agents can respond faster.

Content Recommendation

results = m.search(
    "science fiction books with space exploration themes",
    filters={"user_id": "reader123", "content_type": "book_recommendation"},
    rerank=True,
    top_k=10
)

for result in results["results"]:
    print(f"Recommendation: {result['memory']}")
    print(f"Relevance: {result['score']:.3f}")

Expect high-scoring recommendations that match both the requested theme and any metadata limits you applied.

Verify The Feature Is Working