Reranker-Enhanced Search - Mem0
Reranker-enhanced Search
Reranker-enhanced search adds a second scoring pass after vector retrieval so Mem0 can return the most relevant memories first. Enable it when keyword similarity alone misses nuance or when you need the highest-confidence context for an agent decision.
You’ll use this when…
- Queries are nuanced and require semantic understanding beyond vector distance.
- Large memory collections produce too many near matches to review manually.
- You want consistent scoring across providers by delegating ranking to a dedicated model.
Reranking raises latency and, for hosted models, API spend. Benchmark with production traffic and define a fallback path for latency-sensitive requests.
The Configure it and See it in action snippets below use the Python SDK. The self-hosted TypeScript SDK supports the Cohere, Zero Entropy, Sentence Transformer, Hugging Face, and LLM rerankers; see TypeScript SDK.
TypeScript SDK
The self-hosted TypeScript SDK (mem0ai/oss) ships five rerankers: Cohere, Zero Entropy, Sentence Transformer, Hugging Face, and the LLM reranker. Configure one under reranker, then opt in per search with rerank: true. Keys are camelCase (apiKey, not api_key).Provider SDKs are peer dependencies. Install the one your reranker needs:
pnpm add cohere-ai # cohere
pnpm add zeroentropy # zero_entropy
pnpm add @huggingface/transformers # sentence_transformer, huggingface
# llm_reranker defaults to openai (already a core dependency); install another
# provider's SDK only if you nest a different one under config.llm
Hosted Rerankers (Cohere, Zero Entropy)
Both call a hosted API and read their key from config or the provider’s environment variable (COHERE_API_KEY, ZERO_ENTROPY_API_KEY).
import { Memory } from "mem0ai/oss";
// Cohere reranker (defaults to the rerank-v3.5 model)
const memory = new Memory({
reranker: {
provider: "cohere",
config: { apiKey: process.env.COHERE_API_KEY },
},
});
const results = await memory.search("What are my food preferences?", {
filters: { userId: "alice" },
rerank: true,
});
// Zero Entropy reranker (defaults to the zerank-1 model)
const memory = new Memory({
reranker: {
provider: "zero_entropy",
config: { apiKey: process.env.ZERO_ENTROPY_API_KEY },
},
});
Local Cross-Encoders (Sentence Transformer, Hugging Face)
Both run a cross-encoder locally with Transformers.js: no API key, no network at inference time. Because Transformers.js runs ONNX weights, the default models are the ONNX mirrors of the Python SDK’s defaults (sentence_transformer → Xenova/ms-marco-MiniLM-L-6-v2, huggingface → Xenova/bge-reranker-base). Point model at any ONNX-exported cross-encoder on the Hub to override.
const memory = new Memory({
reranker: {
provider: "sentence_transformer", // or "huggingface"
config: {
// model: "Xenova/bge-reranker-base", // override the default
device: "cpu", // Transformers.js device: "cpu" | "wasm" | "webgpu"
maxLength: 512, // max tokens per query-document pair
normalize: true, // sigmoid-normalize logits to [0, 1] (default)
},
},
});
const results = await memory.search("What movies do I like?", {
filters: { userId: "alice" },
rerank: true,
});
batchSize and showProgressBar are accepted for config parity with the Python SDK but are no-ops in this runtime, because a memory search reranks a small candidate set in a single in-process forward pass. The model is downloaded once and cached in-process on first use.
LLM Reranker
To score with an LLM instead of a dedicated reranker, use the llm_reranker provider. It builds its own LLM from the reranker’s config (defaulting to openai / gpt-4o-mini) rather than reusing the Memory’s main llm:
const memory = new Memory({
reranker: {
provider: "llm_reranker",
config: { apiKey: process.env.OPENAI_API_KEY },
},
});
const results = await memory.search("What movies do I like?", {
filters: { userId: "alice" },
rerank: true,
});
Nest a different provider under config.llm to override the default:
const memory = new Memory({
reranker: {
provider: "llm_reranker",
config: {
llm: {
provider: "anthropic",
config: { apiKey: process.env.ANTHROPIC_API_KEY },
},
},
},
});
Config Reference
| Provider | Default model | Key config fields |
|---|---|---|
cohere |
rerank-v3.5 |
apiKey, model, topK |
zero_entropy |
zerank-1 |
apiKey, model, topK |
sentence_transformer |
Xenova/ms-marco-MiniLM-L-6-v2 |
model, device, maxLength, normalize, topK |
huggingface |
Xenova/bge-reranker-base |
model, device, maxLength, normalize, topK |
llm_reranker |
openai / gpt-4o-mini |
provider, model, apiKey, llm (nested override), topK |
rerank is opt-in per search and a no-op when no reranker is configured. If the reranker call fails, Mem0 logs a warning and returns the original vector-ranked results.
Feature Anatomy
- Initial vector search: Retrieve candidate memories by similarity.
- Reranker pass: A specialized model scores each candidate against the original query.
- Reordered results: Mem0 sorts responses using the reranker’s scores before returning them.
- Optional fallbacks: Toggle reranking per request or disable it entirely if performance or cost becomes a concern.
Supported Providers
- Cohere: Multilingual hosted reranker with API-based scoring.
- Sentence Transformer: Local Hugging Face cross-encoders for GPU or CPU.
- Hugging Face: Bring any hosted or on-prem reranker model ID.
- LLM Reranker: Use your preferred LLM (OpenAI, etc.) for prompt-driven scoring.
- Zero Entropy: High-quality neural reranking tuned for retrieval tasks.
Provider Comparison
| Provider | Latency | Quality | Cost | Local deploy |
|---|---|---|---|---|
| Cohere | Medium | High | API cost | ❌ |
| Sentence Transformer | Low | Good | Free | ✅ |
| Hugging Face | Low–Medium | Variable | Free | ✅ |
| LLM Reranker | High | Very high | API cost | Depends |
Basic Setup
from mem0 import Memory
config = {
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-v3.5",
"api_key": "your-cohere-api-key"
}
}
}
m = Memory.from_config(config)
Confirm results["results"][0]["score"] reflects the reranker output: if the field is missing, the reranker was not applied.
Set top_k to the smallest candidate pool that still captures relevant hits. Smaller pools keep reranking costs down.
Provider-Specific Options
# Cohere reranker
config = {
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-v3.5",
"api_key": "your-cohere-api-key",
"top_k": 10,
"return_documents": True
}
}
}
# Sentence Transformer reranker
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "cuda",
"max_length": 512
}
}
}
# Hugging Face reranker
config = {
"reranker": {
"provider": "huggingface",
"config": {
"model": "BAAI/bge-reranker-base",
"device": "cuda",
"batch_size": 32
}
}
}
# LLM-based reranker
config = {
"reranker": {
"provider": "llm_reranker",
"config": {
"provider": "openai",
"model": "gpt-5-mini",
"api_key": "your-openai-api-key",
"top_k": 5
}
}
}
Keep authentication keys in environment variables when you plug these configs into production projects.
Full Stack Example
config = {
"vector_store": {
"provider": "qdrant",
"config": {
"host": "localhost",
"port": 6333
}
},
"llm": {
"provider": "openai",
"config": {
"model": "gpt-5-mini",
"api_key": "your-openai-api-key"
}
},
"embedder": {
"provider": "openai",
"config": {
"model": "text-embedding-3-small",
"api_key": "your-openai-api-key"
}
},
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-v3.5",
"api_key": "your-cohere-api-key",
"top_k": 15,
"return_documents": True
}
}
}
m = Memory.from_config(config)
A quick search should now return results with both vector and reranker scores, letting you compare improvements immediately.
Async Support
from mem0 import AsyncMemory
async_memory = AsyncMemory.from_config(config)
async def search_with_rerank():
return await async_memory.search(
"What are my preferences?",
filters={"user_id": "alice"},
rerank=True
)
import asyncio
results = asyncio.run(search_with_rerank())
Inspect the async response to confirm reranking still applies; the scores should match the synchronous implementation.
Tune Performance and Cost
# GPU-friendly local reranker configuration
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2",
"device": "cuda",
"batch_size": 32,
"top_k": 10,
"max_length": 256
}
}
}
# Smart toggle for hosted rerankers
def smart_search(query, user_id, use_rerank=None):
if use_rerank is None:
use_rerank = len(query.split()) > 3
return m.search(query, filters={"user_id": user_id}, rerank=use_rerank)
Use heuristics (query length, user tier) to decide when to rerank so high-signal queries benefit without taxing every request.
Handle Failures Gracefully
try:
results = m.search("test query", filters={"user_id": "alice"}, rerank=True)
except Exception as exc:
print(f"Reranking failed: {exc}")
results = m.search("test query", filters={"user_id": "alice"}, rerank=False)
Always fall back to vector-only search: dropped queries introduce bigger accuracy issues than slightly less relevant ordering.
Migrate from v0.x
# Before: basic vector search
results = m.search("query", filters={"user_id": "alice"})
# After: same API with reranking enabled via config
config = {
"reranker": {
"provider": "sentence_transformer",
"config": {
"model": "cross-encoder/ms-marco-MiniLM-L-6-v2"
}
}
}
m = Memory.from_config(config)
results = m.search("query", filters={"user_id": "alice"})
See It in Action
Basic Reranked Search
results = m.search(
"What are my food preferences?",
filters={"user_id": "alice"}
)
for result in results["results"]:
print(f"Memory: {result['memory']}")
print(f"Score: {result['score']}")
Expect each result to list the reranker-adjusted score so you can compare ordering against baseline vector results.
Toggle Reranking Per Request
results_with_rerank = m.search(
"What movies do I like?",
filters={"user_id": "alice"},
rerank=True
)
results_without_rerank = m.search(
"What movies do I like?",
filters={"user_id": "alice"},
rerank=False
)
Log the reranked vs. non-reranked lists during rollout so stakeholders can see the improvement before enforcing it everywhere.
Combine With Metadata Filters
results = m.search(
"important work tasks",
filters={
"AND": [\
{"user_id": "alice"},\
{"category": "work"},\
{"priority": {"gte": 7}}\
]
},
rerank=True,
top_k=20
)
Verify filtered reranked searches still respect every metadata clause: reranking only reorders candidates; it never bypasses filters.
Real-World Playbooks
Customer Support
config = {
"reranker": {
"provider": "cohere",
"config": {
"model": "rerank-v3.5",
"api_key": "your-cohere-api-key"
}
}
}
m = Memory.from_config(config)
results = m.search(
"customer having login issues with mobile app",
filters={"agent_id": "support_bot", "category": "technical_support"},
rerank=True
)
Top results should highlight tickets matching the login issue context so agents can respond faster.
Content Recommendation
results = m.search(
"science fiction books with space exploration themes",
filters={"user_id": "reader123", "content_type": "book_recommendation"},
rerank=True,
top_k=10
)
for result in results["results"]:
print(f"Recommendation: {result['memory']}")
print(f"Relevance: {result['score']:.3f}")
Expect high-scoring recommendations that match both the requested theme and any metadata limits you applied.
Verify The Feature Is Working
- Inspect result payloads for both
score(vector) and reranker scores; mismatched fields indicate the reranker didn’t execute. - Track latency before and after enabling reranking to ensure SLAs hold.
- Review provider logs or dashboards for throttling or quota warnings.
- Run A/B comparisons (rerank on/off) to validate improved relevance before defaulting to reranked responses.