Build an AI Agent for Customer Service That Remembers Every Customer
Build an AI Agent for Customer Service That Remembers Every Customer
Quick Takeaways:
73% of customers expect to move between channels without having to repeat themselves. 53% of consumers say they always have to repeat their issue when transferred between agents.
The fix is one shared
user_idin Mem0 that all channel endpoints write to and read from. Channel identity is a routing concern. Memory identity is unified.Mem0's scoping model (
user_id,agent_id,run_id) gives you precise control: cross-channel facts live atuser_idscope, session state lives atrun_idscope, and per-channel behavior tuning lives atagent_idscope.The demo in this post is a FastAPI backend with three endpoints (
POST /call,POST /email,POST /chat) and a Streamlit UI, all reading and writing the same memory store, with the customer's email as the shared key.Memory extracted from a phone transcript is available in the next chat session automatically.
Here is the scenario: A customer calls your support line, explains their billing issue for four minutes, and resolves the call. Three days later, they email about a receipt. The agent asks them to describe the issue again. That is what happens when your channels don't share memory.
73% of customers expect to start on one channel and finish on another without repeating themselves( survey). 53% of consumers say they always have to repeat their issue when transferred between agents. Only 13% of companies report that customer data, history, and context carry over fully across interactions and channels ( Deloitte's 2024 GCS survey). The standard response is to invest in a unified CRM dashboard, give agents a better inbox, and call it omnichannel. This kind of system does nothing for the AI agent generating the response in real time.
In this blog, we’ll understand that whether you're building an AI customer service bot or a full agent pipeline, you'll end up with a FastAPI backend, three or more channel endpoints, and a Streamlit UI that all read and write the same Mem0 memory store.
Here is a glimpse of what we’ll build:
💡 You can review the full code on the GitHub repository.
Why siloed channels produce bad AI customer support agents
Most support architectures are multichannel, not omnichannel customer support. They're available on multiple platforms (email, chat, phone), but each channel runs its own stack. Its own queue, its own agent prompt, and its own history store. When a customer switches, they're starting over.
The difference between multichannel and omnichannel is whether context moves with the customer. In 2026, the main failure point is no longer channel access. It's channel switching. Customers can usually find a way in. The problem is what happens when they move from a phone call to an email to a chat. If the AI agent only sees the current channel and not the customer's history across all channels, it asks redundant questions, misses prior resolutions, and starts fresh every time.
This is a memory architecture problem. A bigger context window doesn't solve the cross-session problem: even 1M tokens only helps if the right history is in the window. When a customer switches channels, there's no guarantee any prior history is present at all.
Four specific things break when channels are siloed:
Re-asking: The agent asks for information the customer already gave on another channel. This is the 53% problem.
Context mismatch: The email agent refers to a billing status the phone agent already resolved.
Tone drift: The phone call was urgent and empathetic. The chatbot reply is formal and generic. The customer feels like they're talking to a different company.
Duplicate escalations: The chat agent escalates an issue the phone agent already escalated. Engineering gets two tickets about the same problem.
The fix is conceptually simple: Give every channel endpoint the same identity key and the same memory store. When the phone call ends, extract the facts. When the email comes in, extract more. When the chat opens, retrieve everything relevant. One user_id, all channels, shared retrieval.
How does Mem0 scope memory across channels?
Mem0's scoping model maps onto a multi-channel support use case. There are four identity dimensions available on every add() and search() call.
user_idis the cross-channel with cross-session scope. Facts stored here persist across every channel and every session for the user. This is where cross-channel support memory lives: "user is on the annual plan," "payment failed on May 14," "needs receipt for expense report." Use the customer's email as theuser_idso it resolves naturally across channels regardless of which one the customer used to contact you.agent_idis the per-channel-agent scope. The phone agent, email agent, and chat agent each get their ownagent_id. We pass it for agent-level provenance, but note that combininguser_idandagent_idfilters in a singlesearch()call can be tricky. For the practical channel audit trail,metadata={"channel": channel}is the more reliable tag.run_idis the session scope. It is useful for scoping a single conversation in-progress state, or for cleaning up a session's temporary context after it closes. Pass arun_idper session if you need to isolate or expire session-level facts.app_idis the application-level scope. It is used for shared organizational context like product catalog, policy documents, and known bugs. Not used in this demo but useful in production when you want all agents to share institutional knowledge.
For a multi-channel support agent, the practical mapping looks like this:
| Scope | What you store | Example |
|---|---|---|
user_id |
Cross-channel facts | "Annual plan, card 4821, $199 charge, May 14" |
agent_id |
Agent-level provenance | Phone vs email vs chat agent. Use metadata.channel for audit queries |
run_id |
In-session ephemeral state | Current chat session, this call's draft response |
app_id |
Org-wide shared knowledge | Refund policy, known billing bugs, escalation paths |
One important constraint to know before writing any queries: Mem0's search() requires entity parameters inside a filters={} dict, not as top-level keyword arguments. For cross-channel retrieval, always query at user_id scope.
Correct: entity params inside filters={}
result = mem0.search(query,filters={"user_id": user_id},top_k=6)
Wrong: will raise ValueError at runtime
result = mem0.search(query,user_id=user_id,top_k=6) # ValidationError
💡 The multi-agent memory systems post documents the scoping model in full detail. Worth reading before you build anything beyond a single-agent setup.
The Architecture
The architecture has two layers. A FastAPI backend exposes three endpoints. A Streamlit UI runs the demo with a tab per channel. All three endpoints share the same user_id (the customer's email) and the same Mem0 memory store.
Call transcript → POST /call ─┐
Email body → POST /email ─┤─→ Mem0 (user_id = customer email)
Chat message → POST /chat ─┘ ↑
retrieves on /chat
The customer's email is the shared memory key. The phone agent extracts it from the call transcript. The email agent reads it from the From: header. The chat agent asks for it on first contact, then uses it for all subsequent memory reads and writes.
Project structure
Cross-Channel-Support-Memory-with-Mem0/
├── multichannel_support.py # FastAPI backend + shared memory helpers
├── streamlit_app.py # Streamlit UI with Call / Email / Chat tabs
├── test_walkthrough.py # CLI walkthrough for direct API testing
├── requirements.txt
└── .env
Setup
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
Fill in .env:
First, go to app.mem0.ai , sign up for free, and copy your API key from the dashboard.
MEM0_API_KEY=your_mem0_key # from app.mem0.ai
# Azure OpenAI (used in this demo)
OPENAI_PROVIDER=azure
AZURE_OPENAI_API_KEY=your_azure_key
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com
AZURE_OPENAI_DEPLOYMENT=your-gpt-4o-deployment
AZURE_OPENAI_API_VERSION=2024-10-21
# Or standard OpenAI
# OPENAI_PROVIDER=openai
# OPENAI_API_KEY=your_openai_key
# OPENAI_MODEL=gpt-4o
Run the Streamlit demo:
streamlit run streamlit_app.py
Inside the FastAPI backend
The backend has three endpoints and a set of shared helpers. All memory reads and writes flow through two functions: store_memories() and retrieve_memories().
The shared memory helpers
def store_memories(
messages: list[dict],
user_id: str,
agent_id: str,
run_id: Optional[str],
channel: str,
) -> int:
"""
Write messages to Mem0 under the shared user_id.
metadata channel tag lets you filter by channel for analytics
without splitting the memory store.
"""
kwargs = dict(
user_id=user_id,
agent_id=agent_id,
metadata={"channel": channel},
)
if run_id:
kwargs["run_id"] = run_id
result = mem0.add(messages, **kwargs)
return len(result.get("results", []))
def retrieve_memories(user_id: str, query: str, limit: int = 6) -> list[str]:
"""
Retrieve cross-channel memories scoped to this user.
Entity params go inside filters={} - not as top-level kwargs.
"""
result = mem0.search(
query,
filters={"user_id": user_id},
top_k=limit,
)
return [r["memory"] for r in result.get("results", []) if "memory" in r]
Note: This is the retrieval step. In production, this runs before every agent response across every channel. Start free on Mem0 and store your first memories for free.
The metadata={"channel": channel} tag on every write is worth noting. It doesn't affect retrieval (search() ignores metadata by default), but it lets you audit which channel contributed which facts in the Mem0 dashboard.
One note on the V3 API: add() can return an async queued response with status and event_id rather than a results list. The demo uses result.get("results", []) as a safe fallback, but don't rely on a memory count for production logic.
POST /call - Ingest a voice transcript
The handle_call function receives a voice call transcript (as plain text) and stores extracted facts under the user's cross-channel memory. The email is extracted from the transcript using a regex, then used as user_id.
@app.post("/call", response_model=MemoryResponse)
def handle_call(req: CallRequest):
if not req.transcript.strip():
raise HTTPException(status_code=400, detail="transcript cannot be empty")
customer_email = require_customer_email(text=req.transcript)
messages = [{"role": "user", "content": req.transcript}]
store_memories(
messages=messages,
user_id=customer_email,
agent_id=PHONE_AGENT_ID,
run_id=req.run_id,
channel="phone",
)
return MemoryResponse(
status="ok",
message=f"Call transcript processed and saved to Mem0 for {customer_email}.",
)
POST /email - Ingest an inbound support email
This step receives an inbound support email where the subject and body are combined, so Mem0 extracts intent from both. The customer_email is used as user_id for cross-channel memory.
@app.post("/email", response_model=MemoryResponse)
def handle_email(req: EmailRequest):
combined = f"Subject: {req.subject}\n\n{req.body}"
messages = [{"role": "user", "content": combined}]
stored = store_memories(
messages=messages,
user_id=req.customer_email.lower(), # email as the shared key
agent_id=EMAIL_AGENT_ID,
run_id=req.run_id,
channel="email",
)
return MemoryResponse(
status="ok",
memories_stored=stored,
message=f"Email processed. {stored} memories extracted.",
)
POST /chat - Memory-augmented real-time chat
The chat endpoint is where the cross-channel payoff is visible. It retrieves all prior memory across channels before generating a response, then stores the new turn for future channels to see.
@app.post("/chat", response_model=ChatResponse)
def handle_chat(req: ChatRequest):
found_email = extract_email(req.message)
user_id = found_email or req.user_id
if not user_id:
return ChatResponse(
reply="Hi, please share your account email to get started.",
context_used=[],
)
memories = retrieve_memories(user_id=user_id, query=req.message)
system_prompt = build_support_prompt(channel="chat", memories=memories)
reply = generate_support_reply(
system_prompt=system_prompt,
user_message=req.message,
)
store_memories(
messages=[
{"role": "user", "content": req.message},
{"role": "assistant", "content": reply},
],
user_id=user_id,
agent_id=CHAT_AGENT_ID,
run_id=req.run_id,
channel="chat",
)
return ChatResponse(reply=reply, context_used=memories)
Conclusion
The reason most support agents have channel amnesia is not that memory is hard. It's that most architectures treat channels as independent systems with independent history stores. The fix is architectural i.e, one user_id, shared across every channel, reading and writing to the same Mem0 store.
The FastAPI backend in this post has three endpoints in a compact single-file backend (multichannel_support.py). The Streamlit UI makes the before/after visible in a way you can demo in 60 seconds. If you want to move beyond the demo layer, the Next.js + Mem0 customer support agent walkthrough covers building a production-ready frontend on top of the same memory architecture.