Build an AI Agent for Customer Service That Remembers Every Customer

Build an AI Agent for Customer Service That Remembers Every Customer

Quick Takeaways:

Here is the scenario: A customer calls your support line, explains their billing issue for four minutes, and resolves the call. Three days later, they email about a receipt. The agent asks them to describe the issue again. That is what happens when your channels don't share memory.

73% of customers expect to start on one channel and finish on another without repeating themselves( survey). 53% of consumers say they always have to repeat their issue when transferred between agents. Only 13% of companies report that customer data, history, and context carry over fully across interactions and channels ( Deloitte's 2024 GCS survey). The standard response is to invest in a unified CRM dashboard, give agents a better inbox, and call it omnichannel. This kind of system does nothing for the AI agent generating the response in real time.

In this blog, we’ll understand that whether you're building an AI customer service bot or a full agent pipeline, you'll end up with a FastAPI backend, three or more channel endpoints, and a Streamlit UI that all read and write the same Mem0 memory store.

Here is a glimpse of what we’ll build:

💡 You can review the full code on the GitHub repository.

Why siloed channels produce bad AI customer support agents

Most support architectures are multichannel, not omnichannel customer support. They're available on multiple platforms (email, chat, phone), but each channel runs its own stack. Its own queue, its own agent prompt, and its own history store. When a customer switches, they're starting over.

The difference between multichannel and omnichannel is whether context moves with the customer. In 2026, the main failure point is no longer channel access. It's channel switching. Customers can usually find a way in. The problem is what happens when they move from a phone call to an email to a chat. If the AI agent only sees the current channel and not the customer's history across all channels, it asks redundant questions, misses prior resolutions, and starts fresh every time.

This is a memory architecture problem. A bigger context window doesn't solve the cross-session problem: even 1M tokens only helps if the right history is in the window. When a customer switches channels, there's no guarantee any prior history is present at all.

Four specific things break when channels are siloed:

  1. Re-asking: The agent asks for information the customer already gave on another channel. This is the 53% problem.

  2. Context mismatch: The email agent refers to a billing status the phone agent already resolved.

  3. Tone drift: The phone call was urgent and empathetic. The chatbot reply is formal and generic. The customer feels like they're talking to a different company.

  4. Duplicate escalations: The chat agent escalates an issue the phone agent already escalated. Engineering gets two tickets about the same problem.

The fix is conceptually simple: Give every channel endpoint the same identity key and the same memory store. When the phone call ends, extract the facts. When the email comes in, extract more. When the chat opens, retrieve everything relevant. One user_id, all channels, shared retrieval.

How does Mem0 scope memory across channels?

Mem0's scoping model maps onto a multi-channel support use case. There are four identity dimensions available on every add() and search() call.

For a multi-channel support agent, the practical mapping looks like this:

Scope What you store Example
user_id Cross-channel facts "Annual plan, card 4821, $199 charge, May 14"
agent_id Agent-level provenance Phone vs email vs chat agent. Use metadata.channel for audit queries
run_id In-session ephemeral state Current chat session, this call's draft response
app_id Org-wide shared knowledge Refund policy, known billing bugs, escalation paths

One important constraint to know before writing any queries: Mem0's search() requires entity parameters inside a filters={} dict, not as top-level keyword arguments. For cross-channel retrieval, always query at user_id scope.

Correct: entity params inside filters={}

result = mem0.search(query,filters={"user_id": user_id},top_k=6)

Wrong: will raise ValueError at runtime

result = mem0.search(query,user_id=user_id,top_k=6)  # ValidationError

💡 The multi-agent memory systems post documents the scoping model in full detail. Worth reading before you build anything beyond a single-agent setup.

The Architecture

The architecture has two layers. A FastAPI backend exposes three endpoints. A Streamlit UI runs the demo with a tab per channel. All three endpoints share the same user_id (the customer's email) and the same Mem0 memory store.

Call transcript → POST /call ─┐

Email body → POST /email ─┤─→ Mem0 (user_id = customer email)

Chat message → POST /chat ─┘ ↑

retrieves on /chat

The customer's email is the shared memory key. The phone agent extracts it from the call transcript. The email agent reads it from the From: header. The chat agent asks for it on first contact, then uses it for all subsequent memory reads and writes.

Project structure

Cross-Channel-Support-Memory-with-Mem0/

├── multichannel_support.py # FastAPI backend + shared memory helpers

├── streamlit_app.py # Streamlit UI with Call / Email / Chat tabs

├── test_walkthrough.py # CLI walkthrough for direct API testing

├── requirements.txt

└── .env

Setup

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Fill in .env:

First, go to app.mem0.ai , sign up for free, and copy your API key from the dashboard.

MEM0_API_KEY=your_mem0_key # from app.mem0.ai
# Azure OpenAI (used in this demo)
OPENAI_PROVIDER=azure
AZURE_OPENAI_API_KEY=your_azure_key
AZURE_OPENAI_ENDPOINT=https://your-resource.openai.azure.com
AZURE_OPENAI_DEPLOYMENT=your-gpt-4o-deployment
AZURE_OPENAI_API_VERSION=2024-10-21
# Or standard OpenAI
# OPENAI_PROVIDER=openai
# OPENAI_API_KEY=your_openai_key
# OPENAI_MODEL=gpt-4o

Run the Streamlit demo:

streamlit run streamlit_app.py

Inside the FastAPI backend

The backend has three endpoints and a set of shared helpers. All memory reads and writes flow through two functions: store_memories() and retrieve_memories().

The shared memory helpers

def store_memories(
    messages: list[dict],
    user_id: str,
    agent_id: str,
    run_id: Optional[str],
    channel: str,
) -> int:
    """
    Write messages to Mem0 under the shared user_id.
    metadata channel tag lets you filter by channel for analytics
    without splitting the memory store.
    """
    kwargs = dict(
        user_id=user_id,
        agent_id=agent_id,
        metadata={"channel": channel},
    )
    if run_id:
        kwargs["run_id"] = run_id
    result = mem0.add(messages, **kwargs)
    return len(result.get("results", []))

def retrieve_memories(user_id: str, query: str, limit: int = 6) -> list[str]:
    """
    Retrieve cross-channel memories scoped to this user.
    Entity params go inside filters={} - not as top-level kwargs.
    """
    result = mem0.search(
        query,
        filters={"user_id": user_id},
        top_k=limit,
    )
    return [r["memory"] for r in result.get("results", []) if "memory" in r]

Note: This is the retrieval step. In production, this runs before every agent response across every channel. Start free on Mem0 and store your first memories for free.

The metadata={"channel": channel} tag on every write is worth noting. It doesn't affect retrieval (search() ignores metadata by default), but it lets you audit which channel contributed which facts in the Mem0 dashboard.

One note on the V3 API: add() can return an async queued response with status and event_id rather than a results list. The demo uses result.get("results", []) as a safe fallback, but don't rely on a memory count for production logic.

POST /call - Ingest a voice transcript

The handle_call function receives a voice call transcript (as plain text) and stores extracted facts under the user's cross-channel memory. The email is extracted from the transcript using a regex, then used as user_id.

@app.post("/call", response_model=MemoryResponse)

def handle_call(req: CallRequest):
    if not req.transcript.strip():
        raise HTTPException(status_code=400, detail="transcript cannot be empty")
    customer_email = require_customer_email(text=req.transcript)
    messages = [{"role": "user", "content": req.transcript}]
    store_memories(
        messages=messages,
        user_id=customer_email,
        agent_id=PHONE_AGENT_ID,
        run_id=req.run_id,
        channel="phone",
    )
    return MemoryResponse(
        status="ok",
        message=f"Call transcript processed and saved to Mem0 for {customer_email}.",
    )

POST /email - Ingest an inbound support email

This step receives an inbound support email where the subject and body are combined, so Mem0 extracts intent from both. The customer_email is used as user_id for cross-channel memory.

@app.post("/email", response_model=MemoryResponse)

def handle_email(req: EmailRequest):
    combined = f"Subject: {req.subject}\n\n{req.body}"
    messages = [{"role": "user", "content": combined}]
    stored = store_memories(
        messages=messages,
        user_id=req.customer_email.lower(), # email as the shared key
        agent_id=EMAIL_AGENT_ID,
        run_id=req.run_id,
        channel="email",
    )
    return MemoryResponse(
        status="ok",
        memories_stored=stored,
        message=f"Email processed. {stored} memories extracted.",
    )

POST /chat - Memory-augmented real-time chat

The chat endpoint is where the cross-channel payoff is visible. It retrieves all prior memory across channels before generating a response, then stores the new turn for future channels to see.

@app.post("/chat", response_model=ChatResponse)

def handle_chat(req: ChatRequest):
    found_email = extract_email(req.message)
    user_id = found_email or req.user_id
    if not user_id:
        return ChatResponse(
            reply="Hi, please share your account email to get started.",
            context_used=[],
        )
    memories = retrieve_memories(user_id=user_id, query=req.message)
    system_prompt = build_support_prompt(channel="chat", memories=memories)
    reply = generate_support_reply(
        system_prompt=system_prompt,
        user_message=req.message,
    )
    store_memories(
        messages=[
            {"role": "user", "content": req.message},
            {"role": "assistant", "content": reply},
        ],
        user_id=user_id,
        agent_id=CHAT_AGENT_ID,
        run_id=req.run_id,
        channel="chat",
    )
    return ChatResponse(reply=reply, context_used=memories)

Conclusion

The reason most support agents have channel amnesia is not that memory is hard. It's that most architectures treat channels as independent systems with independent history stores. The fix is architectural i.e, one user_id, shared across every channel, reading and writing to the same Mem0 store.

The FastAPI backend in this post has three endpoints in a compact single-file backend (multichannel_support.py). The Streamlit UI makes the before/after visible in a way you can demo in 60 seconds. If you want to move beyond the demo layer, the Next.js + Mem0 customer support agent walkthrough covers building a production-ready frontend on top of the same memory architecture.