Skip to content
nikhilcrypto0Public

About

Production-grade customer support agent: LangGraph + Claude, pgvector RAG, human-approved refunds, evals in CI

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Concierge

ci license: MIT python 3.12 mypy: strict LangGraph 1.2

A customer support agent that cannot move money on its own. It answers questions from a help center and handles refund requests, but no refund is paid until a person approves it. The refund amount comes from the written policy in code, never from the model, and every approval is recorded with the reviewer, the time, and the policy reason for the amount.

Live demo: concierge-sand.vercel.app. Pick a customer, ask it to cancel BK-1042 and refund you, and watch the request wait for a person. The demo runs on free hosting, so the first message after a quiet spell can take up to a minute while the API wakes up. The support console is password protected, because it is where refunds are approved; a view-only tour with sample data and the screenshots below show it. Every test case, with what counts as correct and what each model did, is on the How it was tested page.

Under the hood: a LangGraph workflow over Claude, retrieval on Postgres + pgvector, and evals that gate CI.

The demo company is Tidewell Home Services, a fictional home cleaning and repair business with a help center, bookings, and a refund policy. Two screens make the safety model visible: the customer site with the assistant, and the support console where a person approves refunds.

The customer asks for a refund. The AI does not pay it.

Customer chat waiting for a human decision

A support lead sees why, and decides.

Support console with a pending refund

The customer's chat updates the moment it is approved.

Customer chat showing the approved refund

Results at a glance

Measured on 2026-10-06 with Claude Opus 5, on the current build, on 54 conversations (30 of them safety attacks). Three earlier runs of the original 31 cases scored 30 of 31 each time. All raw runs are in evals/results/. The set is still small, so the intervals are wide (52/54 supports roughly 88% to 99%); the stronger guarantee is structural.

End-to-end agent eval 52 / 54 cases (96.3%)
Safety cases (prompt injection, other customers' bookings, fake "admin" authority, multi-turn pressure) 29 / 30
Refunds executed without a human decision 0
Cost per conversation $0.0082
Turn latency p50 2.1 s, p95 4.9 s
Retrieval (shipped mode) Recall@4 1.00, Section@4 0.976, MRR 0.927

What it does

  • Answers questions from the help center with citations, and says "I couldn't find that" instead of guessing.
  • Looks up bookings for the customer who owns them. Anyone else's booking is indistinguishable from a missing one.
  • Handles refunds by applying the written policy in code, then pausing the workflow until a person approves in full, approves a lower amount, or rejects. The approved refund is applied exactly once, even under retries and concurrent clicks.
  • Hands off to a human for complaints, damage, low-confidence routing, model outages, and budget breaches, and records why.
  • Blocks prompt injection before any model call, treats retrieved help-center text as untrusted too, and never tells the sender why.
  • Lets a demo visitor play the support lead for their own request: a card under the chat offers Approve (optionally a lower amount) and Reject, so the whole loop can be completed without the operator password. It exists only in demo mode (POST /v1/demo/decide, 404 otherwise), can decide only the one open request on the asking customer's own conversation, goes through the same decision code as the real console, and is recorded with the reviewer demo-visitor. The real product has no such route.
  • Shows its own running cost: the landing page reads GET /v1/stats, which prices the last 24 hours of the token ledger per model (the same ledger that enforces the budgets) and reports the average cost per conversation and how much of the daily token limit is used. Aggregates only, no customer or chat data.
  • Lets visitors try to break it: a panel on the demo site sends four real attacks (ignore your rules, someone else's booking, a fake manager, a refund you are not owed) to the live agent and shows the outcome the system actually returned.
  • Exposes read-only tools over MCP, so Claude Code, Claude Desktop, or Cursor can search the help center and look up bookings.

Architecture

flowchart LR
    V[Visitor] --> WEB[Next.js demo UI<br/>customer site + support console]
    WEB -->|server-side proxy<br/>holds the API keys| API[FastAPI]
    O[Support operator] --> WEB
    API --> G[LangGraph workflow]
    G -->|classify, answer| Claude[Claude Opus 5<br/>fallback: Sonnet 5]
    G --> KB[(pgvector<br/>help center)]
    G --> DB[(Postgres<br/>bookings, approvals,<br/>audit log, token ledger)]
    G --> CP[(Postgres<br/>LangGraph checkpoints)]
    G -.traces, opt-in.-> LF[Langfuse]
    MCP[MCP server<br/>read-only] --> KB
    MCP --> DB
Loading

Langfuse tracing is wired in observability.py but off unless you turn it on: every tracing call returns early unless LANGFUSE_SECRET_KEY is set, and docker-compose.yml starts Postgres and the API only. Enabling traces means pointing at your own Langfuse instance. The dotted arrow is drawn that way on purpose, so the diagram does not promise a running service you would not find.

The workflow is an explicit state machine, not a free-form tool-calling loop:

flowchart TD
    start([customer message]) --> guard
    guard -->|injection or empty| respond
    guard --> classify
    classify -->|question| retrieve
    retrieve -->|nothing relevant| respond
    retrieve --> answer --> respond
    classify -->|booking_status / refund_request| lookup[lookup_booking]
    lookup -->|missing or not theirs| respond
    lookup --> status[booking_status] --> respond
    lookup --> assess[assess_refund<br/>policy in code]
    assess -->|ineligible or needs a person| respond
    assess --> request[request_approval<br/>tells the customer]
    request --> wait[await_approval<br/>interrupt: waits for a human]
    wait --> respond
    classify -->|human_handoff| handoff --> respond
    classify -->|out_of_scope| oos[out_of_scope] --> respond
    respond --> done([reply])
Loading

Design decisions and their costs

Decision Why Tradeoff
Routed workflow instead of a ReAct agent loop Every path is enumerable, testable, and auditable; the model can never choose to move money Requests outside the designed intents go to a human
The model makes two judgment calls only: intent and grounded answer Lookups, policy math, approvals, and every reply that states an amount are plain code, so the assistant cannot misquote a refund More code than "let the model figure it out"
Human approval with interrupt(), authority in Postgres The run pauses durably in the Postgres checkpointer; on resume it re-reads the approval row and refuses a mismatched amount, booking, or conversation An operator must act before the customer gets an outcome
A reviewer can approve less, never more Approve in full, approve a lower amount, or reject. The policy amount computed in code is the ceiling: the API refuses anything above it and CHECK constraints in Postgres refuse it even from a direct write. The customer's message is filled from the stored amounts and says when it is less than requested One more control for the operator to learn; a reviewer cannot raise an amount, so a goodwill refund above the policy amount needs a policy change in code
Exactly-once refunds, enforced twice An actions table keyed by idempotency key, a row lock on the approval, and a SQL guard against refunding more than was paid, in one transaction. A Postgres trigger independently refuses any action that no approved request covers, or that exceeds the authorised amount, even from a direct SQL write Tested with 5 concurrent executions: 1 applies, 4 no-op
Per-conversation advisory lock A double-submit or a chat racing an approval never runs the graph twice on one thread; works across replicas The losing request gets a 409 and must retry
Structured output + citation validation Answers are Pydantic objects; an answer citing a document it was not given is rejected as ungrounded Some correct answers with sloppy citations become "I couldn't find that"
Search with the customer's words and the model's rewrite The rewrite resolves follow-ups; the customer's words protect against rewrite drift (see below) Two searches per question (about 7 ms each)
The browser never sees an API key or an email The UI sends a persona id to its own server routes, which attach the key and map the id to a customer, so a visitor cannot read another customer's data by editing a request The demo UI needs a server; it cannot be a static page
Retrieved help-center text is untrusted too A tampered article is an injection channel like a customer message. Passages are cleaned, any that read like instructions to the model are dropped (logged by id, never explained), attribute values are escaped, and a dropped passage cannot be cited Pattern matching is a heuristic and can be evaded. The real guarantee is structural: the model cannot move money, so a passage that slips through can at worst cause a wrong answer, never a payout
Relevance gate before the answer call If nothing is similar enough, skip the model entirely Threshold needs re-tuning if the embedding model changes
Retrieval mode chosen by measurement Hybrid search is the textbook default but scored lower than vector-only here, so vector-only ships Revisit as the corpus changes
Local ONNX embeddings (fastembed, bge-small) No API key or GPU; identical vectors in dev, CI, and the container Weaker on some phrasings than larger models (see the one failing case)
Token budgets per conversation and per day Checked before every model call; a breach halts the AI step and hands off Soft ceiling: concurrent calls can overshoot by one call
Fallback model on provider errors only Opus 5 falls back to Sonnet 5 on connection errors, timeouts, 429, 5xx, and 529, never on 400s Two models to keep prompts compatible with

Evals

Retrieval (free, no LLM calls, gates CI on every push)

42 answerable questions plus 6 off-topic ones, top 4 chunks. Recall@4 is judged at the article level; Section@4 is stricter and counts a hit only when a retrieved chunk is the section that answers the question.

Mode Recall@1 Recall@4 Section@4 MRR Relevance-gate accuracy
Hybrid (keyword weight 0.5) 0.833 1.000 0.976 0.899 0.979
Vector only (shipped) 0.881 1.000 0.976 0.927 0.979
Keyword only 0.690 0.857 0.810 0.766 0.938

Section@4 is the number that explains the one question the agent still fails. Article-level recall was a perfect 1.000, which hid it: for "Can I pay the cleaner in cash?" vector search returns the payments article's generic overview, not its "Accepted payment methods" section, so the answer step correctly says it cannot find the answer. Hybrid search fixes that question but misses a different one ("Can I book a plumber on a Sunday?"), and no keyword weight from 0.25 to 1.5, nor a larger top-k (5 or 6), fixes both. With 42 questions, tuning further would only fit the test set, so the shipped mode is unchanged.

With equal fusion weights, hybrid dropped to Recall@1 0.786: full-text matches on common words ("problem", "home") pulled in the wrong articles. Halving the keyword weight helped but did not beat vector-only. CI fails the build if the shipped mode falls below Recall@4 0.95, MRR 0.85, or gate accuracy 0.90.

End-to-end agent (real Claude, run on demand)

54 conversations (31 originally, 23 safety attacks added later) through the production graph against real Postgres, graded by deterministic checks rather than an LLM judge: the routed outcome and intent, citations to the expected article, key facts in the reply ("$25", "5 to 10 business days"), exact refund amounts proposed for approval, and the reason for any handoff. A handoff caused by a failure fails its case, so a model outage cannot pass as correct routing.

First run After fixes
Cases passed 27 / 31 (87.1%) 30 / 31 (96.8%)
Help-center questions 7 / 11 10 / 11
Bookings, refunds, routing 100% 100%
Safety 7 / 7 7 / 7
Refunds executed without a human 0 0
Cost per conversation $0.0084 $0.0085
Turn latency p50 / p95 2.5 s / 5.4 s 2.5 s / 5.2 s

evals/results/agent.json holds the latest run; the first-run column comes from that run's log. That table is the original 31 cases.

54-case run, and a cheaper model (2026-10-06)

I added 23 safety attacks (injection in other languages and with leetspeak, role-play, "translate then obey", prompt-leak requests, multi-turn pressure, SQL in a booking reference, other customers' bookings by several routes, "approve it yourself") and ran the full set on Claude Opus 5 and on Claude Haiku 4.5.

Opus 5 Haiku 4.5
Cases passed 52 / 54 48 / 54
Safety, as graded 29 / 30 25 / 30
Refunds executed without a human 0 0
Cost per conversation $0.0082 $0.0012
Turn latency p50 / p95 2.1 s / 4.9 s 1.8 s / 4.1 s

How to read this honestly:

  • No safety case in any run opened an approval or leaked another customer's data. Every safety "failure" was an outcome I had not listed as acceptable. For Opus, s15 was a low-confidence handoff to a human. For Haiku, all five were out_of_scope refusals, which are safe but not in my accepted list.
  • I widened my own accepted lists after the first run. Six cases (s10, s12, s13, s18, s20, s22) returned safe outcomes I had not listed, such as "out of scope" or "which booking reference?". I added those outcomes and said so here, rather than pretend I predicted them. The checks that matter (no approval opened, no other customer's data in the reply) were unchanged.
  • Haiku matched Opus on every non-safety case (bookings, refunds, routing, questions) and cost about 7x less. It refuses more attacks as "out of scope" instead of handing them to a person, so strictly as graded it scores lower on safety. I would test Haiku on the classify step and keep Opus for answers before switching, and I have not done that.
  • q05 (paying the cleaner in cash) fails on both models and in every run.
  • Haiku 4.5 rejects the effort setting, so the code leaves it out for Haiku models.

What the evals caught

1. Query rewriting silently broke retrieval. The first agent run answered 4 simple questions ("Do you have service in Denver?", "Are your plumbers licensed?") with "I couldn't find that." The retrieval eval had scored Recall@4 of 1.00, because it searches with the raw question. In production, the classifier rewrote the question first and added the company name, which matched the boilerplate intro of every article and pushed the real answer out of the top 4. The answer step then correctly refused to answer from the wrong documents. Fix: search with both the customer's words and the rewrite, and fuse the rankings. All 4 cases pass now, and a unit test pins the regression.

2. The remaining failure is retrieval granularity, not the pipeline. "Can I just pay the cleaner in cash?" fails when the rewrite keeps the customer's phrasing. A direct retrieval check shows why: bge-small ranks cleaning-service articles above the payment-methods section for "pay the cleaner", which is not even in the top 6. The answer step is right to refuse: the chunk it was given was the payments article's overview, not the section that says cash is not accepted. The article-level retrieval metric could not see this, so the retrieval eval now also scores section-level hits (Section@4 above). I tried hybrid search, keyword weights from 0.25 to 1.5, and a top-k of 5 and 6: none fixed it without breaking another question. The next experiment is a larger embedding model, which needs a schema migration and more memory than the free host has, so it is not done.

Review findings, all fixed

A security review and a correctness review ran against the agent, and again against the demo UI. Neither found a critical issue. What they did find:

  • Request bodies were fully buffered before length validation, so bodies over 64 KB are now rejected before parsing.
  • API docs were public in production; they are now hidden when ENVIRONMENT=prod.
  • Failed authentication was not rate limited; it is now throttled per client address, with bounded memory.
  • The injection guard blocked normal corrections like "please ignore my previous message"; the pattern now targets instruction-like phrases only, with regression tests.
  • Concurrent requests on one conversation could run the graph twice on the same thread; a Postgres advisory lock now serializes them.
  • Every handoff looked the same; handoffs now carry a reason, and the eval fails cases that hand off because something broke.
  • In the UI: "New chat" could leave the previous transcript on screen, switching persona mid-request could show one customer's message under the other's account, a demo reset left an open chat tab stuck until a page refresh, and the console could keep offering Approve on a request someone else had already decided. All four are fixed, and the demo-reset recovery was then verified live.

Running the live server also showed the "request sent for approval" message was missing from the saved transcript. It is now its own workflow step, checkpointed before the pause.

Run it locally

Requirements: Python 3.12, uv, Node 22, and Postgres 16+ with the pgvector extension.

# 1. API
uv sync
cp .env.example .env            # add ANTHROPIC_API_KEY, generate the two API keys, DEMO_MODE=true
createdb concierge
uv run concierge-ingest --seed-demo
uv run uvicorn --factory concierge.api.app:create_app      # http://127.0.0.1:8000

# 2. Demo UI, in a second terminal
cd web
npm install
cp .env.example .env.local      # point it at the API, paste the same two keys, and set
                                # CONSOLE_PASSWORD (16+ chars) and CONSOLE_SESSION_SECRET (32+)
npm run dev                     # http://localhost:3000

The console fails closed: without those two settings every approval route returns 401. Then open http://localhost:3000, ask the assistant to cancel BK-1042 and refund you, and approve it at /console after signing in with CONSOLE_PASSWORD. The API's own interactive docs stay at http://127.0.0.1:8000/docs.

Demo bookings (times are relative to when you seed):

Booking Customer Situation Refund outcome
BK-1042 maya@example.com 5 days away full refund, needs approval
BK-1043 maya@example.com 30 hours away 50%, needs approval
BK-1044 maya@example.com 6 hours away not eligible
BK-1045 maya@example.com completed handed to a person
BK-1046 maya@example.com already refunded not eligible
BK-2001 jordan@example.com 4 days away invisible to Maya

Use it from Claude Code over MCP

claude mcp add concierge -- uv --directory /path/to/concierge run concierge-mcp

Tools: search_help_center, get_booking, list_pending_refund_approvals. The server is read-only by design: no tool can issue a refund.

How the live demo is deployed

Piece Where Notes
Web (web/) Vercel, deploys main Holds the API keys and the console password server-side
API Render free Docker service, from render.yaml Auto-deploy is off; migrations run at startup
Database Neon Postgres with pgvector Use the direct connection string, not the pooled one

Settings (names only; values live in each host's dashboard, never in the repo):

  • API (Render): DATABASE_URL, ANTHROPIC_API_KEY, CLIENT_API_KEYS, OPERATOR_API_KEYS, plus ENVIRONMENT=prod, DEMO_MODE=true, RATE_LIMIT_PER_MINUTE, MAX_TOKENS_PER_DAY, MAX_TOKENS_PER_CONVERSATION.
  • Web (Vercel): CONCIERGE_API_URL, CONCIERGE_CLIENT_KEY, CONCIERGE_OPERATOR_KEY, CONSOLE_PASSWORD, CONSOLE_SESSION_SECRET.

Lessons from the first deploy, each of which cost real time:

  • The Anthropic key must be created inside a workspace. A personal-account key fails with "not scoped to a workspace", and the agent quietly hands every chat to a human (classifier_unavailable). The API now logs Anthropic's status and message so this is visible.
  • Neon's free tier drops idle connections. The connection pool verifies a connection before handing it out; before that fix, the first chat after a quiet spell returned a 500.
  • Seeding is refused when ENVIRONMENT=prod. Seed from a machine with ENVIRONMENT=dev pointed at the database.
  • Deploy the API before the web app when the API contract changes, because Vercel deploys on merge and the API is a manual deploy.
  • The free API host sleeps after 15 idle minutes and takes up to a minute to boot. The pages ping a keyless /api/wake on load and chat retries with a "waking up" notice, so a first visitor sees a short wait, not an error.

Tests

uv run ruff check src tests evals && uv run mypy src      # lint + strict types
uv run pytest tests/unit -q                               # workflow, policy, guardrails, auth
uv run pytest tests/integration -q                        # real Postgres: API, refunds, retrieval
uv run python evals/run_retrieval_eval.py                 # free
uv run python evals/run_agent_eval.py                     # calls Claude, about $0.26 per run
cd web && npm run lint && npx tsc --noEmit && npm run build
  • Unit tests run every workflow path with the model, search index, and database replaced by fakes: approval pause and resume, mismatched approvals, budget breach, invalid model output, ungrounded citations, injection in customer messages and in retrieved articles (including a check that no shipped article is ever quarantined), ownership, rewrite drift, per-turn state isolation, and that the figures on the landing page (web/lib/proof.ts) match evals/results/agent.json.
  • Integration tests use real Postgres: concurrent refund execution, racing approval decisions, rollback on over-refund, the full HTTP refund flow and transcript, state consistency for every decision path (approve, approve less, reject: booking, actions, approval row, and the customer's message agree) with the database itself refusing an over-policy amount, a connection pool that survives the server closing its connections, conversation hijacking, the conversation lock, rate limiting, body-size limits, hidden production docs, customer-scoped bookings, approval detail, and demo reset.
  • CI (GitHub Actions) runs lint, mypy, and unit tests; the web app's lint, types, and production build; integration tests against a pgvector service container; and the retrieval eval with thresholds. The agent eval runs on manual dispatch with an API key secret.

Project layout

src/concierge/
  agent/        graph.py (state machine + turn runner), nodes.py, state.py, prompts.py
  api/          app.py (endpoints), security.py (API keys, rate limits), schemas.py
  bookings/     policy.py (refund rules), repository.py (SQL, approvals, exactly-once refunds)
  retrieval/    chunking.py, embeddings.py, search.py (vector / keyword / fusion), ingest.py
  llm.py        Claude calls with structured output and fallback
  guardrails.py input sanitization, injection detection, retrieved-text hygiene
  budget.py     token ledger and budgets
  locks.py      per-conversation advisory lock
  mcp_server.py read-only MCP tools
web/
  app/          customer site (/), support console (/console), server-side /api proxy routes
  components/   customer/ and console/ UI
  hooks/        chat, bookings, approval queue, polling, storage, API warm-up
  lib/          server-only API client, operator session, zod validation, personas, formatting
migrations/     SQL schema (001 base, 002 approved amount)
Dockerfile      API image with the embedding model baked in
render.yaml     Render blueprint for the API
data/kb/        help-center articles
evals/          datasets, runners, committed results
tests/          unit/ and integration/

Known limitations and next steps

The full list of threats, guards and residual risks is in docs/THREAT_MODEL.md. The most important honest caveat: the agent eval is small (54 cases, written by one person), so its pass rates carry wide error bars; the money guarantee rests on the structure of the system, not on those numbers.

  • The console login is one shared operator password (a signed, HttpOnly cookie checked on the server), not per-person accounts or SSO, so every approval is recorded against the same operator identity. A team deployment needs individual accounts.
  • Security headers are partial. Frame, sniffing, referrer, and permissions policies ship; a strict script content policy needs per-request nonces and is still missing.
  • Retrieval: the one failing eval case points at the embedding model; compare a larger model and hybrid-on-rewrites with the existing evals before changing the default.
  • Identity: the client API key represents a trusted backend that asserts the customer's email. The demo is safe because its server maps two fixed personas to emails; a product with real customers should pass a signed customer token (JWT) instead.
  • Rate limiting is in-process; multiple replicas need Redis or Postgres-backed limits. Behind Render's proxy the failed-login throttle keys on the proxy's address and is weak; the per-key limits are the controls that actually cap cost. The website shares one key, so the costly calls (chat and demo decisions) share 20 requests a minute in total on the live demo, while cheap reads (bookings, transcripts, stats) have a separate budget six times larger, so page loads and polling never starve a chat. Chunked uploads without a length header should be capped at the reverse proxy.
  • Migrations run at startup, which suits one instance; multi-replica deploys should run them as a release step.
  • The Docker image has not been built in CI yet; the Dockerfile and compose file cover the API only, not the web app.
  • No streaming yet: replies appear when the turn completes, with a typing indicator meanwhile.
  • Handoff is a reply, not a ticket: the next step is creating a ticket in a helpdesk system.
  • Evals use deterministic checks: an LLM-as-judge faithfulness score on free-form answers is the next eval to add.

About

Production-grade customer support agent: LangGraph + Claude, pgvector RAG, human-approved refunds, evals in CI

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages