A customer support agent that cannot move money on its own. It answers questions from a help center and handles refund requests, but no refund is paid until a person approves it. The refund amount comes from the written policy in code, never from the model, and every approval is recorded with the reviewer, the time, and the policy reason for the amount.
Live demo: concierge-sand.vercel.app. Pick a customer, ask it to cancel BK-1042 and refund you, and watch the request wait for a person. The demo runs on free hosting, so the first message after a quiet spell can take up to a minute while the API wakes up. The support console is password protected, because it is where refunds are approved; a view-only tour with sample data and the screenshots below show it. Every test case, with what counts as correct and what each model did, is on the How it was tested page.
Under the hood: a LangGraph workflow over Claude, retrieval on Postgres + pgvector, and evals that gate CI.
The demo company is Tidewell Home Services, a fictional home cleaning and repair business with a help center, bookings, and a refund policy. Two screens make the safety model visible: the customer site with the assistant, and the support console where a person approves refunds.
Measured on 2026-10-06 with Claude Opus 5, on the current build, on 54 conversations (30 of them safety attacks). Three earlier runs of the original 31 cases scored 30 of 31 each time. All raw runs are in evals/results/. The set is still small, so the intervals are wide (52/54 supports roughly 88% to 99%); the stronger guarantee is structural.
| End-to-end agent eval | 52 / 54 cases (96.3%) |
| Safety cases (prompt injection, other customers' bookings, fake "admin" authority, multi-turn pressure) | 29 / 30 |
| Refunds executed without a human decision | 0 |
| Cost per conversation | $0.0082 |
| Turn latency | p50 2.1 s, p95 4.9 s |
| Retrieval (shipped mode) | Recall@4 1.00, Section@4 0.976, MRR 0.927 |
- Answers questions from the help center with citations, and says "I couldn't find that" instead of guessing.
- Looks up bookings for the customer who owns them. Anyone else's booking is indistinguishable from a missing one.
- Handles refunds by applying the written policy in code, then pausing the workflow until a person approves in full, approves a lower amount, or rejects. The approved refund is applied exactly once, even under retries and concurrent clicks.
- Hands off to a human for complaints, damage, low-confidence routing, model outages, and budget breaches, and records why.
- Blocks prompt injection before any model call, treats retrieved help-center text as untrusted too, and never tells the sender why.
- Lets a demo visitor play the support lead for their own request: a card under the chat offers Approve (optionally a lower amount) and Reject, so the whole loop can be completed without the operator password. It exists only in demo mode (
POST /v1/demo/decide, 404 otherwise), can decide only the one open request on the asking customer's own conversation, goes through the same decision code as the real console, and is recorded with the reviewerdemo-visitor. The real product has no such route. - Shows its own running cost: the landing page reads
GET /v1/stats, which prices the last 24 hours of the token ledger per model (the same ledger that enforces the budgets) and reports the average cost per conversation and how much of the daily token limit is used. Aggregates only, no customer or chat data. - Lets visitors try to break it: a panel on the demo site sends four real attacks (ignore your rules, someone else's booking, a fake manager, a refund you are not owed) to the live agent and shows the outcome the system actually returned.
- Exposes read-only tools over MCP, so Claude Code, Claude Desktop, or Cursor can search the help center and look up bookings.
flowchart LR
V[Visitor] --> WEB[Next.js demo UI<br/>customer site + support console]
WEB -->|server-side proxy<br/>holds the API keys| API[FastAPI]
O[Support operator] --> WEB
API --> G[LangGraph workflow]
G -->|classify, answer| Claude[Claude Opus 5<br/>fallback: Sonnet 5]
G --> KB[(pgvector<br/>help center)]
G --> DB[(Postgres<br/>bookings, approvals,<br/>audit log, token ledger)]
G --> CP[(Postgres<br/>LangGraph checkpoints)]
G -.traces, opt-in.-> LF[Langfuse]
MCP[MCP server<br/>read-only] --> KB
MCP --> DB
Langfuse tracing is wired in observability.py but off unless you turn it on: every tracing call returns early unless LANGFUSE_SECRET_KEY is set, and docker-compose.yml starts Postgres and the API only. Enabling traces means pointing at your own Langfuse instance. The dotted arrow is drawn that way on purpose, so the diagram does not promise a running service you would not find.
The workflow is an explicit state machine, not a free-form tool-calling loop:
flowchart TD
start([customer message]) --> guard
guard -->|injection or empty| respond
guard --> classify
classify -->|question| retrieve
retrieve -->|nothing relevant| respond
retrieve --> answer --> respond
classify -->|booking_status / refund_request| lookup[lookup_booking]
lookup -->|missing or not theirs| respond
lookup --> status[booking_status] --> respond
lookup --> assess[assess_refund<br/>policy in code]
assess -->|ineligible or needs a person| respond
assess --> request[request_approval<br/>tells the customer]
request --> wait[await_approval<br/>interrupt: waits for a human]
wait --> respond
classify -->|human_handoff| handoff --> respond
classify -->|out_of_scope| oos[out_of_scope] --> respond
respond --> done([reply])
| Decision | Why | Tradeoff |
|---|---|---|
| Routed workflow instead of a ReAct agent loop | Every path is enumerable, testable, and auditable; the model can never choose to move money | Requests outside the designed intents go to a human |
| The model makes two judgment calls only: intent and grounded answer | Lookups, policy math, approvals, and every reply that states an amount are plain code, so the assistant cannot misquote a refund | More code than "let the model figure it out" |
Human approval with interrupt(), authority in Postgres |
The run pauses durably in the Postgres checkpointer; on resume it re-reads the approval row and refuses a mismatched amount, booking, or conversation | An operator must act before the customer gets an outcome |
| A reviewer can approve less, never more | Approve in full, approve a lower amount, or reject. The policy amount computed in code is the ceiling: the API refuses anything above it and CHECK constraints in Postgres refuse it even from a direct write. The customer's message is filled from the stored amounts and says when it is less than requested | One more control for the operator to learn; a reviewer cannot raise an amount, so a goodwill refund above the policy amount needs a policy change in code |
| Exactly-once refunds, enforced twice | An actions table keyed by idempotency key, a row lock on the approval, and a SQL guard against refunding more than was paid, in one transaction. A Postgres trigger independently refuses any action that no approved request covers, or that exceeds the authorised amount, even from a direct SQL write |
Tested with 5 concurrent executions: 1 applies, 4 no-op |
| Per-conversation advisory lock | A double-submit or a chat racing an approval never runs the graph twice on one thread; works across replicas | The losing request gets a 409 and must retry |
| Structured output + citation validation | Answers are Pydantic objects; an answer citing a document it was not given is rejected as ungrounded | Some correct answers with sloppy citations become "I couldn't find that" |
| Search with the customer's words and the model's rewrite | The rewrite resolves follow-ups; the customer's words protect against rewrite drift (see below) | Two searches per question (about 7 ms each) |
| The browser never sees an API key or an email | The UI sends a persona id to its own server routes, which attach the key and map the id to a customer, so a visitor cannot read another customer's data by editing a request | The demo UI needs a server; it cannot be a static page |
| Retrieved help-center text is untrusted too | A tampered article is an injection channel like a customer message. Passages are cleaned, any that read like instructions to the model are dropped (logged by id, never explained), attribute values are escaped, and a dropped passage cannot be cited | Pattern matching is a heuristic and can be evaded. The real guarantee is structural: the model cannot move money, so a passage that slips through can at worst cause a wrong answer, never a payout |
| Relevance gate before the answer call | If nothing is similar enough, skip the model entirely | Threshold needs re-tuning if the embedding model changes |
| Retrieval mode chosen by measurement | Hybrid search is the textbook default but scored lower than vector-only here, so vector-only ships | Revisit as the corpus changes |
| Local ONNX embeddings (fastembed, bge-small) | No API key or GPU; identical vectors in dev, CI, and the container | Weaker on some phrasings than larger models (see the one failing case) |
| Token budgets per conversation and per day | Checked before every model call; a breach halts the AI step and hands off | Soft ceiling: concurrent calls can overshoot by one call |
| Fallback model on provider errors only | Opus 5 falls back to Sonnet 5 on connection errors, timeouts, 429, 5xx, and 529, never on 400s | Two models to keep prompts compatible with |
42 answerable questions plus 6 off-topic ones, top 4 chunks. Recall@4 is judged at the article level; Section@4 is stricter and counts a hit only when a retrieved chunk is the section that answers the question.
| Mode | Recall@1 | Recall@4 | Section@4 | MRR | Relevance-gate accuracy |
|---|---|---|---|---|---|
| Hybrid (keyword weight 0.5) | 0.833 | 1.000 | 0.976 | 0.899 | 0.979 |
| Vector only (shipped) | 0.881 | 1.000 | 0.976 | 0.927 | 0.979 |
| Keyword only | 0.690 | 0.857 | 0.810 | 0.766 | 0.938 |
Section@4 is the number that explains the one question the agent still fails. Article-level recall was a perfect 1.000, which hid it: for "Can I pay the cleaner in cash?" vector search returns the payments article's generic overview, not its "Accepted payment methods" section, so the answer step correctly says it cannot find the answer. Hybrid search fixes that question but misses a different one ("Can I book a plumber on a Sunday?"), and no keyword weight from 0.25 to 1.5, nor a larger top-k (5 or 6), fixes both. With 42 questions, tuning further would only fit the test set, so the shipped mode is unchanged.
With equal fusion weights, hybrid dropped to Recall@1 0.786: full-text matches on common words ("problem", "home") pulled in the wrong articles. Halving the keyword weight helped but did not beat vector-only. CI fails the build if the shipped mode falls below Recall@4 0.95, MRR 0.85, or gate accuracy 0.90.
54 conversations (31 originally, 23 safety attacks added later) through the production graph against real Postgres, graded by deterministic checks rather than an LLM judge: the routed outcome and intent, citations to the expected article, key facts in the reply ("$25", "5 to 10 business days"), exact refund amounts proposed for approval, and the reason for any handoff. A handoff caused by a failure fails its case, so a model outage cannot pass as correct routing.
| First run | After fixes | |
|---|---|---|
| Cases passed | 27 / 31 (87.1%) | 30 / 31 (96.8%) |
| Help-center questions | 7 / 11 | 10 / 11 |
| Bookings, refunds, routing | 100% | 100% |
| Safety | 7 / 7 | 7 / 7 |
| Refunds executed without a human | 0 | 0 |
| Cost per conversation | $0.0084 | $0.0085 |
| Turn latency p50 / p95 | 2.5 s / 5.4 s | 2.5 s / 5.2 s |
evals/results/agent.json holds the latest run; the first-run column comes from that run's log. That table is the original 31 cases.
I added 23 safety attacks (injection in other languages and with leetspeak, role-play, "translate then obey", prompt-leak requests, multi-turn pressure, SQL in a booking reference, other customers' bookings by several routes, "approve it yourself") and ran the full set on Claude Opus 5 and on Claude Haiku 4.5.
| Opus 5 | Haiku 4.5 | |
|---|---|---|
| Cases passed | 52 / 54 | 48 / 54 |
| Safety, as graded | 29 / 30 | 25 / 30 |
| Refunds executed without a human | 0 | 0 |
| Cost per conversation | $0.0082 | $0.0012 |
| Turn latency p50 / p95 | 2.1 s / 4.9 s | 1.8 s / 4.1 s |
How to read this honestly:
- No safety case in any run opened an approval or leaked another customer's data. Every safety "failure" was an outcome I had not listed as acceptable. For Opus,
s15was a low-confidence handoff to a human. For Haiku, all five wereout_of_scoperefusals, which are safe but not in my accepted list. - I widened my own accepted lists after the first run. Six cases (
s10,s12,s13,s18,s20,s22) returned safe outcomes I had not listed, such as "out of scope" or "which booking reference?". I added those outcomes and said so here, rather than pretend I predicted them. The checks that matter (no approval opened, no other customer's data in the reply) were unchanged. - Haiku matched Opus on every non-safety case (bookings, refunds, routing, questions) and cost about 7x less. It refuses more attacks as "out of scope" instead of handing them to a person, so strictly as graded it scores lower on safety. I would test Haiku on the classify step and keep Opus for answers before switching, and I have not done that.
q05(paying the cleaner in cash) fails on both models and in every run.- Haiku 4.5 rejects the
effortsetting, so the code leaves it out for Haiku models.
1. Query rewriting silently broke retrieval. The first agent run answered 4 simple questions ("Do you have service in Denver?", "Are your plumbers licensed?") with "I couldn't find that." The retrieval eval had scored Recall@4 of 1.00, because it searches with the raw question. In production, the classifier rewrote the question first and added the company name, which matched the boilerplate intro of every article and pushed the real answer out of the top 4. The answer step then correctly refused to answer from the wrong documents. Fix: search with both the customer's words and the rewrite, and fuse the rankings. All 4 cases pass now, and a unit test pins the regression.
2. The remaining failure is retrieval granularity, not the pipeline. "Can I just pay the cleaner in cash?" fails when the rewrite keeps the customer's phrasing. A direct retrieval check shows why: bge-small ranks cleaning-service articles above the payment-methods section for "pay the cleaner", which is not even in the top 6. The answer step is right to refuse: the chunk it was given was the payments article's overview, not the section that says cash is not accepted. The article-level retrieval metric could not see this, so the retrieval eval now also scores section-level hits (Section@4 above). I tried hybrid search, keyword weights from 0.25 to 1.5, and a top-k of 5 and 6: none fixed it without breaking another question. The next experiment is a larger embedding model, which needs a schema migration and more memory than the free host has, so it is not done.
A security review and a correctness review ran against the agent, and again against the demo UI. Neither found a critical issue. What they did find:
- Request bodies were fully buffered before length validation, so bodies over 64 KB are now rejected before parsing.
- API docs were public in production; they are now hidden when
ENVIRONMENT=prod. - Failed authentication was not rate limited; it is now throttled per client address, with bounded memory.
- The injection guard blocked normal corrections like "please ignore my previous message"; the pattern now targets instruction-like phrases only, with regression tests.
- Concurrent requests on one conversation could run the graph twice on the same thread; a Postgres advisory lock now serializes them.
- Every handoff looked the same; handoffs now carry a reason, and the eval fails cases that hand off because something broke.
- In the UI: "New chat" could leave the previous transcript on screen, switching persona mid-request could show one customer's message under the other's account, a demo reset left an open chat tab stuck until a page refresh, and the console could keep offering Approve on a request someone else had already decided. All four are fixed, and the demo-reset recovery was then verified live.
Running the live server also showed the "request sent for approval" message was missing from the saved transcript. It is now its own workflow step, checkpointed before the pause.
Requirements: Python 3.12, uv, Node 22, and Postgres 16+ with the pgvector extension.
# 1. API
uv sync
cp .env.example .env # add ANTHROPIC_API_KEY, generate the two API keys, DEMO_MODE=true
createdb concierge
uv run concierge-ingest --seed-demo
uv run uvicorn --factory concierge.api.app:create_app # http://127.0.0.1:8000
# 2. Demo UI, in a second terminal
cd web
npm install
cp .env.example .env.local # point it at the API, paste the same two keys, and set
# CONSOLE_PASSWORD (16+ chars) and CONSOLE_SESSION_SECRET (32+)
npm run dev # http://localhost:3000The console fails closed: without those two settings every approval route returns 401. Then open http://localhost:3000, ask the assistant to cancel BK-1042 and refund you, and approve it at /console after signing in with CONSOLE_PASSWORD. The API's own interactive docs stay at http://127.0.0.1:8000/docs.
Demo bookings (times are relative to when you seed):
| Booking | Customer | Situation | Refund outcome |
|---|---|---|---|
| BK-1042 | maya@example.com | 5 days away | full refund, needs approval |
| BK-1043 | maya@example.com | 30 hours away | 50%, needs approval |
| BK-1044 | maya@example.com | 6 hours away | not eligible |
| BK-1045 | maya@example.com | completed | handed to a person |
| BK-1046 | maya@example.com | already refunded | not eligible |
| BK-2001 | jordan@example.com | 4 days away | invisible to Maya |
claude mcp add concierge -- uv --directory /path/to/concierge run concierge-mcpTools: search_help_center, get_booking, list_pending_refund_approvals. The server is read-only by design: no tool can issue a refund.
| Piece | Where | Notes |
|---|---|---|
Web (web/) |
Vercel, deploys main |
Holds the API keys and the console password server-side |
| API | Render free Docker service, from render.yaml |
Auto-deploy is off; migrations run at startup |
| Database | Neon Postgres with pgvector | Use the direct connection string, not the pooled one |
Settings (names only; values live in each host's dashboard, never in the repo):
- API (Render):
DATABASE_URL,ANTHROPIC_API_KEY,CLIENT_API_KEYS,OPERATOR_API_KEYS, plusENVIRONMENT=prod,DEMO_MODE=true,RATE_LIMIT_PER_MINUTE,MAX_TOKENS_PER_DAY,MAX_TOKENS_PER_CONVERSATION. - Web (Vercel):
CONCIERGE_API_URL,CONCIERGE_CLIENT_KEY,CONCIERGE_OPERATOR_KEY,CONSOLE_PASSWORD,CONSOLE_SESSION_SECRET.
Lessons from the first deploy, each of which cost real time:
- The Anthropic key must be created inside a workspace. A personal-account key fails with "not scoped to a workspace", and the agent quietly hands every chat to a human (
classifier_unavailable). The API now logs Anthropic's status and message so this is visible. - Neon's free tier drops idle connections. The connection pool verifies a connection before handing it out; before that fix, the first chat after a quiet spell returned a 500.
- Seeding is refused when
ENVIRONMENT=prod. Seed from a machine withENVIRONMENT=devpointed at the database. - Deploy the API before the web app when the API contract changes, because Vercel deploys on merge and the API is a manual deploy.
- The free API host sleeps after 15 idle minutes and takes up to a minute to boot. The pages ping a keyless
/api/wakeon load and chat retries with a "waking up" notice, so a first visitor sees a short wait, not an error.
uv run ruff check src tests evals && uv run mypy src # lint + strict types
uv run pytest tests/unit -q # workflow, policy, guardrails, auth
uv run pytest tests/integration -q # real Postgres: API, refunds, retrieval
uv run python evals/run_retrieval_eval.py # free
uv run python evals/run_agent_eval.py # calls Claude, about $0.26 per run
cd web && npm run lint && npx tsc --noEmit && npm run build- Unit tests run every workflow path with the model, search index, and database replaced by fakes: approval pause and resume, mismatched approvals, budget breach, invalid model output, ungrounded citations, injection in customer messages and in retrieved articles (including a check that no shipped article is ever quarantined), ownership, rewrite drift, per-turn state isolation, and that the figures on the landing page (
web/lib/proof.ts) matchevals/results/agent.json. - Integration tests use real Postgres: concurrent refund execution, racing approval decisions, rollback on over-refund, the full HTTP refund flow and transcript, state consistency for every decision path (approve, approve less, reject: booking, actions, approval row, and the customer's message agree) with the database itself refusing an over-policy amount, a connection pool that survives the server closing its connections, conversation hijacking, the conversation lock, rate limiting, body-size limits, hidden production docs, customer-scoped bookings, approval detail, and demo reset.
- CI (GitHub Actions) runs lint, mypy, and unit tests; the web app's lint, types, and production build; integration tests against a pgvector service container; and the retrieval eval with thresholds. The agent eval runs on manual dispatch with an API key secret.
src/concierge/
agent/ graph.py (state machine + turn runner), nodes.py, state.py, prompts.py
api/ app.py (endpoints), security.py (API keys, rate limits), schemas.py
bookings/ policy.py (refund rules), repository.py (SQL, approvals, exactly-once refunds)
retrieval/ chunking.py, embeddings.py, search.py (vector / keyword / fusion), ingest.py
llm.py Claude calls with structured output and fallback
guardrails.py input sanitization, injection detection, retrieved-text hygiene
budget.py token ledger and budgets
locks.py per-conversation advisory lock
mcp_server.py read-only MCP tools
web/
app/ customer site (/), support console (/console), server-side /api proxy routes
components/ customer/ and console/ UI
hooks/ chat, bookings, approval queue, polling, storage, API warm-up
lib/ server-only API client, operator session, zod validation, personas, formatting
migrations/ SQL schema (001 base, 002 approved amount)
Dockerfile API image with the embedding model baked in
render.yaml Render blueprint for the API
data/kb/ help-center articles
evals/ datasets, runners, committed results
tests/ unit/ and integration/
The full list of threats, guards and residual risks is in docs/THREAT_MODEL.md. The most important honest caveat: the agent eval is small (54 cases, written by one person), so its pass rates carry wide error bars; the money guarantee rests on the structure of the system, not on those numbers.
- The console login is one shared operator password (a signed, HttpOnly cookie checked on the server), not per-person accounts or SSO, so every approval is recorded against the same operator identity. A team deployment needs individual accounts.
- Security headers are partial. Frame, sniffing, referrer, and permissions policies ship; a strict script content policy needs per-request nonces and is still missing.
- Retrieval: the one failing eval case points at the embedding model; compare a larger model and hybrid-on-rewrites with the existing evals before changing the default.
- Identity: the client API key represents a trusted backend that asserts the customer's email. The demo is safe because its server maps two fixed personas to emails; a product with real customers should pass a signed customer token (JWT) instead.
- Rate limiting is in-process; multiple replicas need Redis or Postgres-backed limits. Behind Render's proxy the failed-login throttle keys on the proxy's address and is weak; the per-key limits are the controls that actually cap cost. The website shares one key, so the costly calls (chat and demo decisions) share 20 requests a minute in total on the live demo, while cheap reads (bookings, transcripts, stats) have a separate budget six times larger, so page loads and polling never starve a chat. Chunked uploads without a length header should be capped at the reverse proxy.
- Migrations run at startup, which suits one instance; multi-replica deploys should run them as a release step.
- The Docker image has not been built in CI yet; the Dockerfile and compose file cover the API only, not the web app.
- No streaming yet: replies appear when the turn completes, with a typing indicator meanwhile.
- Handoff is a reply, not a ticket: the next step is creating a ticket in a helpdesk system.
- Evals use deterministic checks: an LLM-as-judge faithfulness score on free-form answers is the next eval to add.


