ClauseHound:拒绝猜测的宠物保险决策引擎
ClauseHound: a pet-insurance decision engine that refuses to guess
ClauseHound 是一款宠物保险决策引擎,针对"买保险还是存钱""终身成本对比""理赔金额分析""带病换保"四个决策,从保险公司自身条款文本中给出带原文引用的答案。其数据存于 Sanity,含 5 家保险公司、28 条保障条款、30 条免责条款、18 个等待期等 106 份文档,金额计算由代码以整数美分完成,LLM 只做解释。无法核实的问题会明确拒绝回答,条款与免责冲突时并列展示且不裁决。
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.
🔗 Demo: https://clausehound-1rmpnbeb5-lingzhoudesign-gmailcoms-projects.vercel.app
🗄️ Sanity project ID: yijqzehr (dataset production)
🏷️ Tag: #sanitychallenge
What I Built
Your Bernese Mountain Dog needs hip surgery at age three. You file the claim, confident — and get denied. The 12-month orthopedic waiting period was in Section 4 all along. You just never found it before you bought the policy.
That's the moment ClauseHound is built for — except before it happens. ClauseHound is a pet-insurance decision engine for people shopping for a policy. Not a chatbot that answers trivia about pet insurance: it takes the four decisions shoppers actually agonize over and answers each one from the insurers' own policy text, with the exact clause cited:
- Insure vs. save — "Should I get pet insurance or just put $60/month in a savings account?" ClauseHound runs the deterministic math (premiums vs. vet bills over time, in integer cents, in code — the LLM never does arithmetic) and tells you plainly when saving wins. It says "save" when saving wins.
- Lifetime cost comparison — paste two quotes and it computes the true 10-year cost per plan: premiums + deductible + reimbursement math, side by side, exact to the cent.
- Claim-payment analysis — "My dog needs a $4,200 cruciate surgery in month 8 — what would each insurer actually pay?" Answered per-carrier, from the coverage clauses, exclusions, and waiting periods that govern that claim.
- Switching guidance — "My dog has a pre-existing skin condition; which insurers will still cover her if I switch?" The highest-stakes question in pet insurance, answered carrier by carrier, with the pre-existing-condition clauses cited.
One grounding contract runs through all four:
- Verdict first, citations follow. Every answer opens with the decision, then shows its work.
- No evidence, no answer. Ask about something the corpus doesn't cover and you get "I couldn't verify this in the policy documents" — not a guess, a refusal, with pointers to what is answerable.
-
Conflicts stay unresolved. When a coverage clause and an exclusion both match, both surface side by side, marked
unresolved. The agent never picks a winner. - Owner-reported facts stay visibly separate from verified policy facts. What you told it and what the policy says are never blended.
- Carrier-specific claims require policy citations. No citation, no claim.
Demo
No login needed. The landing page frames the four decisions; each job card (and the hero) opens the chat modal with its question already prefilled — you press Send, it never auto-sends.
Try these in order:
- Insure vs. save — press Send on the prefilled question. It asks three quick questions that change the answer (dog's age, monthly savings, breed/size) — nothing more. Watch the loading stages narrate the pipeline ("Pulling the relevant policy text…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…"), then the verdict: for a healthy young dog with no breed risks, it will tell you saving wins, with the math shown.
- The pre-existing switch — "My dog has a pre-existing skin condition. Which insurers will still cover her?" Five carriers, five cited answers, one table.
-
The conflict — "Are cruciate ligament injuries covered in the first year?" Some carriers cover, some exclude as pre-existing within the waiting period. Both clauses surface, side by side,
unresolved.

The landing page frames the four decisions, not a chat box.

Each job card opens the modal with its question prefilled. Nothing auto-sends.

The answered state: "There's no single winner — it depends on when the emergency hits." Verdict first, then the month-by-month math, then the cited clauses.
Code
Next.js + TypeScript app. The interesting parts:
-
app/lib/mcp.ts— Sanity Context MCP client (JSON-RPC over HTTPS). The app never queries Sanity directly; every fact comes through the Context endpoint's tools (initial_context,groq_query, …). -
app/lib/cost.ts— deterministic money math in integer cents: premiums, deductibles, reimbursement rates, annual limits, multi-year scenarios. The LLM explains the numbers; it never computes them. (Tested to the cent —cost.test.ts.) -
app/lib/insureVsSave.ts,switching.ts,denial.ts,renewals.ts,appeals.ts— the four decision jobs as typed modules, each with its own test suite. -
app/components/ChatModal.tsx+ChatApp.tsx— the modal chat: prefilled questions, streaming answers with loading stages, citation cards, contradiction panels, cost cards.
The repo is private (lymcho/clausehound, now merged to main); the test suite (1,653 tests) and the eval fixtures are part of the submission evidence below.
How I Used Sanity
Everything ClauseHound knows lives in Sanity as structured content — not blobs of text, but typed documents with relationships, modeled from the insurers' published US sample policies:
-
5 insurers → 5 policy documents → 28 coverage clauses, 30 exclusion clauses, 18 waiting periods, 20 claim steps — 106 documents in the
productiondataset, verified live.
The schema is the trust strategy. Every clause document carries a required sectionRef (e.g. Sec. 3.4) — the citation anchor — and a required policy reference back to its policyDocument, which references its insurer:
insurer → policyDocument → coverageClause / exclusionClause / waitingPeriod / claimStep
(each with required sectionRef)
The agent reads through a Sanity Context MCP endpoint backed by an Agent Context over that dataset. A GROQ content filter at the context level scopes every query to the four answerable types and only documents attached to a real policy:
_type in ["coverageClause", "exclusionClause", "waitingPeriod", "claimStep"]
&& defined(*[_type == "policyDocument" && _id == ^.policy._ref][0])
That filter is the guardrail: the agent cannot wander into irrelevant content no matter what the user asks. Coverage, exclusions, and waiting periods are distinct document types and stay distinct in answers — which is exactly why the conflict view works: a coverageClause and an exclusionClause matching the same question surface as two typed records with their sectionRefs, not as two paragraphs to blend.
The app runs an agentic loop over the MCP tools and narrates the pipeline at real phase boundaries — "Pulling the relevant policy text…", "Analyzing your question against the fine print…", "Merging the insurer-by-insurer findings…", "Verifying every claim against the policy text…". The grounding isn't just enforced, it's visible while you wait.
This is the bet the challenge asks for: the agent only works because the content was structured. Keyword search over the same policies could find the words "hip dysplasia"; it could not keep five carriers' answers from bleeding into each other, keep a coverage clause distinct from the exclusion that overrides it, or compute a 10-year cost from typed premium/deductible/reimbursement fields. The structure is the product.
What the live runs showed
The release is covered by a 19-fixture held-out eval (eval/fixtures.json, fixtures frozen — never edited to make the app pass). Each fixture posts a real user question to the deployed /api/chat and checks the response structurally: HTTP 200, citations resolve to real clause IDs, no dangling markers, carrier scope correct, refusals refuse, contradictions surface as unresolved. The two cost fixtures go further: they paste real quotes and re-derive every dollar figure with an independent port of the integer-cent math — exact match required, no LLM judging.
16/19 passed on the production build, 2026-10-01 (merged to main as 9957cb1; the merge tree is byte-identical to the evaluated feature commit 6f4e452b, so the results stand for production). Live-answer latencies ran 18–95s (streamed); the deterministic cost fixtures verify in under a second.
Honest failures, since they matter more than the score:
- inv-10 (refusal discipline): the fixture asks about something unanswerable (cloning coverage); the app correctly refused but emitted 5 citations alongside the refusal. A refusal should be clean — citations on a refusal are a leak.
- inv-11 (alternative therapy): asked about acupuncture coverage and returned 0 citations. The corpus's coverage of alternative-therapy clauses is thin; the app should either find the clause or say "couldn't verify" with pointers — it did neither cleanly enough to pass.
-
cfl-01 (conflict surfacing): "What is the waiting period for cruciate ligament treatment?" — the expected
unresolvedcontradiction between carriers' cruciate clauses didn't surface in the eval run. A manual re-run minutes later did surface it (Lemonade 6 months vs. Nationwide 12-month exclusion), so the behavior is flaky across runs rather than absent — which is arguably worse. This is the highest-priority fix: the conflict view is the feature.
Known limitations the eval doesn't hide:
- The corpus is five carriers' published sample policies (US). It doesn't cover every plan variant, rider, or state-specific endorsement.
- The fixtures are authored by me; they measure agreement with my labels, not independent accuracy or retrieval quality.
- Answers stream and can take up to ~2 minutes on hard multi-carrier questions.
Sanity Project Details
- Project ID:
yijqzehr, datasetproduction(106 documents, queried live 2026-10-01) - Document types:
insurer,policyDocument,coverageClause,exclusionClause,waitingPeriod,claimStep - Agent reads through the Sanity Context MCP endpoint (
SANITY_AGENT_CONTEXT_URL), backed by an Agent Context over the production dataset; the GROQ content filter above scopes all retrieval to answerable, policy-attached documents.
No login is needed to try ClauseHound. Live answers are streamed; the deterministic cost comparisons run without the model.
Demo only — research aid, not professional insurance advice. Verify with your insurer before buying.
来源:Google AI:DEV 作者专属(RSS) · dev.to