跳到正文
原文
Google AI:DEV 作者专属(RSS)· Tushar Agarwal·· 4 小时前AI 评分40

你的 AI 智能体值这些 token 吗?我们用 TigerGraph 做了实测

Is your AI agent worth its tokens? We measured it with TigerGraph

AI 导读

在 TigerGraph Agentic GraphRAG 黑客松中,作者构建 OCCAM 路由系统,用同一模型(Gemini Flash,temperature 0)和 100 道奥运问题对比四条流水线。

正文

Title: Is your AI agent worth its tokens? We measured it with TigerGraph
Tags: ai, rag, graph, python


Everyone is bolting agents onto retrieval. Almost nobody asks what they cost.

For the TigerGraph Agentic GraphRAG Hackathon, the guidebook states the real question: it is not whether agentic produces a better answer, but whether the extra reasoning and retrieval steps are worth the extra token cost. We built OCCAM to answer that with numbers.

Named for Occam's razor: entities should not be multiplied beyond necessity, and neither should retrieval steps.

What we ran

The same 100 public questions about Olympic events (2,951 Wikipedia documents, 2,210 event pages plus 740 distractors), through four pipelines, with one model (Gemini Flash, temperature 0):

Pipeline Accuracy Tokens / question
RAG 63% 3,259
GraphRAG 97% 838
Agentic GraphRAG 100% 890
OCCAM (routed) 100% 64

All four pipelines execute the same query plan against the same graph. Only the author of the plan differs, so any difference in the numbers comes from the reasoning strategy, not the plumbing.

Where each approach breaks

  • RAG fails structurally on counting. "How many cycling events had more than 30 competitors?" needs a median of 11 documents and up to 43. No top-10 retrieval holds that evidence. Aggregation accuracy: 10%.
  • GraphRAG gives up on ambiguity. When the graph returns a shortlist instead of an answer, a single pass stops. It lost 3 of 100 questions that way.
  • The agent pays on every question. It reaches 100%, but spends about 890 tokens each time, including on the 97 questions that did not need it.

The agent tier changed the answer on 3 of 100 questions and broke none. That is the finding: agents are decisive on about 3% of questions and overhead on the rest.

The router

OCCAM sends each question to the cheapest tier that can answer it:

  1. Tier 0: a rule maps the question to a graph query. No LLM call.
  2. Tier 1: one planning call.
  3. Tier 2: the full agent loop, re-planning with the executor's complaint.
  4. Tier 3: hybrid vector + BM25 document retrieval, then an LLM reader.

A tier only hands over when it can name what it was missing, and that gap is recorded in the trace. The controller never silently skips evidence. 98 of 100 questions were answered at tier 0 with zero tokens.

Where TigerGraph comes in

The graph holds Games, Events, Venues, Sports, Athletes and NOCs, with edges like HAS_EVENT, AT_VENUE, WON_GOLD and PREV_EDITION. Chunk embeddings live on a Chunk vertex as a TigerVector attribute, so a similarity search can run inside a traversal.

Each agent tool is an installed GSQL query on a TigerGraph Savanna workspace (4.2.5). Aggregation runs entirely in the database: count_above walks every event of a sport at one Games and returns the count, so the model never sees the 8 to 43 documents involved. A parity script runs the same questions through the database and our reference implementation: 78 checked, 0 mismatched.

Honest limits

  • Tier 0 is a cache for the five question shapes in this benchmark. On six paraphrased questions that no rule matches, all six fell to tier 1 and were answered correctly, at about 830 tokens each.
  • The benchmark token and latency figures come from the local reference store, not Savanna. They agree where both have a route.
  • One model, one run, no variance bars.

We also report evidence recall next to accuracy. RAG scores 63% accuracy against 76% recall, and on superlatives its recall is 0.16: it names a plausible winner without retrieving the documents that settle it. Accuracy alone would have hidden that.

Try it

The takeaway: the value is not in having an agent. It is in knowing which questions need one.

来源:Google AI:DEV 作者专属(RSS) · dev.to