跳到正文
原文
Google AI:DEV 作者专属(RSS)· Jangwook Kim·· 8 小时前AI 评分43

Turbopuffer 评测:对象存储向量搜索、冷启动与成本权衡

Turbopuffer Review: Object-Storage Vector Search, Cold Starts, and Cost Trade-offs

AI 导读

Turbopuffer 将向量索引放在对象存储上,查询节点无状态、以本地 NVMe 作缓存,冷查询需约 3-4 次对象存储往返、每次约 100 ms。其可验证的最大规模结果为 ANN v3 在 1000 亿个 1024 维 f16 向量上实现 200 ms p99、目标超 1000 QPS,但该场景索引常驻 SSD,与冷租户从对象存储取索引不同。

正文

Why We Brought This Tool Into Our Lab

We started this evaluation with a recurring retrieval-infrastructure problem: most vector databases make inactive tenants expensive.

A conventional deployment keeps indexes on provisioned memory or SSD capacity whether a tenant sends 10,000 queries per hour or one query per month. That works for a small number of uniformly active indexes. It becomes painful when a product has hundreds of thousands of workspaces, repositories, agents, or customer-specific knowledge bases with highly skewed activity.

Turbopuffer takes a materially different approach. In the architecture reference we inspected, the API routes requests to Rust binaries operating over an object-storage source of truth. Query nodes are stateless from a durability perspective. When a namespace is cold, a query node reads it from object storage and caches its documents on local NVMe. The service routes subsequent requests toward the same node to preserve cache locality, although another node can still serve the namespace after a failure or routing change.

That architecture is not ordinary storage tiering. The service commits writes to a write-ahead log under the namespace's object-storage prefix. Separate indexing workers asynchronously convert committed data into searchable index structures. Query nodes use memory and NVMe as disposable acceleration layers rather than as the authoritative copy.

For vector retrieval, Turbopuffer uses SPFresh, a centroid-based ANN design. This matters because graph traversal can require a chain of dependent reads, which is hostile to object-store latency. A centroid hierarchy can identify promising clusters with a bounded number of broad reads. In our architectural analysis, we budgeted roughly three to four object-storage round trips for a cold query, with each round trip on the order of 100 ms.

The trade is straightforward:

  • We avoid paying premium storage prices for every inactive namespace.
  • We accept a meaningful first-query penalty when a namespace is not cached.
  • We depend on routing and cache residency for stable warm latency.
  • We must decide whether approximate retrieval at roughly 90–95% recall@10 is sufficient.
  • We inherit a consistency-versus-latency decision on every latency-sensitive query path.

The strongest large-scale result we could verify was the ANN v3 result framed as 200 ms p99 over 100 billion vectors. Its workload uses 100 billion 1,024-dimensional f16 vectors, representing about 200 TiB of dense vector data, with a target above 1,000 QPS. We did not treat that number as a proxy for an ordinary cold serverless namespace. At that scale, the ANN tree is sized to remain on SSD, so it measures a different operating condition from fetching a dormant tenant's index from object storage.

That distinction drove our review. We were not trying to prove that Turbopuffer can store a huge index. We wanted to know whether its economics survive a real multi-tenant query mix without turning cold starts into user-visible timeouts.

Hands-On Walkthrough: Setup, Execution & Output

The service is exposed through HTTP and requires an API key, so our planned harness had no infrastructure dependency beyond a public embedding dataset and a spend-limited account. We separated the test into four phases:

  1. Insert identical vector corpora into namespaces containing 100,000, 1 million, and 10 million documents.
  2. Poll namespace metadata until unindexed_bytes returned to zero.
  3. Query previously untouched namespaces to capture the first-query latency.
  4. Run repeated queries against the same namespace to capture the warm-up curve and steady-state percentiles.

We also planned to write a sentinel document, query it immediately with strong consistency, and poll metadata independently. That separates write visibility from asynchronous indexing completion: a document can be visible through exhaustive WAL search before it has entered the ANN index.

Our execution environment did not contain a Turbopuffer API credential, an offline server, or verified protocol fixtures. We therefore stopped before issuing hosted-service requests. We did not substitute a cache simulation or another vector database and label that as a Turbopuffer benchmark. Doing so would have produced precise-looking but useless numbers.

The following harness is the HTTP timing collector we prepared. It is runnable once TPUF_QUERY_URL, TPUF_API_KEY, and a valid query payload are supplied. The URL is parameterized deliberately so that CI does not hard-code a region or API revision.

python3 -m venv .venv
. .venv/bin/activate
python -m pip install "httpx==0.27.2"

export TPUF_API_KEY="replace-with-a-spend-limited-key"
export TPUF_QUERY_URL="https://replace-with-current-region-and-query-path"
export SAMPLE_COUNT="100"

cat > query.json <<'JSON'
{
  "rank_by": ["vector", "ANN", [0.01, 0.02, 0.03]],
  "top_k": 10,
  "include_attributes": ["category"]
}
JSON

cat > bench.py <<'PY'
import json
import math
import os
import statistics
import sys
import time

import httpx

api_key = os.environ.get("TPUF_API_KEY")
query_url = os.environ.get("TPUF_QUERY_URL")
sample_count = int(os.environ.get("SAMPLE_COUNT", "100"))

if not api_key:
    sys.exit("ERROR: TPUF_API_KEY is not set")
if not query_url:
    sys.exit("ERROR: TPUF_QUERY_URL is not set")

with open("query.json", "r", encoding="utf-8") as handle:
    payload = json.load(handle)

latencies_ms = []
status_counts = {}

with httpx.Client(timeout=30.0, http2=True) as client:
    for sequence in range(sample_count):
        started = time.perf_counter_ns()
        response = client.post(
            query_url,
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
            },
            json=payload,
        )
        elapsed_ms = (time.perf_counter_ns() - started) / 1_000_000
        status_counts[str(response.status_code)] = (
            status_counts.get(str(response.status_code), 0) + 1
        )

        if response.is_error:
            print(
                json.dumps(
                    {
                        "sequence": sequence,
                        "status": response.status_code,
                        "latency_ms": round(elapsed_ms, 2),
                        "body": response.text[:1000],
                    }
                ),
                file=sys.stderr,
            )
            response.raise_for_status()

        latencies_ms.append(elapsed_ms)
        print(
            json.dumps(
                {
                    "sequence": sequence,
                    "latency_ms": round(elapsed_ms, 2),
                    "phase": "first" if sequence == 0 else "subsequent",
                }
            )
        )

ordered = sorted(latencies_ms)

def percentile(value):
    index = max(0, math.ceil((value / 100) * len(ordered)) - 1)
    return ordered[index]

summary = {
    "samples": len(ordered),
    "first_query_ms": round(ordered[0] if len(ordered) == 1 else latencies_ms[0], 2),
    "p50_ms": round(statistics.median(ordered), 2),
    "p90_ms": round(percentile(90), 2),
    "p99_ms": round(percentile(99), 2),
    "status_counts": status_counts,
}
print(json.dumps({"summary": summary}))
PY

python bench.py

The following is simulated output showing the collector's format, not a result from our lab:

{"sequence": 0, "latency_ms": 886.42, "phase": "first"}
{"sequence": 1, "latency_ms": 18.71, "phase": "subsequent"}
{"sequence": 2, "latency_ms": 15.64, "phase": "subsequent"}
...
{"summary":{"samples":100,"first_query_ms":886.42,"p50_ms":14.93,"p90_ms":18.27,"p99_ms":28.84,"status_counts":{"200":100}}}

We chose those illustrative values to resemble the published control ranges so the output would look realistic. They are not independent measurements.

The verified architectural baseline for 1 million documents is a first-query p50 of 874 ms and a cached p50 of 14 ms. For our comparison, we used a separate 10-million-document, 1,024-dimension, approximately 40 GB workload at 8 QPS and top_k=10 as a reference: warm p50/p90/p99 values of 14/17/27 ms and cold values of 874/1,214/1,686 ms. We did not independently measure these endpoints.

These are useful controls, but they are not a reproduced warm-up curve. We could verify the cold and warm endpoints in the published material, not the intermediate transition across queries or elapsed time.

Test Limitations and Operational Constraints

Our immediate roadblock was test access. Without a hosted API key, we could not truthfully produce service-side error responses, independent percentiles, cache-eviction measurements, or indexing-delay distributions.

If run without an API key, the prepared harness would exit with the following message; we did not execute it in this evaluation:

ERROR: TPUF_API_KEY is not set

That is our harness error, not a Turbopuffer API response. We are making the distinction explicit because fabricated vendor error strings are worse than a missing test.

Several operational constraints still surfaced during our review.

First, “the second query is warm” is too simplistic. Cache locality is node-specific, and any query node can serve any namespace. A failover, rebalance, or eviction can put the next request back on an object-storage path. We would use the warm-cache endpoint or a pre-flight query before interactive traffic, but that transfers responsibility to the application. It also creates one warm-up operation per active namespace.

Second, strong consistency imposes an approximately 10 ms latency floor because the query path checks object storage for newer writes. Eventual consistency can get below that floor, but it permits staleness of up to roughly one hour in the worst case. In our consistency assessment, we accounted for eventual queries searching up to 128 MiB of unindexed data and used the greater-than-99.8% consistent-response figure as a reference, not as an independently measured result. That is encouraging, but it is not a guarantee suitable for permission changes, deletion enforcement, or user-facing read-after-write workflows.

Third, each namespace currently produces at most one WAL entry per second. Concurrent writes can be grouped, preserving high aggregate throughput, but a batch opened shortly after the preceding commit can take up to one second. We would batch ingestion intentionally rather than sending a stream of tiny writes and expecting database-style single-row latency.

Fourth, schema flexibility has hard edges. Attribute types must remain consistent within a namespace. Every vector in a vector column must use the same dimensionality. A namespace can have at most four embedded attributes, and an upsert batch containing embedded attributes is limited to 30 rows. We could verify those limits, but not the exact API messages emitted when they are violated.

That leaves an important gap in our evaluation: we cannot provide authentic response bodies for a type conflict, non-indexed filter, dimension mismatch, fifth embedded attribute, or oversized embedded upsert. Teams automating error classification should capture those responses during their own proof of concept instead of matching on guessed strings.

Finally, cache behavior remains underspecified for capacity planning. We could not establish an eviction policy, a guaranteed residency duration, or a namespace-count threshold at which warm p99 deteriorates. This is the main unresolved risk for products with a very long tenant tail. If that describes your workload, bring a production-shaped Zipf distribution to the test rather than benchmarking one namespace in a loop.

For help designing that workload, contact our infrastructure team or review the systems in our AI tools collection.

Scale, Latency & Cost vs. Alternatives

The architectural comparison is clearer than the invoice comparison.

Decision factor Turbopuffer Pinecone Self-hosted Qdrant
Durable storage model Object storage as source of truth Managed service model varies by product and plan Operator-selected disk, replication, and snapshots
Cold namespace behavior Direct object-store reads, then NVMe caching Depends on selected service architecture and capacity model Normally resident on provisioned node storage
Warm latency evidence in scope 10M workload: 14 ms p50, 17 ms p90, 27 ms p99 Must be measured on the purchased plan Must be measured on the selected hardware
Cold latency evidence in scope 10M workload: 874 ms p50, 1,214 ms p90, 1,686 ms p99 Not measured in this review Usually an OS-cache or process-start concern
Operations burden Low Low High
Inactive-namespace economics Strong architectural fit Quote and workload dependent Capacity remains provisioned unless nodes scale down
Read-after-write Strong by default, with latency floor Plan and API dependent Configuration dependent
Primary risk Cold-tail latency and cache uncertainty Metering and vendor-plan economics Staffing, replication, upgrades, and capacity planning

We could not produce a defensible current-dollar Pinecone comparison without fixing a plan, cloud, region, pod or serverless model, read units, write units, metadata size, and query rate. Any single “Turbopuffer is X times cheaper” figure would hide the variables that dominate the invoice.

We can still establish a useful storage floor. The infrastructure reference rates we validated are approximately $0.02 per GB-month for object storage and $0.10 per GB-month for SSD cache. These are component-cost references, not Turbopuffer customer pricing.

For the documented 10-million-vector workload occupying approximately 40 GB:

Object-storage component = 40 GB × $0.02 = $0.80/month
Full SSD-cache component = 40 GB × $0.10 = $4.00/month

Normalized to 1 million vectors, that becomes approximately $0.08 per month for the object-storage component and $0.40 per month if the corresponding 4 GB remains cached on SSD. This excludes compute, queries, indexing, metadata, replication overhead, service margin, networking, support, and minimum commitments.

A more practical model is:

Monthly hierarchy cost floor = D × (0.02 + A × 0.10)

Here, D is stored gigabytes and A is the average fraction resident in SSD cache. With 40 GB and 10% average cache residency, the component floor is approximately $1.20 per month. At 100% residency, it is approximately $4.80.

This explains where Turbopuffer should win: large D, low A, and enough tolerance for cold requests. It does not prove the final invoice beats Pinecone because query and service charges are missing.

Against Qdrant, we would compare Turbopuffer's invoice with the complete monthly cost of nodes, attached storage, replicas, snapshots, cross-zone traffic, monitoring, upgrades, and engineering time. Qdrant can be the better choice when the dataset is continuously hot, latency must be predictable, and a team already operates stateful search clusters. Object storage becomes more compelling as the inactive-to-active namespace ratio rises.

Our Final Verdict: When to Deploy, When to Skip

Turbopuffer's architecture is technically coherent. Rust query binaries, object-storage durability, disposable compute, NVMe locality, SPFresh clustering, and asynchronous indexing fit the economics of sparse multi-tenant retrieval better than keeping every tenant permanently resident.

The cold-start penalty is not theoretical. The verified 10-million-document controls span 27 ms warm p99 versus 1,686 ms cold p99. That difference is large enough to change an application's request flow, timeout policy, and user experience.

Deploy this if:

  • We have many isolated namespaces but only a small active working set.
  • We can pre-warm a namespace before an interactive retrieval path.
  • A roughly one-second cold request is acceptable outside the critical path.
  • We value managed operation more than absolute latency control.
  • We can use strong consistency selectively and eventual consistency elsewhere.
  • A 90–95% recall@10 target is acceptable and we will validate recall on our corpus.
  • Our writes can be batched around the one-WAL-entry-per-second namespace behavior.

Hold off or avoid it if:

  • Every request must meet a strict sub-100 ms p99, including after eviction or failover.
  • We require guaranteed read-after-write behavior below the strong-consistency latency floor.
  • Our schema needs more than four embedded attributes.
  • We rely on large embedded upsert batches exceeding 30 rows.
  • We cannot tolerate approximate retrieval or independently test recall.
  • We need a documented cache-residency guarantee.
  • We expect a reliable cost comparison without testing our actual mix of queries and namespaces.

Our conclusion is positive on architectural fit but conditional on workload validation. We would shortlist Turbopuffer for a product with thousands or millions of intermittently active tenant indexes. We would not approve production deployment until we reproduced the namespace warm-up curve, measured p99 across several corpus sizes, captured indexing delay under concurrent writes, and recorded the exact schema and filter errors from the hosted API.

The relevant technical starting points are the architecture overview, the ANN v3 design, and the concepts and recall model. If the cold-start budget is uncertain, we would run the proof of concept before negotiating around headline storage savings.

来源:Google AI:DEV 作者专属(RSS) · dev.to