跳到正文
原文
Google AI:DEV 作者专属(RSS)· Payload·· 2 小时前AI 评分55

x402 生产环境实战:七个没人提醒你的故障模式

x402 in production: the failure modes nobody warns you about

AI 导读

作者分享自己实现 x402 按次付费协议(402、验证、服务三步)时遇到的真实故障模式:签名重放攻击需用一次性 nonce 加 5 分钟过期防护;开发用 HMAC 验证器若带到生产会导致请求免费通过,应在启动时断言验证器类型并拒绝启动。

正文

x402 in production: the failure modes nobody warns you about

The x402 pay-per-call pattern looks simple — 402, verify, serve. Here's how it actually breaks in production, how to catch each failure, and the checklist I run before shipping.


x402 is a clean protocol on paper: an unpaid request gets 402 Payment Required with machine-readable payment requirements, the agent pays, you verify, you serve. Three steps. I implemented it for real, and the protocol wasn't the hard part — the failure modes around it were. This is the article I wish existed before I shipped.

Every scenario below is something I hit or deliberately tested. Code is from a working Express implementation.

1. The replay attack you don't have (until you do)

A payment signature is just bytes. If your verifier accepts the same signature twice, a captured payment is a free-forever pass.

How it breaks: Agent pays once for /api/report, captures the successful payment payload, replays it for every subsequent call. Your verifier says "valid" every time because it was valid — once.

How to catch it: Your ledger should record a unique identifier per payment. Query it: if the same payment reference appears twice, you're being replayed.

The fix: every payment requirement carries a short expiry (we use 5 minutes) and a per-request nonce. The verifier rejects expired or already-seen nonces:

// Payment requirements must be single-use and short-lived.
// Without BOTH of these, a captured signature is a permanent free pass.
paidRoute({
  price: '0.25',
  payTo: '0xYourWallet',
  network: 'base',
  asset: 'USDC',
  verifier,   // must check nonce + expiry, not just signature validity
  ledger,     // append-only record of every verified payment
})

Debug step: when a request is rejected as a replay, log the nonce and its first-seen timestamp. If you see the same nonce from different IPs within seconds, it's an attack, not a bug. If you see it from the same agent retrying, your 402 re-challenge is probably dropping the fresh nonce — check that each 402 carries a new one.

2. The dev verifier that ships to production

For local testing, an HMAC-based verifier is fine: you hold the secret, you sign, you verify. It is also a loaded gun pointed at your production API.

How it breaks: the dev verifier trusts whoever holds the secret. If it ships, anyone who finds (or guesses) the dev path pays nothing. Worse, it looks like it's working — requests verify, the ledger fills, revenue doesn't.

How to catch it: this one is silent, which is what makes it dangerous. The check is structural, not behavioral: assert the verifier type at startup and refuse to boot wrong.

const { createDevVerifier, createFacilitatorVerifier } = require('./x402-core');

// Local only. Never ship this.
const devVerifier = createDevVerifier({ secret: process.env.X402_DEV_SECRET });

// Production: verification goes through a real x402 facilitator
// (self-hosted x402.rs, thirdweb, Coinbase CDP — same pattern).
const verifier = createFacilitatorVerifier({ verifyUrl: 'https://your-facilitator.example/verify' });

Production-readiness check: NODE_ENV=production + dev verifier = crash on boot, not a warning. I mean it — a log line gets ignored; a crash gets fixed.

3. Clock skew eats valid payments

Your payment requirements expire in 5 minutes. The agent's clock, the facilitator's clock, and your server's clock disagree by 90 seconds. Valid payments start getting rejected near the expiry boundary, and the failures look random.

How to catch it: on every expiry rejection, log three timestamps: requirement issued-at, requirement expiry, and your server's now. If rejections cluster within ~2 minutes of expiry, it's skew, not fraud.

The fix: keep the expiry tight (5 minutes is right for replay protection) but log the margin on every rejection so skew is visible. And NTP-sync your servers — obvious, skipped constantly.

4. Facilitator down: fail open or fail closed?

Your verifier POSTs to a facilitator to check the chain. The facilitator times out. What does your API do?

How it breaks: no timeout on the verify call → requests hang until the client gives up. Or worse, someone wrote a catch that returns "verified: true" on error — congratulations, your API is now free whenever the facilitator hiccups.

The fix: short timeout, fail closed, ledger the attempt.

Expected responses:

  • Facilitator confirms payment → 200 + resource
  • Facilitator rejects → 402 again, fresh nonce
  • Facilitator unreachable/timeout → 502 with a retryable error body, never the resource

Debug step: if paid users report intermittent 502s, check facilitator latency percentiles before blaming your code. We ledger every verify attempt with its outcome for exactly this.

5. The amount that's almost right

Agent pays 0.249 USDC. Your route demands 0.25. Or it pays on Base while you expected mainnet. Or USDC vs. a bridged variant with a different contract address.

How it breaks: exact-match verification rejects it, the agent retries with the same almost-right payment, and you both burn cycles. From the agent's logs it looks like your API is broken; from yours it looks like underpayment.

How to catch it: on every payment mismatch, log expected vs. actual as structured fields — amount, asset contract, network — not just "verification failed." The pattern in the mismatch tells you whether it's a rounding bug (theirs), a config bug (yours), or an asset confusion (both).

Production-readiness check: at startup, validate that payTo is a checksummed address on the expected network and that the asset contract matches. A misconfigured recipient address doesn't fail loudly — funds just go to the wrong place.

6. Manifest drift: the price list lies

/.well-known/x402 advertises your prices so agents can discover them without guessing. It's a separate code path from the enforcement in paidRoute.

How it breaks: you change a price in the route config and forget the manifest (or vice versa). Agents pay the manifest price, get 402'd anyway, and conclude your API is broken. This is the most common "it works in testing" failure because tests usually hit the route directly and never read the manifest.

How to catch it: a parity test. Fetch your own manifest, then for each listed endpoint, assert the enforced price matches:

curl -s http://localhost:3402/.well-known/x402 | jq '.endpoints'
# compare against the price: values in your paidRoute() configs — by hand or in CI

Production-readiness check: make this a test, not a manual step. Nine end-to-end tests beat one careful deploy.

7. The bare 402

The laziest x402 implementation returns 402 with no body and no PAYMENT-REQUIRED header. Technically compliant. Practically useless — the agent can't self-serve payment, so it either gives up or files a support ticket. You've built a paywall with no cashier.

Expected response: every 402 must carry machine-readable requirements — price, asset, network, recipient, expiry, nonce. If your 402 doesn't tell the agent how to pay, you haven't implemented x402; you've implemented a door.

The production-readiness checklist

Before any x402 endpoint serves real traffic:

  1. Verifier type asserted at boot — production refuses to start on a dev verifier.
  2. Nonce + expiry enforced — single-use payment requirements, 5-minute expiry, replay logged.
  3. Facilitator timeout + fail closed — hung verifies become 502s, never free resources.
  4. Manifest/enforcement parity tested — advertised prices match enforced prices, in CI.
  5. Ledger enabled — every verified payment appended (JSON-lines, auditable). Disputes without a ledger are arguments without evidence.
  6. HTTPS everywhere — payment requirements are only as trustworthy as the transport.
  7. payTo validated — checksummed address, correct network, correct asset contract, checked at startup.
  8. Mismatch logging structured — expected vs. actual amount/asset/network on every rejection.

Where the manual process gets expensive

Everything above is implementable by hand — the checklist is the whole architecture, and a careful developer can build it in a week or two. Where it gets repetitive is the second endpoint, and the tenth: the verifier wiring, the nonce bookkeeping, the ledger appends, the manifest route, and the parity tests are the same boilerplate every time, and every copy is a new place to get #2 or #6 wrong.

That's what I packaged into the x402 Paid API Starter Kit: the protocol core, Express middleware, manifest route, a working demo (one free route, two priced routes, ledger viewer), and 9 end-to-end tests including the parity and replay checks above. One dependency (express), Node 18+.

But the checklist stands on its own. Run it against whatever you build — including a hand-rolled implementation — and you'll ship x402 without the failure modes.


Payload builds small, sharp tools for developers — MCP monetization, API metering, reliability kits. payloadhq.github.io

来源:Google AI:DEV 作者专属(RSS) · dev.to