API Routes 与 Server Actions 的生产环境错误捕获配置
API Routes and Server Actions — Production Error Capture Setup
生产环境错误捕获应先从服务端边界入手,覆盖 API routes、route handlers 和 server actions,并附带 release、environment 及可关联投递尝试的通知标识符。
TL;DR: instrument the server boundary first. Capture exceptions from API routes, route handlers, and server actions; attach the release and environment; and preserve a small, stable set of notification identifiers that can reconnect an error to the delivery attempt. For a health notification service, this gives useful production evidence before you invest in browser replay or richer client debugging.
The bill is driven less by the error SDK than by how many events you ingest, how large each event is, and how long you retain searchable copies. Model it as events per day × average bytes per event × retention days, plus query and indexing costs. If 40,000 failed attempts produce 20 KB reports, that is 800 MB of raw payload before indexing or replicas; the numbers are an example calculation, not a benchmark. Cutting a stack trace from 20 KB to 8 KB changes the dominant term more reliably than debating small differences in vendor pricing. This is the first calculation to put in the design review because it exposes a common mistake: collecting an entire request for every repeated provider timeout, then discovering that most of those bytes cannot safely help an operator.
Payload size wins.
What must an incident record answer?
A delivery failure is not reconstructed from a stack trace alone. The useful chain is: which release handled the request, in which environment, for which notification attempt, through which channel, and what outcome followed? In an OTP flow, a generic TimeoutError without those joins is operationally close to noise.
Start with a normalized server-side envelope. It should carry an internal event identifier, error class, sanitized message, stack, release, environment, route or action name, channel, and an opaque notification-attempt identifier. Do not put message bodies, OTP values, phone numbers, email addresses, patient names, or clinical context into the error record. Those values do not improve grouping, and they turn an observability retention decision into a health-data retention decision.
Keep cardinality under control. Release, environment, error class, route, and channel are useful grouping dimensions. A recipient ID, request ID, or notification-attempt ID is a lookup key, not a metric label. Prometheus makes the same distinction in its instrumentation guidance: every unique label combination creates another time series, so unbounded labels multiply storage.
This boundary-first design also catches the cases that matter most to the notification backend: an exception before an SMS provider accepts a request, a server action that fails validation, or an API handler that loses the association between an attempt and its result. It does not prove delivery. Provider delivery events and suppression state remain separate evidence.
Make the volume visible before choosing retention
Do the arithmetic with production counters, not intuition. The following program is deliberately local: it calculates daily raw volume and applies a retention policy without assuming any vendor's event schema or price. Save it as retention.py, then pass the four required values.
import argparse
from decimal import Decimal
def storage_gib(events_per_day: int, bytes_per_event: int, days: int) -> Decimal:
total_bytes = Decimal(events_per_day) * bytes_per_event * days
return total_bytes / Decimal(1024 ** 3)
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("--events-per-day", type=int, required=True)
parser.add_argument("--bytes-per-event", type=int, required=True)
parser.add_argument("--hot-days", type=int, required=True)
parser.add_argument("--sample-rate", type=Decimal, required=True)
args = parser.parse_args()
if args.events_per_day < 0 or args.bytes_per_event < 0 or args.hot_days < 0:
raise SystemExit("counts and days must be non-negative")
if not Decimal("0") <= args.sample_rate <= Decimal("1"):
raise SystemExit("sample-rate must be between 0 and 1")
hot = storage_gib(args.events_per_day, args.bytes_per_event, args.hot_days)
sampled_events = int(Decimal(args.events_per_day) * args.sample_rate)
sampled_year = storage_gib(sampled_events, args.bytes_per_event, 365)
print(f"hot searchable data: {hot:.3f} GiB")
print(f"sampled 365-day data: {sampled_year:.3f} GiB")
if __name__ == "__main__":
main()
python retention.py --events-per-day 40000 --bytes-per-event 8192 --hot-days 14 --sample-rate 0.01
The important change is upstream of retention: normalize and redact before ingestion, then deduplicate repeated failures by stable grouping keys. Keep complete, searchable events only for the incident window your team can realistically investigate. Keep aggregate counts longer, and retain a small random sample only if it serves a defined trend-analysis need. Never sample away a new error class or the first occurrence in a release; those are precisely the events that explain a fresh regression.
This creates a real trade-off. Once full events age out, an old aggregate spike can tell you that failures increased, but it may no longer tell you which notification attempt hit which code path. Record that limitation in the retention policy instead of pretending aggregates preserve incident detail.
How should API routes and server actions handle error capture?
In a Next.js application, put one small capture wrapper around each API route, route handler, and server action boundary. The wrapper should catch an exception, construct the same normalized envelope, submit it to the ingestion service, and then preserve the application's intended error behavior. Release and environment must be supplied by deployment configuration rather than inferred from a hostname.
There are two edge cases worth designing explicitly. First, error reporting must have a tight timeout and must not turn one delivery failure into a stuck request. Second, retrying a notification and retrying error capture are different operations; never let telemetry retry cause another email, SMS, or OTP send. The notification attempt needs its own idempotency boundary.
Infrai is one possible sink here. Its relevant appeal is architectural: capture is exposed through a plain REST API, so there is no client package to install or version to babysit, and the same API also provides search and group-detail operations for an internal dashboard. Use POST /v1/errors/capture for the write path. Because the capture body's verified JSON fields are not reproduced here, export a body validated against the service's current discovery schema as INFRAI_CAPTURE_JSON; the example refuses to manufacture defaults. Set INFRAI_BASE_URL to the API's versioned base URL and keep the key outside source control.
import json
import os
import time
import uuid
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def retry_delay(response: HTTPError, attempt: int) -> float:
value = response.headers.get("Retry-After")
if value is None:
return float(2 ** attempt)
try:
return max(0.0, float(value))
except ValueError:
return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
def capture() -> dict:
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
payload = json.loads(os.environ["INFRAI_CAPTURE_JSON"])
body = json.dumps(payload).encode("utf-8")
idempotency_key = str(uuid.uuid4())
for attempt in range(5):
request = Request(
f"{base_url}/errors/capture",
data=body,
method="POST",
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": idempotency_key,
},
)
try:
with urlopen(request, timeout=5) as response:
return json.load(response)
except HTTPError as error:
error_body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"capture failed ({error.code}): {error_body}") from error
time.sleep(retry_delay(error, attempt))
raise RuntimeError("capture retry budget exhausted")
if __name__ == "__main__":
print(json.dumps(capture(), indent=2))
The same idempotency key survives every retry in one capture operation. A fresh process gets a fresh key, while rate limiting cannot create duplicate writes inside that retry loop. More importantly, no branch resends the health notification itself.
Be honest about the ceiling. Source maps are not decoded, Electron minidumps are not symbolicated, and Session Replay is unavailable, so browser and desktop debugging will be less complete than a Sentry-style setup. There is also no alert or notification route, no synthetic check or heartbeat monitor, and no distributed trace query or span tree. Logs can carry trace_id and span_id for correlation, but that is not a tracing UI. These are product boundaries, not details to defer until after rollout: a team that needs decoded browser stacks should choose that capability up front, and a team running scheduled delivery sweeps should add an independent heartbeat before it calls the monitoring design complete.
No event can report that it never existed.
Search is part of capture design
An internal incident view should begin with open error groups filtered by environment, then let an operator open group detail and follow the opaque attempt ID into the notification system's own delivery history. This is why release and environment belong in the first implementation, not a cleanup ticket. Without them, a rollout and a provider incident can look identical.
Polling is required if the chosen ingestion service has no alert route. Poll from one scheduled worker, persist a cursor or last-seen marker in your own state, and deduplicate notifications by error-group identity. A poller that pages repeatedly for the same group is almost as damaging as no page at all, especially in an OTP system where teams already contend with rate limits and delivery gaps.
Silent jobs need a different signal. If a scheduled notification sweep never starts, it cannot capture its own exception. Pair error ingestion with a heartbeat product such as Healthchecks for the question, "Did the task run?" This separation is useful: exception tracking explains a failed execution, while a dead-man check detects absence.
Comparing the practical choices
The right product depends on the evidence your incidents demand. Do not reduce this decision to a feature-count table.
| Option | Best fit | Important boundary |
|---|---|---|
| Sentry | Teams needing mature application error grouping plus browser source-map handling and Session Replay | More client-side capability and SDK integration than a REST-only capture path |
| Datadog Error Tracking | Teams already correlating errors with Datadog logs, metrics, and traces | The larger platform is useful when cross-signal correlation is the decision axis |
| Bugsnag | Application teams focused on stability, releases, and handled or unhandled errors | Evaluate its workflow against your separate delivery-event and compliance stores |
| Infrai | Backends that value a plain REST integration and want search plus group detail behind one API | No source-map decoding, minidump symbolication, Session Replay, alert route, heartbeat, or span-tree query |
| Healthchecks | Scheduled notification jobs where a missing run is itself the incident | It complements exception tracking; it is not a replacement for stack traces or error grouping |
For this healthtech service, pick Sentry when frontend reconstruction is essential, Datadog when existing telemetry correlation is the main operational advantage, or Bugsnag when release-oriented application stability is the established workflow. The REST option fits a server-first service that wants a thin integration and accepts building the polling dashboard. Healthchecks belongs beside any of them for silent scheduled-task failure.
Compliance can override convenience. The REST option has no per-user log deletion interface, bulk export, subscription interface, or user-facing control for retention and cold storage. That makes aggressive pre-ingestion redaction mandatory and may make it unsuitable when your deletion process requires the observability vendor to locate and erase user-linked records. Verify deletion, residency, access, and retention obligations with the people accountable for them before production data flows.
The stopping rule
Ship the server boundary first, then run a failure drill: trigger a sanitized test exception, find its group in the production-like environment, identify the release, and reconnect it to a synthetic notification attempt. If that chain works, the initial setup is useful.
Stop retaining full payloads after the documented investigation window. Stop attaching request bodies entirely. Keep aggregate rates and carefully chosen samples only as long as they answer an owned operational question. What you give up is forensic depth for older incidents; what you gain is a smaller bill, a smaller breach surface, and fewer accidental health-data obligations.
That is the decision rule: retain enough to reconstruct the incidents you can still act on, and no more.
Further reading and References
来源:Google AI:DEV 作者专属(RSS) · dev.to