从遥测到自动事件响应:用 OpenTelemetry、Grafana 与 Codex CLI 为 FastAPI 应用搭建可观测性
From Telemetry to Automated Incident Response: My DevOps & Observability Journey
一位开发者在 DataTalksClub AI Dev Tools Zoomcamp 中,把基于 SQLite 的 FastAPI Order Tracker 改造成可观测系统,用 OpenTelemetry、Prometheus、Loki、Tempo 和 Grafana 采集指标、日志与链路,并针对 5xx 响应配置告警。
Description: A hands-on learning journey building observability and automated incident response into a FastAPI order-tracking application using OpenTelemetry, Prometheus, Loki, Tempo, Grafana, and Codex CLI.
From Telemetry to Automated Incident Response
As part of the DataTalksClub AI Dev Tools Zoomcamp, I worked on a small but practical challenge: turning a simple FastAPI application into an observable system with automated incident response.
The project started as a basic Order Tracker application backed by SQLite.
By the end, it had evolved into a system that could:
- collect metrics, logs, and traces
- visualize telemetry in Grafana
- detect application errors through alerting
- send incidents to an incident-response service
- provide evidence for investigation
- trigger a coding agent to investigate and remediate an issue
The most valuable lesson for me was that observability is not just about collecting data.
It is about connecting detection → investigation → remediation → verification.
The Problem
A basic application can work perfectly during development and still fail in production.
For an order-tracking API, a failure might look simple:
GET /api/orders/express-1002
returns a server error.
Without observability, the workflow often becomes:
Something is broken
↓
Check application logs
↓
Search for the error
↓
Try to reproduce it
↓
Find the problematic code
↓
Apply a fix
↓
Restart the application
↓
Test again
This works, but it becomes increasingly difficult as the system grows.
I wanted to build a workflow where telemetry could help connect these steps.
The Solution
The final architecture introduced several components:
This architecture separates telemetry collection, storage, visualization, alerting, and incident response.
Instrumenting the Application
The first step was adding OpenTelemetry instrumentation to the FastAPI application.
I wanted each order lookup to produce useful telemetry, including:
- HTTP route
- HTTP status code
- request information
- logs
- traces
- request metrics
For example:
GET /api/orders/standard-1001
→ 200
and:
GET /api/orders/standard-1002
→ 404
The important part was not simply knowing that a request happened.
The telemetry needed enough context to answer:
What endpoint failed, what was the HTTP status, and what happened during the request?
Adding the Observability Stack
The next step was connecting the application to an observability stack:
- OpenTelemetry Collector — receives and routes telemetry
- Prometheus — metrics
- Loki — logs
- Tempo — traces
- Grafana — visualization and alerting
The OpenTelemetry Collector became the central telemetry pipeline.
Conceptually:
Application
│
│ OTLP
▼
OpenTelemetry Collector
│
├── Metrics → Prometheus
│
├── Logs → Loki
│
└── Traces → Tempo
Grafana then provides a single interface for exploring these signals.
This was one of the important lessons from the project:
Metrics tell you that something is happening.
Logs help explain what happened.
Traces help show where it happened.
Using the three together provides much more useful operational context than relying on a single signal.
Grafana Alerting
After telemetry was available, I configured Grafana alerting for server-side failures.
The goal was to detect 5xx responses.
A 404 response such as:
GET /api/orders/standard-1002
→ 404
should not trigger the 5xx incident workflow.
This distinction was useful because it forced me to think about alert conditions rather than simply alerting on every failed request.
The alert also needed to provide enough information for the next stage of the workflow.
Automated Incident Response
The next step was adding a dedicated incident-response service.
The service exposes:
POST /alerts
When Grafana sends an alert, the responder stores the incident information and gathers useful evidence.
The idea is to transform:
Grafana Alert
into:
Incident
├── endpoint
├── status
├── logs
├── traces
└── investigation context
This creates a bridge between observability and automated remediation.
Connecting an AI Coding Agent
The most interesting part of the exercise was connecting the incident workflow to a coding agent.
Instead of simply notifying a developer:
🚨 500 error detected
the system can provide the coding agent with context about the incident.
The intended workflow becomes:
5xx detected
↓
Grafana alert
↓
Incident-response service
↓
Collect evidence
↓
Coding agent investigates
↓
Identify root cause
↓
Apply fix
↓
Restart application
↓
Verify behavior
This changes the role of AI from simply generating code to participating in an operational workflow.
Finding the Bug
The final exercise intentionally exposed a bug in the Express order lookup.
The problematic implementation calculated the estimated delivery date using:
estimated_at = placed_at.replace(day=placed_at.day + 2)
This looks reasonable at first glance.
However, it assumes that the resulting day exists in the same month.
For an order created near the end of a month, that assumption can fail.
The fix was to use date arithmetic instead:
estimated_at = placed_at + timedelta(days=2)
This allows Python's datetime implementation to correctly handle month boundaries.
The important part was not only finding the line of code.
The observability and incident-response workflow provided the path from:
HTTP failure
↓
Telemetry
↓
Alert
↓
Incident evidence
↓
Root-cause investigation
↓
Code change
↓
Verification
What I Learned
1. Observability is more than logging
Before this project, it was easy to think of observability as:
"Add some logs."
The project changed that perspective.
A useful observability system combines multiple signals and makes them actionable.
2. Alerts need context
An alert saying:
Something went wrong
is not particularly useful.
An actionable incident should provide enough context to start an investigation.
For example:
Endpoint
HTTP status
Logs
Trace
Time window
Dashboard context
3. Automation should be connected to verification
Automatically changing code is not enough.
A remediation workflow should also verify that the application works after the change.
That makes the workflow closer to:
Detect → Investigate → Fix → Verify
rather than simply:
Detect → Fix
4. AI agents need operational context
A coding agent becomes much more useful when it receives evidence from the system instead of being asked to investigate blindly.
Telemetry can provide the context needed for the agent to understand what actually happened.
Technology Stack
The project brought several tools together:
| Area | Technology |
|---|---|
| API | FastAPI |
| Runtime | Python |
| Database | SQLite |
| Telemetry | OpenTelemetry |
| Telemetry Pipeline | OpenTelemetry Collector |
| Metrics | Prometheus |
| Logs | Loki |
| Traces | Tempo |
| Visualization | Grafana |
| Containers | Docker / Docker Compose |
| AI-assisted remediation | Codex CLI |
Verification
The completed workflow was validated through the homework tasks.
| Step | Scenario | Result |
|---|---|---|
| Q1 | Application health check | 200 |
| Q2 | Existing order lookup | 200 |
| Q3 | Missing order lookup | 404 |
| Q4 | 5xx alert behavior | Normal for 404 |
| Q5 | Incident-response workflow | Completed |
| Q6 | Express order incident | Root cause identified and fixed |
The important outcome was not just that each individual component worked.
The complete workflow worked as a chain.
Final Takeaway
The biggest lesson from this project was that observability becomes much more valuable when it is connected to action.
A mature workflow can look like:
Application
↓
Telemetry
↓
Observability
↓
Alerting
↓
Incident Response
↓
AI-assisted Investigation
↓
Remediation
↓
Verification
This project gave me hands-on experience connecting these pieces together rather than learning them as isolated technologies.
For me, that's the real value of Learning in Public: documenting not only what I built, but also the problems I encountered, the root causes I found, and what changed in my understanding along the way.
Acknowledgments
This project was completed as part of the DataTalksClub AI Dev Tools Zoomcamp — Homework 4: DevOps and Observability for AI-Built Apps.
Thanks to the DataTalksClub community for creating practical exercises that connect software development, DevOps, observability, and AI-assisted engineering.
Github: order-tracker-v2
来源:Google AI:DEV 作者专属(RSS) · dev.to
