跳到正文
原文
Google AI:DEV 作者专属(RSS)· Varun Sharma·· 2 小时前AI 评分62

作者开源只诊断不部署的 AWS DevOps Agent 并分享生产经验

Building an AWS DevOps Agent That Diagnoses but Never Deploys

AI 导读

作者开源了一个 AWS DevOps Agent(github.com/sharma-the-karma/aws-devops-agent),在 GitHub Actions 流水线失败时拉取日志、用 Amazon Bedrock 上的 Claude 诊断并在 PR 上给出建议 diff,但坚持不自动部署。

正文

A pipeline fails, someone gets tagged, and you sit there waiting for ten megabytes of terminal output to render in your browser. Then you scroll.

The moment that pushed me to build this was an ECS deployment pipeline that failed late at night because an AWS Service Control Policy had quietly stripped an update permission. It took me forty minutes of squinting at a browser window full of repetitive Docker build steps just to spot the single AccessDeniedException line buried at line 2,100.

The cause is almost always boring like that. An IAM role is missing an action it newly needs. A Terraform lock is left over from an interrupted run. A Docker cache miss produces a build architecture mismatch. A package registry times out. The fix is usually three lines of policy or a bumped timeout. The real cost is digging through thousands of lines of build noise to find those three lines' worth of signal.

So I built something to do the digging. When a GitHub Actions run fails, an agent pulls the logs, diagnoses the failure with Claude on Amazon Bedrock, and posts a suggested diff on the pull request. It is plain Python, GitHub Actions, AWS OIDC federation, and Bedrock. There is no agent framework, no vector database, and nothing looping unattended against production. The code is open source at github.com/sharma-the-karma/aws-devops-agent.

How it works

It is event-driven and serverless, so it only runs when a pipeline fails. That means no standing infrastructure and no long-lived AWS keys sitting in repository secrets.

flowchart LR
    A[GitHub Actions CI Workflow] -->|Fails| B[Triage Trigger]
    B --> C[Log Distiller Engine]
    C -->|High-Signal Context| D[AWS Bedrock Converse API]
    D -->|Claude| E[Structured Diagnostic Report]
    E --> F[PR Comment with Diff]

    subgraph AWS Cloud
        D
        G[IAM OIDC Role] -.->|Federates| B
    end

A CI job fails on a pull request, which fires a separate triage workflow on the workflow_run completion event. That runner asks GitHub's OIDC provider for a temporary token and uses it to assume an IAM role in AWS. A distiller script downloads the failed step's logs, strips the ANSI color codes, and pulls out the part around the failure. That chunk goes to Bedrock's Converse API, and the diagnosis comes back as a PR comment, which gets edited in place on later failures instead of piling up.

What I learned

Raw logs wreck the model's reasoning

My first version was as naive as it gets: take the full output of the failed step and paste it into the prompt. It went badly in three ways. A normal container build log runs past 10,000 lines, and the real error might sit on line 1,840, so the model was busy commenting on package install warnings instead of the fatal exit code. Pushing 50,000 tokens through a model is also slow when someone is waiting on feedback, and paying to process the same dependency downloads on every failed run adds up.

The fix was a small distiller, src/log_distiller.py. It strips ANSI codes, looks for known failure anchors like AccessDeniedException, tracebacks, Terraform errors, and exit codes, and keeps a window of lines around each hit.

import re
from typing import Dict, Any

ANSI_ESCAPE_PATTERN = re.compile(r"\x1B(?:[@-Z\\-_]|\[[0-?]*[ -/]*[@-~])")

ERROR_INDICATORS = [
    re.compile(r"error[:\s]", re.IGNORECASE),
    re.compile(r"traceback \(most recent call last\):", re.IGNORECASE),
    re.compile(r"failed with exit code \d+", re.IGNORECASE),
    re.compile(r"accessdeniedException", re.IGNORECASE),
    re.compile(r"terraform error[:\s]", re.IGNORECASE),
]

def extract_failure_context(raw_log: str, window_before: int = 15, window_after: int = 35) -> Dict[str, Any]:
    clean_lines = ANSI_ESCAPE_PATTERN.sub("", raw_log).splitlines()
    selected_indices = set()

    for idx, line in enumerate(clean_lines):
        for pattern in ERROR_INDICATORS:
            if pattern.search(line):
                start = max(0, idx - window_before)
                end = min(len(clean_lines), idx + window_after + 1)
                selected_indices.update(range(start, end))
                break

    sorted_indices = sorted(selected_indices)
    extracted = [f"L{i+1:04d}: {clean_lines[i]}" for i in sorted_indices]

    return {
        "distilled_log": "\n".join(extracted),
        "total_lines_analyzed": len(clean_lines),
        "lines_extracted": len(extracted)
    }

The difference was dramatic. On one multi-stage Docker build failure, the raw step log ran past 12,000 lines—roughly 48,000 tokens. Slicing with the anchor window brought it down to 580 tokens. That dropped inference latency from nearly 15 seconds to under two, and the model was finally looking at the actual error instead of hallucinating about npm peer dependency warnings. It is still crude, since the error[:\s] pattern is broad and will occasionally grab harmless lines, but it is far better than sending everything.

Don't give the agent permanent keys

Never create long-lived AWS access keys (AKIA...) for a CI pipeline or an agent. If someone opens a pull request that changes the workflow file or runs untrusted code, those keys are exposed.

Use an IAM OIDC identity provider for GitHub instead. The workflow gets a short-lived JWT from GitHub and trades it with AWS STS for temporary credentials. The trust policy below limits the role to this repository:

data "aws_iam_policy_document" "github_oidc_trust" {
  statement {
    effect  = "Allow"
    actions = ["sts:AssumeRoleWithWebIdentity"]

    principals {
      type        = "Federated"
      identifiers = [aws_iam_openid_connect_provider.github.arn]
    }

    condition {
      test     = "StringEquals"
      variable = "token.actions.githubusercontent.com:aud"
      values   = ["sts.amazonaws.com"]
    }

    condition {
      test     = "StringLike"
      variable = "token.actions.githubusercontent.com:sub"
      values   = ["repo:sharma-the-karma/aws-devops-agent:*"]
    }
  }
}

The wildcard on sub allows any ref in the repo, so in a real setup I would tighten it to the default branch (repo:sharma-the-karma/aws-devops-agent:ref:refs/heads/main), especially since workflow_run executes in the context of the default branch. The permissions on the role are deliberately tiny: bedrock:InvokeModel and nothing else that matters. The agent cannot write to S3, change security groups, or deploy anything.

Use the Converse API

Bedrock gives you InvokeModel and the newer Converse. With InvokeModel you send vendor-specific request bodies, so switching from Claude to another model means rewriting your payload. Converse uses one schema across models and handles system prompts, tool definitions, and token limits cleanly.

import boto3

SYSTEM_PROMPT = """You are an AWS DevOps and Infrastructure Architect.
Inspect failing CI/CD logs, diagnostic traces, and infrastructure errors.
Provide a concise, practical triage assessment with actionable solutions.

Structure your response into:
1. Root Cause Breakdown (what failed and its category)
2. Immediate Fix / Patch (exact code snippet or command)
3. Preventive Action & Blast Radius (how to prevent it and potential side-effects)
"""

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="anthropic.claude-3-5-sonnet-20241022-v2:0",
    messages=[
        {
            "role": "user",
            "content": [{"text": f"CI/CD Failure Trace:\n{distilled_log}"}]
        }
    ],
    system=[{"text": SYSTEM_PROMPT}],
    inferenceConfig={
        "maxTokens": 2048,
        "temperature": 0.2,
        "topP": 0.9,
    }
)

diagnostic_report = response["output"]["message"]["content"][0]["text"]

I used anthropic.claude-3-5-sonnet-20241022-v2:0 here for depth of reasoning on complex Terraform graphs, but for routine CI triage where you just need fast stack trace extraction, anthropic.claude-3-5-haiku-20241022-v1:0 is three times faster, significantly cheaper, and more than capable.

I keep the temperature at 0.2. For stack trace analysis you want consistent, boring answers, not creative ones.

Update the PR comment instead of adding a new one

If an engineer pushes four commits in a row while fixing a broken build, nobody wants four bot comments on the PR. They get ignored. So I put an HTML comment tag in the markdown header and edit the existing comment when it is there:

COMMENT_TAG = "<!-- aws-devops-agent-triage -->"

def post_or_update_pr_triage_comment(repo, pr_number: int, analysis_markdown: str):
    body = f"{COMMENT_TAG}\n### AWS DevOps Agent Triage Report\n\n{analysis_markdown}"
    pull_request = repo.get_pull(pr_number)

    for comment in pull_request.get_issue_comments():
        if COMMENT_TAG in comment.body:
            comment.edit(body)
            return "Updated existing triage comment."

    pull_request.create_issue_comment(body)
    return "Created new triage comment."

The comment needs pull-requests: write on the workflow's GITHUB_TOKEN. After that, each new failure quietly rewrites the same comment.

Never let the agent apply infrastructure changes

There is a lot of excitement about agents that write code, commit it, and apply cloud changes with no human in the loop. I think that is a bad idea for infrastructure, and here is the kind of case that makes me say so: a Terraform run fails because an ECS service cannot attach to a subnet. The agent correctly spots the subnet misconfiguration and proposes a replacement. But in Terraform, changing certain networking attributes forces the resource to be destroyed and recreated. If the agent applies that fix on its own, it can take down a live service or a database.

So the line is deliberate. The agent investigates, isolates the error, and suggests a diff. A human reads it, checks the blast radius, and decides whether to deploy. The point is to remove the tedious log-reading, not to remove engineering judgment.

What a report looks like

Here is the output on a PR where an ECS service update failed because of a missing IAM permission:

AWS DevOps Agent Triage Report

1. Root Cause Breakdown
The deployment failed with AccessDeniedException while calling ecs:UpdateServicePrimaryTaskSet. The execution role assumed by the CI runner lacks this action in its attached policy. This is an infrastructure / IAM permissions issue.

2. Immediate Fix / Patch
Add the missing action to your CI deployer policy in infra/iam_deployer.tf:

 statement {
   actions = [
     "ecs:UpdateService",
     "ecs:DescribeServices",
+    "ecs:UpdateServicePrimaryTaskSet"
   ]
   resources = [aws_ecs_service.api_service.arn]
 }

3. Preventive Action & Blast Radius
Review deployment policies whenever you introduce CodeDeploy or blue/green task sets on ECS services. Applying this change is non-destructive and requires no resource recreation.

Try it yourself

Clone the repo:

git clone https://github.com/sharma-the-karma/aws-devops-agent.git
cd aws-devops-agent

Deploy the IAM role:

cd infra
terraform init
terraform apply -var="github_org=sharma-the-karma" -var="github_repo=aws-devops-agent"

Add the role ARN from the output as AWS_ROLE_ARN in your repository secrets (Settings > Secrets and variables > Actions).

To test locally against the sample log without calling anything for real:

cd src
pip install -r requirements.txt
python triage_agent.py --log-file sample_failure.log --dry-run

Closing thoughts

The place AI earns its keep in DevOps, in my experience, is a narrow one: shortening the time between "something broke" and "I understand why". Trim the logs, use short-lived credentials, and keep a human on the apply button, and you get that without much risk or machinery.

The code is at github.com/sharma-the-karma/aws-devops-agent. Adapt it to your own pipelines however you like.

来源:Google AI:DEV 作者专属(RSS) · dev.to