跳到正文
原文
Google AI:DEV 作者专属(RSS)· Lars Winstand·· 7 小时前AI 评分60

GPT-5.4 智能体在生产环境变懒?排查后发现是 3 个工作流 bug 在教它提前放弃

We thought our GPT-5.4 agent got lazier in production — it was a 3-bug workflow teaching it to quit

AI 导读

作者的 n8n 智能体在 staging 表现正常,上线后 GPT-5.4 输出变短、工具调用减少、答案更自信但错误,最初怀疑模型退化,实际是 3 个工作流 bug:研究类任务重试上限从 6 降到 2、一个分支把符合 schema 的部分输出当成功、OpenAI 兼容 API 路径奖励首个可接受答案而非最佳答案。

正文

We thought our GPT-5.4 agent got lazier in production — it was a 3-bug workflow teaching it to quit

We had an n8n agent that looked great in staging.

It would:

  • search
  • pull docs
  • compare sources
  • verify claims
  • write a grounded answer

Then we shipped it.

In production, the same task started ending after one shallow pass.

Symptoms were exactly what people usually call “model laziness”:

  • shorter outputs
  • fewer tool calls
  • more confident wrong answers
  • less evidence of multi-step reasoning

Our first instinct was to blame GPT-5.4.

That was the wrong diagnosis.

The real issue was boring and very fixable:

  1. retry cap dropped from 6 to 2
  2. one n8n branch treated a partial answer as success
  3. our OpenAI-compatible API path was rewarding the first acceptable-looking response instead of the best one

Once we fixed the workflow, quality came back.

That changed how I think about “lazy” agents in production.

Most of the time, the model did not suddenly get worse. Your orchestration started ending runs early.

The production symptom that fooled us

In staging, the agent trace looked like this:

  1. retrieve sources
  2. inspect docs
  3. compare claims
  4. verify one weak point
  5. synthesize answer
  6. return final output

In production, it looked more like this:

  1. retrieve one weak source
  2. write answer anyway

Same task class. Same model family. Very different behavior.

And because the final answer still looked polished, it passed casual review more often than it should have.

That is the dangerous part.

A broken agent rarely looks broken in an obvious way. It often looks efficient.

The 3 bugs that made GPT-5.4 look lazy

1) Retry cap dropped from 6 to 2

This was the biggest quality hit.

For simple classification, 2 retries can be fine.

For research, debugging, document synthesis, or anything with tool use, 2 is often a trap. One bad retrieval result plus one tool hiccup and the agent is out of budget.

Example of the kind of config drift that causes this:

{
  "task_type": "research",
  "max_retries": 2,
  "timeout_seconds": 20
}

That looks harmless until your workflow depends on search + fetch + verify.

2) A branch treated partial output as success

In our n8n flow, one formatting branch said, effectively:

  • if output matches schema
  • and output length is above a minimum
  • mark run successful

That means the agent could skip retrieval depth and still win.

This is how you accidentally train an agent to stop early.

Pseudo-logic:

const passed =
  isValidJson(response) &&
  response.answer.length > 280;

if (passed) {
  return "success";
}

That is not quality control.

That is a shallow-answer reward function.

3) The API path rewarded “acceptable” over “best”

This one is common in OpenAI-compatible stacks.

If your app accepts the first plausible answer and never checks whether the expected tool path ran, the orchestration layer starts selecting for speed, not depth.

That can happen whether you are routing to GPT-5.4, Claude Opus 4.6, or Grok 4.20.

The model is not “choosing to be lazy” in some abstract sense.

Your workflow is telling it:

if you look done quickly enough, you pass

Why agents look worse in production than in staging

Because production has constraints that staging often hides.

In a clean test harness, a model usually gets:

  • one prompt
  • predictable context
  • generous timeout
  • no weird branch logic
  • no flaky tools

In production, the agent sits inside a box made of:

  • retry limits
  • timeout settings
  • parser requirements
  • tool wrappers
  • success conditions
  • fallback branches
  • queue pressure
  • rate limiting

That box matters more than people want to admit.

A strong model inside a bad loop will look worse than a decent model inside a clean loop.

The fastest way to tell if the model is actually the problem

Run the same task through the same scaffold and change one variable at a time.

Not “same prompt, different environment.”

Actually the same scaffold:

  • same system prompt
  • same tool definitions
  • same retry budget
  • same max output settings
  • same evaluator
  • same API path
  • same success criteria

If you compare production n8n against a clean notebook script, you are not isolating the model.

You are changing the entire experiment.

The debugging signal that mattered most: stop reason

Not vibes.

Not output length alone.

Not “this answer feels thinner.”

Stop reasons told us far more than final-answer scoring.

For Anthropic agents, useful stop reasons include values like:

  • end_turn
  • max_tokens
  • tool_use
  • pause_turn

For OpenAI-compatible workflows, inspect whether:

  • the expected tool calls actually fired
  • the reasoning path was used when expected
  • the run ended because the model finished
  • or because your orchestration layer decided it had enough

If you only evaluate the final answer, you are debugging blind.

What convinced us it was the workflow, not GPT-5.4

We ran the same broken production scaffold against multiple model families.

What we saw:

  • GPT-5.4 looked bad under the production n8n flow
  • GPT-5.4 looked fine under the staging flow
  • Claude Opus 4.6 also looked bad under the broken production flow
  • Grok 4.20 looked bad too

That pattern matters.

When three strong models all become “lazy” in the same way, the workflow is usually guilty.

Here is the mental model I use now:

If this changes Suspect
One model regresses, others stay stable model or provider issue
All models regress under one workflow orchestration bug
Output gets shorter after retry/timeout changes early stopping
JSON validity improves while answer quality drops parser-first reward problem

The false leads we chased first

We spent too long blaming the model layer.

Our guesses were reasonable:

  • maybe GPT-5.4 regressed
  • maybe the OpenAI Responses API settings changed behavior
  • maybe reasoning effort was too low
  • maybe output token limits were clipping the answer

Those are all real failure modes.

They just were not the main problem here.

The actual issue was simpler:

staging rewarded grounded completion
production rewarded acceptable formatting

Agents optimize for whatever your workflow rewards.

If your automation says “close enough,” GPT-5.4, Claude Opus 4.6, and Grok 4.20 will all start looking suspiciously eager to be done.

Workflow patterns that create fake laziness

These are the ones I would audit first.

Parser-first design

If the main objective is valid JSON, many agents will satisfy the parser before they satisfy the task.

Example smell:

if (schema.safeParse(output).success) {
  return success;
}

That should almost never be the whole success condition for a research task.

Direct-answer escape hatches

If n8n, Make, Zapier, OpenClaw, LangGraph, or your custom loop allows a final answer before retrieval or verification, expect shallow completions.

For research-class tasks, tool use often should not be optional.

Tiny retry budgets on multi-step tasks

This one is everywhere.

People use one global retry budget for everything:

max_retries: 2

That might be fine for:

  • classification
  • extraction
  • light transformations

It is usually bad for:

  • web research
  • code debugging
  • document comparison
  • multi-source synthesis

Evaluating only the final message

This is the worst one.

If you only score the final text, you hide:

  • skipped tool calls
  • failed searches
  • parser shortcuts
  • premature exits
  • timeout-driven summaries pretending to be conclusions

My strong opinion: this single habit causes teams to think they are comparing models when they are actually comparing orchestration mistakes.

What we changed

We did not switch models.

We changed the workflow.

Fix 1: restore retry budget

We moved the retry cap back from 2 to 6 for research-class tasks.

agent_profiles:
  classification:
    max_retries: 2
  research:
    max_retries: 6
  debugging:
    max_retries: 6

Fix 2: require retrieval for research tasks

We added a hard gate.

If the task is research, at least one retrieval step must happen before a run can pass.

Pseudo-code:

function validateRun(run) {
  if (run.taskType === "research" && run.toolCalls.search < 1) {
    return { ok: false, reason: "missing_required_retrieval" };
  }

  if (!run.outputSchemaValid) {
    return { ok: false, reason: "invalid_schema" };
  }

  return { ok: true };
}

Fix 3: stop treating formatting as quality

We changed success conditions so a run could not pass on formatting alone.

That meant checking trajectory, not just output shape.

Fix 4: compare traces, not just answers

We reviewed traces side by side in our observability stack and checked stop reasons across both the OpenAI-compatible path and the Anthropic path.

That made the difference obvious fast.

A practical checklist for debugging “lazy” agents

If you think your production agent got worse, this is the order I would check things in.

1) Diff staging vs production config

diff staging-agent.yaml production-agent.yaml

Look for changes in:

  • retry caps
  • timeouts
  • token/output limits
  • tool availability
  • branch logic
  • fallback behavior
  • evaluator thresholds

2) Inspect stop reasons

Log them explicitly.

{
  "run_id": "abc123",
  "model": "gpt-5.4",
  "stop_reason": "end_turn",
  "tool_calls": 0,
  "task_type": "research"
}

If research tasks are ending with zero tool calls and still passing, that is your bug.

3) Count tool calls per task type

A simple metric catches a lot:

SELECT
  task_type,
  AVG(tool_call_count) AS avg_tool_calls,
  AVG(retry_count) AS avg_retries,
  AVG(output_chars) AS avg_output_chars
FROM agent_runs
WHERE created_at >= NOW() - INTERVAL '7 days'
GROUP BY task_type;

If tool-call counts collapse after a deploy, investigate the workflow before blaming the model.

4) Compare multiple model families under the same scaffold

This is where an OpenAI-compatible API setup helps.

If GPT-5.4, Claude Opus 4.6, and Grok 4.20 all fail the same way under one loop, the loop is probably broken.

5) Score trajectory, not just answer text

Track things like:

  • retrieval attempted
  • verification attempted
  • tool errors encountered
  • stop reason
  • retries used
  • output grounded in cited material

Why this matters more when you run lots of agents

This kind of bug gets expensive fast when you are running automations all day.

Not just in dollars. In bad outputs, hidden regressions, and wasted debugging time.

Teams running agents in n8n, Make, Zapier, OpenClaw, or custom OpenAI-compatible stacks usually hit the same wall:

they start by asking “which model is best?”

Then eventually they realize the more useful question is:

“what exactly is our workflow rewarding?”

That is also why predictable API infrastructure matters.

When you can swap models without rewriting your stack, compare traces cleanly, and run lots of evals without per-token anxiety, it gets much easier to find orchestration bugs instead of arguing about vibes.

That is a big part of why Standard Compute is interesting for agent teams: it is a drop-in OpenAI-compatible API, so you can keep your existing SDKs and workflows, route across GPT-5.4, Claude Opus 4.6, and Grok 4.20, and test agent behavior without every debugging session turning into a billing event.

For teams running automations 24/7, flat monthly pricing is not just a finance preference. It changes how aggressively you can evaluate, compare, and fix agent systems.

The rule I use now

Before blaming GPT-5.4 for getting lazy:

  • inspect the trajectory
  • check stop reasons
  • compare the exact same scaffold
  • verify required tool steps happened
  • make sure your success condition is rewarding quality, not just completion

Sometimes a model really does regress.

More often, production taught your agent that quitting early is the winning move.

来源:Google AI:DEV 作者专属(RSS) · dev.to