跳到正文
原文
Google AI:DEV 作者专属(RSS)· Gabriele Pieretti·· 4 小时前AI 评分36

Laravel 13.33/13.34 如何应对 Docker 中队列任务超时与崩溃

Slow Jobs, Killer Containers: Laravel Queue Timeouts and Crashes in Docker

AI 导读

Laravel 13.34.0 新增 CountCrashesAsExceptions 属性,让被 OOM 杀死或部署中断的队列任务也计入 maxExceptions,避免任务在 retryUntil 到期前无限重试并重复调用付费模型。

正文

Gabriele Pieretti

A job that sends an email takes half a second. A job that calls a generative model takes as long as the provider feels like, and sometimes a lot longer. Run that job in a Docker container and it has three ways to die: the worker timeout fires, the kernel kills it for memory, or a deploy pulls the container out from under it. Until this month Laravel treated those three deaths differently and didn't tell you. Laravel 13.33 and 13.34 change two of them. Here's which two, and what's still on you.

One job, three ways to die

Picture a job that sends an image to a model and waits for the answer, running under php artisan queue:work inside a container. In production it can end like this:

  • Worker timeout. The job exceeds --timeout (or the job's $timeout). Laravel uses pcntl and SIGALRM; when the alarm fires, the worker has always killed itself with SIGKILL. The timeout is recorded against the job (which fails immediately if it has $failOnTimeout) and the process is gone.
  • Container OOM. The process goes over the container's memory limit (mem_limit or deploy.resources.limits.memory in Compose). The kernel sends SIGKILL: no exception, no failed() callback, no application log. If php is PID 1, the container exits with 137.
  • Deploy. docker compose up recreates the container. It sends SIGTERM, the worker tries to finish its current job, and after stop_grace_period — ten seconds by default — comes SIGKILL. For a job waiting on a model, ten seconds isn't much.

In the last two cases the job stays reserved until retry_after expires, then becomes available again and another worker picks it up. As far as the framework is concerned nothing happened: no exception was thrown, so no exception was counted.

The gap: maxExceptions can't see crashes

With $tries the damage is bounded: every time the job is popped the attempt counter goes up, and eventually it fails. But if you call rate-limited external APIs you often can't use $tries, because the RateLimited middleware releases the job and burns an attempt even when the job never actually ran. So the usual setup becomes $tries = 0, a $maxExceptions cap, and retryUntil() as a safety net.

The problem is that maxExceptions only counts real exceptions. A job that pushes the worker into OOM, gets killed, returns to the queue and pushes the next worker into OOM increments nothing. It loops until retryUntil() runs out, burning CPU and — if it calls a paid model — billed requests. The author of PR #61737 describes hitting exactly this with a worker being OOM-killed on Laravel Cloud.

The fix is opt-in, per job:

#[MaxExceptions(3), CountCrashesAsExceptions]
class GeneratePreview implements ShouldQueue
{
    public $tries = 0;

    public function retryUntil(): DateTime
    {
        return now()->addMinutes(30);
    }
}

The mechanism is simple and honest. When the worker picks up the job it writes a marker to the cache, and deletes it when the attempt ends. If the marker is still there on the next attempt, the previous one died badly, and that death counts as one exception. It costs two cache calls per attempt, and it needs a cache shared between workers: with the array driver, or file inside separate containers, the marker dies with the container. It shipped in 13.34.0; instead of the attribute you can also set public $countCrashesAsExceptions = true;.

A timeout that doesn't kill the worker

PR #61591, shipped in 13.33, adds a static flag:

// AppServiceProvider::boot()
Worker::$killOnTimeout = false;

With it off, the worker doesn't kill itself when the timeout fires: it throws a TimeoutExceededException inside the job. The job gets a chance to clean up, close what it opened, even catch the exception and decide what to do. The worker survives and moves on to the next job without re-bootstrapping the framework and reopening every connection.

I'd be careful with it, for two reasons spelled out in the discussion on PR #61622. First, an exception thrown from a signal handler only propagates once PHP gets back to executing instructions; a job stuck inside a blocking system call never gets there. Second, if the exception does unwind, the worker carries on with whatever state the interrupted job left behind. SIGKILL is brutal, but it guarantees that state dies with the process. And a broad try/catch (\Throwable) inside handle() will swallow the exception and make the timeout disappear altogether — which is why the default stays true.

Line up your timeouts, inside out

Neither flag replaces the thing that actually matters: a strict ordering of timeouts, from the innermost to the outermost. In years of running Laravel queues this is the part I've seen get wrong most often, and rarely out of ignorance — the four numbers just live in four different files and nobody reads them side by side.

  1. The HTTP timeout on the model call comes first. Http::timeout(60) or your SDK's equivalent. It's the only one that turns waiting into a normal, catchable, counted exception.
  2. The job's $timeout, a few seconds higher. It only fires if the first one wasn't enough.
  3. retry_after on the queue connection, above the job timeout. If it's lower, a second worker picks the job up while the first is still running it. The docs say so; with a paid model it means paying twice for the same answer.
  4. stop_grace_period on the Compose worker service, above the job timeout, so a deploy waits for the running job instead of killing it.
services:
  worker:
    image: app:latest
    command: php artisan queue:work --timeout=90 --memory=256
    stop_grace_period: 120s
    deploy:
      resources:
        limits:
          memory: 512M

A note on --memory: it stops the worker cleanly once memory goes over the threshold, but it checks between jobs. It won't save you from a single job that balloons past the container limit while decoding an image. For that you need a container limit with headroom, and — now — CountCrashesAsExceptions so you don't end up in a loop.

Your payload is a file, even if you don't call it one

One last point, from Miraviso, the SaaS I'm building for hair salons. There, the haircut preview is generated by Gemini on an EU server, with consent, and is never written to disk. That stack is FastAPI, not Laravel, but the rule translates word for word: put an image in a job payload and you've written it to your queue backend — Redis with persistence, a jobs table, and on failure failed_jobs. Retries read it back from there. I wrote about these undeclared copies in a piece on sensitive data in logs; queues follow the same logic. If the promise is "never on disk", an async job is the wrong home for that data, and automatic retries need rethinking accordingly.

The rest is hygiene: a PR that counts crashes, a flag that makes timeouts manageable, and four numbers in the right order. Of those, the four numbers are the only thing no upgrade will ever do for you.


Originally published at gabrielepieretti.dev. I write about Laravel, privacy-by-design and building a vertical B2B SaaS solo — more here.

来源:Google AI:DEV 作者专属(RSS) · dev.to