跳到正文
原文
Google AI:DEV 作者专属(RSS)· Christian Anderson·· 3 小时前AI 评分43

家庭实验室盘点:41 个容器、一块 6 GB GPU,以及我的 AI 智能体都跑在哪

Homelab census: 41 containers, one 6 GB GPU, and where my AI agents run

AI 导读

一位开发者盘点其家庭实验室:四台机器上共 41 个 Docker 容器和 9 个 LXC 容器,全部运行在 26 个 CPU 线程和 71 GB 内存上,AI 路由由一块 6 GB 显存的 RTX 2060 决定。

正文

Today I asked my LLM box what it was doing, and it told me it had a 27-billion-parameter model loaded. Total footprint 18.3 GB. Amount of that on the graphics card: 0.5 GB.

So the 27B was "running on the GPU" in the same sense that I am running a marathon when I walk to the shop. The other 17.8 GB sat in system RAM and did its arithmetic on the CPU, one patient token at a time.

That seemed like a good moment for a proper census: what is actually running, on what, and which AI jobs go to local hardware, a paid cloud API or a flat subscription. The short version: 41 Docker containers and 9 LXC containers across four boxes, and a single 6 GB card makes most of the routing decisions for me.

(Background, already written up: the small-company org chart, the six-job agent team and why the coding agent moved to OpenCode and OpenRouter.)

The census

Measured on 24 September.

Box Hardware What it runs
Proxmox node 1 4 cores, ~8 GB RAM 5 LXC containers: three single-purpose service boxes, a Docker host (12 containers) and a Docker "proxy node" (4 containers)
Proxmox node 2 4 cores, 16 GB RAM 4 LXC containers: the Hermes agent box (native, no Docker), Obico (4 containers), Vaultwarden (1), Langfuse (6)
ZimaBlade 2 cores, 16 GB RAM, fanless 14 Docker containers on ZimaOS
LLM box i7-10875H (16 threads), 31 GB RAM, RTX 2060 with 6 GB VRAM Ollama with 21 model tags, and ComfyUI sharing the same GPU. No containers.

The arithmetic, since I got it wrong the first time: node 1 has 12 + 4 = 16 Docker containers, node 2 has 4 + 1 + 6 = 11, so 27 on the Proxmox side, plus 14 on the ZimaBlade. That's 41 Docker containers and 9 LXC containers, plus three native services: Ollama, ComfyUI and Hermes itself.

All of that sits on 26 CPU threads and 71 GB of RAM, if you add up four boxes that were never designed to be added up.

What's in the Docker hosts

The node 1 Docker host is the toolbox: crawl4ai for page fetching, SearXNG for search, ESPHome, Wyoming Whisper and Piper for speech-to-text and text-to-speech, a small web app plus its public demo, and Honcho, which is four containers on its own (API, database, deriver, Redis). The proxy node carries Home Assistant, Mosquitto, and the NetBird server and dashboard.

On node 2, Langfuse, which traces the agent's model calls, is six containers (ClickHouse, web, worker, MinIO, Postgres, Redis). That's more containers than the thing it observes.

The ZimaBlade is the photo library and the networking plumbing: Immich (server, machine learning, Postgres, Redis), the tunnel connector and reverse proxy, a Homepage dashboard, and omp, a terminal coding agent that points back at the LLM box, plus a few small utilities.

Two things the census doesn't show

Node 2's guests all live on a single 5400rpm hard drive. The agent moved there for the extra RAM, and it was the right call, but every reboot is a queue of cold-starting databases fighting over one spindle. The first time it happened the agent's dashboard answered 502 for about 15 minutes. The fix was stopping Langfuse, the heaviest starter and the one thing nobody misses for ten minutes. The dashboard came up 21 seconds later.

The move itself left a twin behind. The original agent box on node 1 was never stopped: same MAC, same IP, same VPN identity. What I called "flaky networking" for days was two containers taking turns answering for one address. A census would have caught it on day one.

The AI side, and the card that decides it

Here's the rule the whole routing falls out of: the agent framework I run refuses any model with less than a 64K context window, and the LLM box has 6 GB of VRAM.

Those two facts don't get along. The best local model I've measured for the job is qwen3.5 9B, built as a 64K variant. At 64K context it's 7.66 GB in total, of which 3.93 GB lands on the card. That's 51% on the GPU and the rest on the CPU, and no amount of context tuning fixes it: at 8K context it's still only 64% on the GPU, because Ollama pins roughly the same slice of VRAM and spills the remainder.

I spent much of August trying to shop my way out of that. Every candidate was measured and rejected:

  • Smaller qwen2.5 builds fit the card much better, but their context is architecturally 32K. Ollama silently clamps a 64K setting back down, so the framework rejects them outright.
  • The 27B loaded at 4096 context when I benchmarked it, still far below the 64K floor, and in a like-for-like test prefilled at 132 tokens a second against 604 for the 9B, and generated at 3.0 against 10.7. A typical agent turn would have spent over two minutes before producing its first token.
  • qwen2.5-coder, the obvious pick for a coding agent, can't tool-call. It writes the JSON for the tool call into its reply as prose, and the agent prints it and stops. Its 16K and 32K tags are still on disk, a small memorial to a good idea.

So the local model is genuinely the best option on this hardware, and a real agent turn on it still took around a minute and a half when I timed it in August. The same work on a cheap cloud model answers a tool-call probe in 1.5 to 2 seconds.

That gap decides most of the routing.

Who goes where

Job Runs on Why
Coding agent OpenCode Go subscription (deepseek-v4-flash) Flat $10 a month. Usage caps, not a meter
Chief of staff, morning brief, research, writing, travel OpenRouter qwen/qwen3.7-flash Pay per token, about $1.39 a month at measured volume
Security audit and code review Local Ollama only Deliberately no third-party egress
Embeddings, session titles, compression Local Ollama The framework's auxiliary calls stay local even when the main model is remote
Honcho user modelling Local Ollama Background work where latency doesn't matter
Images ComfyUI (SD1.5) on the same card Free, and small enough to share
Speech Whisper and Piper containers, CPU Never touches the GPU at all
Last resort An OpenRouter free model Rate-limited, so it only ever goes last

Paid cloud: OpenRouter

The cloud model became the primary for the general profiles in late August. The number that justified it came from measuring 30 days of real usage: 44.35M input tokens against 421K output. Input outweighs output by about 105 to 1, which means input price sets the bill, and headline "per million" pricing is close to useless for ranking models. On a model at $0.03 per million input tokens, that volume comes to roughly $1.39 a month.

The chain is vendor-diverse on purpose: qwen/qwen3.7-flash, then deepseek/deepseek-v4-flash-0731 as a second paid model from a different vendor, then the local 64K model, then a free model last. A fallback that shares its primary's failure mode isn't a fallback.

A watchdog probes the primary every three hours for two failure modes: the model dying, and the money running out. A healthy model you can't pay for is still an outage, so below a credit floor it moves everything back to local and says so.

It also got something badly wrong. On 7 September every paid model "failed" its health check at once, so the watchdog moved four agents (chief of staff, research, writing and travel) back to the local model. Nothing had died: an upstream PII-redaction filter had turned the city in the probe prompt into [ADDRESS], so no model could answer it properly. Worse, the watchdog only restored automatically after a credit revert, not a failure revert, so those four agents stayed on the slow local model for 17 days until I caught it on 24 September. It now re-probes after any failover and moves back when a model passes. A probe that fails the same way on every vendor is a problem with the probe, not the models.

Subscription: OpenCode Go

The coding agent is the heaviest user by a distance, so it's on OpenCode's Go plan instead: $10 a month, with rolling caps of $12 per five hours, $30 a week and $60 a month. Every response reports a cost of zero. On a subscription the risk isn't a bill, it's hitting a cap halfway through a task, so the watchdog reads Go's usage endpoint and tells a cap from a death. A capped profile gets parked on OpenRouter and put back automatically when the window resets.

One trap worth passing on: Go lives at /zen/go/v1, not /zen/v1. The second is pay-as-you-go and returns "insufficient balance" on a perfectly good Go key, which cost me an hour and a wrong conclusion.

Local: Ollama

Local isn't the cheap option here, it's the private one. The security and review agents see the most sensitive estate detail and never leave the LAN. Everything else local is background work, or the fallback for when the internet or a credit balance lets me down.

Running it taught me more than the cloud side did:

  • A hang never triggers a fallback. Something on the network kept asking Ollama to keep a model loaded forever. With only 6 GB, the scheduler waits for room rather than refusing, so every other model request just sat there, and Ollama logged nothing. Fallback chains fire on errors. This wasn't an error, it was silence, and every local-model profile was dead with nothing reporting it. A small script now unloads anything pinned for more than 48 hours, every 15 minutes.
  • The prompt was the bottleneck, not the model. A captured agent turn was about 18,400 prompt tokens, 74% of them tool definitions. Turning off toolsets the specialist agents had never once called cut their tool schemas by 47%. That was worth more than any model swap available.
  • Measure what's on the wire. For a while every benchmark I ran through the framework's one-shot mode was invalid, because that path never sent "thinking off". The real workloads were fine. My test harness wasn't.

What I'd tell someone starting

Count first. Take the census from the hosts, not from memory. Mine turned up a twin container and a tracing stack bigger than the thing it traces.

Let VRAM make the first cut. Write down the smallest context your agent needs, then check what fraction of a model at that context actually lands on your card. If the answer is half, you have your routing decision.

Rank cloud models on input price. Agent workloads are overwhelmingly prompt. Output price is mostly a distraction.

Keep local for the jobs that must stay local, and as the fallback. Don't make it carry interactive work it can't do quickly.

Assume silence is a failure. A pinned model, an empty ledger, a probe that swallows exceptions: none of them raise an error. Build the check that notices nothing happening.

A 27B model on 0.5 GB of VRAM is a fun number. It's also a fair summary of the whole exercise: the hardware will happily let you do something slow and call it running. The census is how you find out.


🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.

来源:Google AI:DEV 作者专属(RSS) · dev.to