跳到正文
原文
Google AI:DEV 作者专属(RSS)· ElderAI·· 4 小时前AI 评分42

ElderAI 用 100 美元预算微调 ATLAS Code 编程模型的经验

What we learned fine-tuning our own coding model on a $100 budget

AI 导读

ElderAI 在约 100 美元租用 GPU 预算下微调其编程模型 ATLAS Code,目前所有微调版本均未通过质量门槛,邀请制预览仍运行原始 checkpoint。团队采用预先写定并哈希锁定的评测门槛,要求工具调用解析率至少 97%、通用编程最多丢一题,并靠 40% 中途检查提前止损。

正文

We're a small team at ElderAI building ATLAS Code, a coding model for agent tools (Cline, Aider, Continue, Cursor and anything else that takes an OpenAI-compatible base URL). We do our own training runs on rented GPUs with about $100 of prepaid compute. Here's an honest account of the last two days.

The short version: none of our fine-tunes has cleared its quality gate yet. That's why the invite-only preview still runs the starting checkpoint we're trying to improve. We'll swap in a fine-tune once one earns it, and we'll say so when that happens. Here's what we learned while getting there.

1. Write the gate before you look at the results

On a small budget it's tempting to look at a run, find the metric that went up, and call it a win. We stopped letting ourselves do that.

Every run now starts with a small gate file written before any metric is computed. It's hashed, and the launcher refuses to start if the file changes. For our latest run it said:

  • Edits: the fine-tune must produce more byte-exact correct files than the starting checkpoint on our edit test set. A tie fails.
  • General coding: it may lose at most one problem on a standard Python coding benchmark compared to the starting checkpoint, measured in the same job.
  • Tool-call format: at least 97% of its tool calls must parse.

The starting checkpoint's numbers are measured in the same job with the same harness, so we never compare against a number from a different setup.

The gate has already stopped us from fooling ourselves more than once. In the latest run, the fine-tune was ahead on the edit metric at the 40% checkpoint but failed the tool-call parse bar by about one call in a hundred. The gate said stop, so we stopped. (Funny detail: the starting checkpoint also sat right at that bar in the same job. That tells us the bar is very tight, but we don't get to loosen it after seeing the result.)

2. Teaching tool-call format can quietly cost you coding skill

Our first instinct was to pour agent-style data (read a file, call edit_file, finish) into training to make the model better at tool use. Format did improve. But the general coding check slipped a little at the same time, enough to fail the "lose at most one problem" rule on several runs.

What helped:

  • Rehearsal data. We mix plain code-generation examples back in, identical across runs, so the model doesn't drift away from just writing code.
  • Gentler updates. Small LoRA adapters and low learning rates. A larger update that also touched the MLP layers moved the edit metric more, but it's also the riskier setting for general skill.
  • A midcheck at 40%. We evaluate a merged checkpoint partway through and kill the run if general coding has already dropped. Most of our savings came from this.

The tradeoff is real, and on a small model you feel it fast. We don't see it as a bug to patch. It's the main thing we have to manage.

3. "Did the edit apply?" and "is it byte-exact?" are different questions

Our first edit metric was strict: after the model's tool call, the file had to match the real commit byte for byte. Almost everything failed, the starting checkpoint and every fine-tune alike, so we went through every failure by hand. We found:

  • The scorer wasn't buggy. Whitespace, tabs and line endings explained only a handful of failures.
  • The test task was mostly impossible to get exactly right. Each item used a real commit message as the instruction, such as "Increase spacing for quadrature encoders", with the real post-commit file as the target. The actual commit changed spacing=3 to 6. No model can recover the 6 from that message. Many commits also contained unrelated whitespace edits.
  • The model was also often wrong: lazy one-line edits, off-target changes, and old_str values that weren't unique in the file.

So we changed what we measure instead of tuning the scorer until the numbers looked better:

  • edit_applies: every edit call found a unique match and the file actually changed. This is a mechanical check.
  • Whitespace-normalized exact: line endings, trailing spaces and blank lines are ignored, but indentation still counts (it matters in Python).
  • A "precise instruction" test split: the same commits, but the instruction spells out the change exactly. That makes exact match a fair test.
  • Up to 3 tool calls with real tool errors fed back. A model that gets "old_str not found", then retries and gets it right, should get credit for the final file, the way a real agent loop works.

Exact match stays the official number. The others are reported next to it so we can see why something failed.

4. Two small failure modes that showed up everywhere

  • Raw tabs inside JSON strings. When the model quoted tab-indented code inside a tool-call argument, it sometimes wrote a literal tab, which is invalid JSON. This one escaping issue accounted for most of our tool-call parse failures. Before throwing compute at it, we're looking at fixing it with training data: we oversample tab-indented and backslash-heavy files and add a few "wrong call → real error → corrected call" examples, with no training loss on the wrong call.
  • Non-unique old_str. The model picks a snippet that appears twice, and the edit tool correctly refuses. Our training trajectories now use the smallest whole-line snippet that is unique at the time of the call, and every training row is checked to reproduce its target file exactly.

5. Cost caps and watchdogs

Every run has a watchdog on the box that controls it:

  • a hard per-run dollar stop, a projected-cost stop, a wall-clock cap, and a nightly cap;
  • stops if the logs stall, the GPU sits idle, or the loss goes NaN;
  • verifies the machine is actually deleted when a run ends.

A watchdog can also be too strict. Two of our recent attempts were stopped by the projected-cost rule while the job itself was healthy. ETA jitter early in training pushed the projection a few cents over the cap. We paid about $1.18 to learn that, then fixed the projection headroom.

Our biggest single saving is the midcheck early stop. A run that fails at 40% costs about $1.40 instead of about $3.

For scale: our recent attempts (several pilots, two full runs that stopped at midcheck, and the latest gate run) came to about $15 of GPU time in total.

What's next

  • One more gated run with the precise-instruction edit data, with a gate already written and unchanged.
  • If it passes, ATLAS Code gets the fine-tune and we'll publish what changed. If it fails, we'll write that up too.

Try it

ATLAS Code is in invite-only preview behind an OpenAI-compatible /v1 API, with a Playground in your account. New accounts get 200 free credits, and plans are capped: when credits run out, requests stop, with no overage. We don't train on your prompts or code.

If you run Cline, Aider, Continue or a similar agent and want to tell us where it breaks, request access at elderai.cloud. We read everything people send.

来源:Google AI:DEV 作者专属(RSS) · dev.to