资深工程师反思:AI 代写代码后,初级开发者的判断力从何而来
A junior asked me how I knew the code was wrong. I couldn't answer him.
一位资深工程师与入职约八个月的初级开发者结对时,凭直觉判断 AI 生成的代码有误却无法解释原因,十分钟后验证确实有错。他将这种判断力追溯到多年前一次"先确认后落库"的生产事故,认为它来自真实事故的反复打磨而非规则。他担忧 AI 接管生产代码后,初级开发者不再经历写码、上线、深夜救火的磨砺,产生判断力的机制被自动化,却尚无替代方案。
We were pairing on a Thursday. Nothing dramatic. A junior on my team, maybe eight months in, sharp, the kind who actually reads the code instead of pasting it. We'd asked the AI for a chunk of logic, it came back in a few seconds, clean and typed and already passing the tests he'd written.
He went to accept it. I said, "no, don't ship that."
He stopped. Looked at it again. Looked at me. And he asked the most reasonable question in the world.
"Wait — how did you know that's wrong?"
And I opened my mouth to explain, and nothing came out.
I knew. I just couldn't say how.
Here's the thing that rattled me, because it wasn't a dramatic bug and this isn't a dramatic story. The diff was wrong. I was right that it was wrong, and ten minutes later we proved it. That part was fine.
What wasn't fine was that I had no idea how to tell him what I'd done. I hadn't run a checklist. I hadn't spotted a specific line and matched it to a rule. Something about the shape of it had made the back of my neck go cold, and I'd learned, over a lot of years, to trust that exact feeling. But "trust the cold feeling on the back of your neck" is not a thing you can hand to a human being who hasn't grown the neck yet.
I tried anyway. I said something about how it felt off, how the error path looked too convenient, how I'd want to know what happens if this gets called twice. All true. All useless to him, because every one of those was a conclusion, and he was asking for the method, and I didn't have a method. I had a scar that fired before I'd finished reading.
That's the part I've been chewing on since. Not that I couldn't teach him. That the single most valuable thing I do all day is the one thing I have no idea how to transmit.
Two things get bundled into the word "senior"
For years I thought being senior was a pile of knowledge. And some of it is, and that part I can teach. "Validate at the boundary." "Don't trust the client's timestamp." "Wrap the external call." Rules. I can write them on a whiteboard and he'll have them by Friday, and honestly, the AI already has all of them and applies them more consistently than I do.
But the thing I did on Thursday wasn't a rule. It was the opposite of a rule. It was knowing, against a clean diff that satisfied every rule either of us could name, that something was still wrong. Rules tell you what to check. The other thing tells you to keep looking after every check has passed. One is knowledge. The other is judgment, and I have never once been able to put it into words that survive contact with someone who hasn't earned it.
You don't learn judgment. You survive into it — and that's exactly why I couldn't hand it to him.
Nobody taught me the thing I did on Thursday. There was no course, no cert, no senior who sat me down. I got it the only way I think anyone gets it: I was confidently wrong about something that mattered, in front of someone who paid for it, and the lesson got welded on.
Where the cold feeling actually came from
Let me be specific, because I can trace that exact flinch to its source.
Years ago I shipped a write path that told the client "got it" before it had actually saved the row. The code was clean. Genuinely clean — idiomatic, typed, tests around it, error handling buttoned up. It read like someone careful wrote it, because I was careful. By every rule I knew at the time, it was correct.
Then one ordinary day a retry landed at exactly the wrong moment. The "got it" went out, the save never happened, and a paying customer got locked out of their own account with nothing in the logs to say they'd ever been there. Ack before persist. I can still feel the phone call.
That's the source. When I looked at Thursday's diff and my neck went cold, I wasn't running an algorithm. I was pattern-matching against that customer. The AI's clean code looked exactly as correct as my clean code had looked, right up until it cost someone their account. The flinch is that customer, compressed into half a second and welded onto my nervous system by the worst night of that year.
So when the junior asked "how did you know," the honest, full answer was: a person I locked out at 2am four years ago taught me, and I have no idea how to give you that without the person. You can't lecture someone into the flinch. You can only get unlucky enough times that it grows on its own.
And here's the part that actually scares me
For twenty years, the flinch had a reliable supply chain. You got it by doing the grind — writing the production code yourself, shipping it, and being there when it broke. The small failures stacked up into judgment whether you wanted them to or not. You couldn't skip the grind, so you couldn't skip the scars.
Watch what we just did to that supply chain.
The junior doesn't write the production code anymore. The AI does, and it does it well, and I'm glad — the typing was never the hard part. But the typing, the shipping, the breaking, the 2am call: that was the entire mechanism that used to turn a junior into me. We didn't just automate the boring part. We automated the forge. I got my judgment because I had to do, by hand, the exact work the AI now does for him before he ever feels it go wrong. He gets the clean output on day one and never has to earn the neck.
So the one skill that survived AI — the only one, the "no" — is also the one skill the AI quietly stopped manufacturing in the next generation. The machine can hand him the code. It cannot hand him the scar, and the scar was the teacher.
To be clear, this is not "make juniors suffer"
I'm not romanticising pain, and I'm not saying pull the AI and make him hand-write CRUD until he bleeds for it. That's cargo-cult mentorship — the suffering was never the point, the consequence was. And I use AI all day; I'd never go back. The problem isn't that he has it easy. The problem is narrower and weirder than that: we removed the thing that used to produce judgment, and we haven't replaced it with anything.
So that's the actual job now, and it's a design problem, not a vibes problem. I can't give him the flinch. But I can give him the one thing that grew the flinch in me — being wrong somewhere it's cheap to be wrong — on purpose, faster, instead of waiting for a real customer to do it the expensive way.
A few things I actually do with him now:
I make him the skeptic, not the author. The AI writes the diff. His job isn't to write a better one — it's to break the one we got. Before we accept anything, he has to finish the sentence "this loses money when ___." He's wrong most of the time. Doesn't matter. The rep is the distrust, not the catch. That's the muscle, and it only grows when it's his job to doubt.
I manufacture the consequence small. Ship it behind a flag, to one internal user, to a canary. Let it break where breaking is cheap, so the lesson arrives before the customer does. A flinch you got from a staging incident is the same flinch — it just didn't cost anybody their account.
I stop answering "is it right" and make him answer "how is it wrong." The question trains the instinct. "Looks good" trains nothing. If I just tell him what I saw, he learns my conclusion. If I make him hunt for how it breaks, he starts growing his own neck.
Why this is the exact reason I build the way I do
One level up, same problem.
I work on an agent platform, and the whole industry right now wants to cheer for the thing that produces. Look how much it ships, look how clean. But a thing that produces is just doing the skill that stopped being scarce — and worse, it's doing it in a way that signs off on its own work, the same confident "looks correct" over a masterpiece and a disaster alike. It's the AI version of a junior with no flinch: fast, fluent, and completely unable to distrust itself.
So I never let the thing that writes the code be the thing that blesses it. There's an author that produces the diff, cheap and endless. There's a separate skeptic whose entire job is to distrust that diff and try to break it — the flinch given its own seat, institutionalised, so it doesn't depend on anyone having gotten unlucky enough to grow one. And there's a human on the merge button, because somebody still has to own the call, and that human is the one still accumulating the real scars. Author, skeptic, human. That's the whole shape of xenition, and it's my actual answer to the junior's question: you can't teach the flinch, so you build the seat that does its job and you put the person in it until they grow their own.
I still can't tell him how I knew. What I can do is put him where he'll find out the way I did — wrong, early, and somewhere cheap enough that the lesson costs a canary instead of a customer. The knowing was never teachable. The getting-wrong is. That's the only part I can actually hand him.
来源:Google AI:DEV 作者专属(RSS) · dev.to