Engineer’s Heart, No Engineer’s Training
I’ve got an engineer’s heart and none of the training. Turns out that’s the whole problem for whoever comes after me.
TL;DR:
- For most everyday work, the frontier model is basically a commodity now. The gap you actually feel between the top models on normal tasks is small, and shrinking.
- What’s not a commodity is knowing what to build, which path to take, and whether the output is any good. That’s judgment, and judgment is useless if you don’t know what to judge or how.
- You build judgment by doing the work. Slowly, by hand, getting it wrong a lot. Which is exactly the part AI now does for you.
- I learned the slow way because there wasn’t another one. The next generation gets to skip the struggle. I’m not convinced they can skip what the struggle was quietly teaching.
- So the question I keep landing on: how does the next engineer earn judgment when the work that used to teach it is the first thing we automate?
A benchmark, and a lesson I did not order#
I’m not a great engineer. I never trained as one. I’ve got the engineer’s heart, the curiosity, the itch to pull things apart and see how they work, but none of the formal background. My code isn’t great and isn’t terrible. I don’t stand up big systems from scratch on my own. When a real architectural decision shows up, I go find people who do this for a living, and I actually plan before I touch anything. I’m telling you all this up front because the rest of this post leans on it.
A while back I set out to answer a simple question: can local AI actually red team? I ended up with two DGX Sparks and 33 models on a bench. The models turned out to be the easy part. The hard part, the part that ate weeks, was everything wrapped around them: tokenizer regressions, kernel limits, cross-node deadlocks, silent wrong-logits bugs. My own words closing that post were that “the models are the easy part … that’s the actual project.” And if I’m honest, even the harness I built could’ve been engineered better. More consistent, more reproducible, with a lot less of me patching the same thing twice. (More on that in the next part of that series.)
That’s where this post started. Same models everyone else can download. The whole difference was in how carefully the thing around them was built, and whether I knew enough to catch it being quietly wrong.
Does anyone actually feel the difference anymore?#
Here’s a question I keep putting to people who build with AI for normal work. Think about the frontier models that shipped in the last twelve months. Astra, Opus, Sonnet, Daybreak, pick your lineup. For the everyday stuff, drafting a doc, summarizing a thread, refactoring a function, writing a quick script, how much of a difference do you honestly feel between them?
For most tasks? Honestly, not much. They’re all good. The jump from “no AI” to “any decent frontier model” is enormous. The jump from one frontier model to the next, for general work, is something most people would struggle to pick out in a blind test.
There are exceptions, and they matter. In narrow, hard domains the gap is still real. My own red-team benchmark is one of those corners, where correctness is the scarce resource and the models are very much not interchangeable. But that’s a specialist’s problem. For the broad middle, where most of the building actually happens, the model has quietly stopped being the thing that decides whether your work is any good.
So if the model’s basically a commodity, what’s left? For what it’s worth, I run Claude Code as my daily driver because the harness around the model is good. The model it happens to use is almost beside the point.
Everyone shipped a harness#
Look at where the real effort is going. Almost nobody’s training their own frontier model (good luck with that). The effort goes into the harness: the scaffolding that feeds the model context, runs its tools, checks its work, and decides what to do with whatever comes back.
The clearest tell is vulnerability finding. In about a year, a startling number of outfits shipped an open harness for it:
- Visa released VVAH, a harness for autonomous vulnerability discovery, deliberately model-agnostic. Its own words: “no single provider is a hard dependency.”
- Capital One released VulnHunter, which goes the opposite way and is “built and optimized for Claude Opus running in Claude Code.”
- Protect AI has Vulnhuntr, LLMs plus static analysis to find exploitable bugs.
- OpenAI introduced Aardvark, an agentic security researcher that hunts for vulnerabilities in code.
- Google’s Project Zero and DeepMind built Big Sleep, which found a real exploitable bug in SQLite.
- Anthropic open-sourced a Claude-based security review action for pull requests.
- Trail of Bits built Buttercup for DARPA’s AI Cyber Challenge.
I’m not the only one clocking this. Someone put it bluntly in a security group chat the other day:

Now tell me which of those harnesses is the best.
You can’t answer that by looking at the model. A bunch of them run on the same frontier models. Two of them are banks with nearly identical problems that made opposite calls on whether to even hard-wire the model. The only way to rank any of them is to understand how it was built: what it does with context, how it checks a finding, where it’ll quietly lead you off a cliff, and whether you know enough to drive it. The model barely enters into it. The engineering is the whole answer.
You have to know what good even looks like#
This is the uncomfortable part. Judgment isn’t some free-floating talent you either have or you don’t. It’s useless if you don’t know what you’re judging or how. You can stare at two harnesses all day, but if you’ve never built one and never watched one fall over in production, you’re not really judging anything. You’re guessing with extra steps.
I can make those calls, more or less, because I picked up the reference points the slow and annoying way. I came up before any of this, when learning something meant doing it yourself, breaking it, reading the docs you didn’t want to read, and poking at it until it finally made sense. Nobody handed me a working answer in ten seconds. The struggle was irritating. It was also the whole education. Every dumb call I shipped and then had to unwind turned into a pattern I can now spot in about half a second.
And to be clear, I still get it wrong. Plenty of times. Judgment doesn’t make you right. I wish it did. It just means you usually know which questions to ask, and roughly how badly a call can burn you when you get it wrong. I make decisions I regret, I misjudge things, and some of my “yeah, that looks fine” moments age terribly. The reps didn’t buy me a crystal ball. What they bought me is a feel for the shape of what I don’t know, and the sense to drag in an actual engineer when a decision is over my head, before I commit instead of after it’s on fire. That, too, is a skill. It came from reps, not from a prompt.
And yes, I build almost everything with AI now#
You can see the objection coming, because I’ve got it pointed right at myself. I build with AI constantly. I stood up an entire lab that stitches together a coordinator, an offensive tooling layer, and a sandbox. I built ghosttype in a few hours with Claude Code. A conference submission tracker for my team came together in one evening. I’m not giving any of that up, and I’m not going to tell you to either.
So this isn’t a “put the tools down and suffer” argument. Use them. Use them hard. The catch is that the engineering skill underneath doesn’t get to quietly become optional just because the machine is doing the typing now. AI runs on that skill. Take it away and you’re mostly forwarding whatever the model said and hoping for the best. The same tool makes a good engineer faster and a clueless one confidently wrong, and the only thing standing between those two outcomes is the human holding it.
The manual years I put in are exactly what let me look at AI output and feel when it’s off before I can even explain why. That feel is the skill. AI leans on it hard, and it’s the one thing the tool will never hand you for free.
That’s what actually worries me, and it’s slower and less dramatic than the usual AI panic. If nothing forces anyone to build that skill anymore, it doesn’t vanish in a bang. It rots, gently, in us and in the people coming up behind us, until one day a whole team can ship something and not one person in the room can say whether it’s any good. Keeping the skill alive has to become a deliberate choice now, for ourselves and for whoever we’re training, because the job stopped demanding it on its own.
I don’t pretend to have the perfect answer for how. This was my path, and I honestly don’t know if it’s the best one or just the one I got stuck with because of when I happened to be born. But giving the skill up isn’t on the table, so here’s roughly how I try to hold onto it.
So how do you build judgment now?#
If the model’s a commodity and the harness is the craft, and that craft comes from reps you can now skip, then the real problem for the next engineer is simple and a little ugly: nothing is going to make them struggle anymore, and the struggle was where the learning lived.
Some of this is just the old stuff, and the old stuff still works:
- Build your own labs. Stand the thing up by hand at least once, even when AI could scaffold it in sixty seconds.
- Read the actual material. The docs, the source, the RFC. Not the AI’s cheerful summary of the RFC.
- Do the thing by hand before you automate it, so you actually know what the automation is hiding from you.
None of that is new. The part that is new is the one I’d put right at the center.
Learn to explain your choices. Not just what you did with the AI, but why you went this way instead of the other way, and what it costs you. If you can’t defend the choice, you didn’t make it. The model made it, and you just shipped it with your name on it.
This isn’t hypothetical for me. On the red-team bench, I let Claude pick the specific model versions to run through the suite. It picked some duds. We didn’t catch it until later, when the numbers stopped making sense and we had to go back and redo the whole selection. The pick looked perfectly reasonable. It wasn’t, and for a while nobody could actually say why those versions and not others. That gap, between “the AI chose it” and “I can tell you exactly why this was the right call,” is the whole game. The new skill is closing that gap on purpose, every time, instead of letting a confident-sounding answer walk you off somewhere dumb.
So if you’re early, or you’re bringing up people who are: use the AI, then make yourself account for it. Take the harder road on purpose once in a while. Keep enough friction in the loop that something actually sticks. Sand all the friction off and there’s nothing left to learn from.
The part I cannot answer#
I don’t have a clean ending for the question I opened with. The models will keep getting better and cheaper. The code will keep getting easier to crank out. And I’m fairly sure the people who come out of this well won’t be the fastest prompters. They’ll be the ones who can still look at a slick, finished-looking output and say, with actual reasons, “no, this is wrong, and here’s why.”
That ability has only ever come from one place.
“For the things we have to learn before we can do them, we learn by doing them.” - Aristotle, Nicomachean Ethics