The constraint
SMAARiX is small, and the size is not the interesting part. What matters is that the work in front of us was sized for a much larger organisation, and closing that gap by hiring was not the option on the table.
So the question was never whether to use agents. It was which parts of an engineering function I was willing to hand over, and what I would need to see back before I believed any of it.
There is a second thing worth saying, because it shaped everything after. I was not supervising this from above. I was in it — writing the specifications, getting the actual requirement out of people who were describing a different one, then reviewing what came back. That is an uncomfortable place to stand. It is also where I learned the most, because you cannot hide a badly written assignment from yourself when you are the one who has to go and fix it.
What carries over from managing people
The surprise was how little of it was new.
An agent that produces work you have to check is a junior engineer. Not metaphorically — structurally. Whether the arrangement helps or creates a second job for you depends on three things.
- Specification
- Define the work and what will count as done.
- Authority
- Set the limits of what the delegate may do.
- Verification
- Decide what evidence must come back with the result.
Get those wrong with a person and you get rework and resentment. Get them wrong with an agent and you get rework and a large bill. The failure is identical in shape and it arrives faster, which is the useful part. Agents compress the feedback loop on management decisions from quarters to hours.
Specification. The instruction that fails is the one that describes the outcome and assumes the context. A person fills that gap by asking, or by knowing you. An agent fills it by inventing, confidently. So the work has to name the exact files, the acceptance criteria and the boundary of what may be touched — and writing that down has improved the assignments I give people, because most of what I used to leave implicit was never shared understanding. It was me not having decided yet.
Authority. Delegating a task and delegating the authority to act are separate, and conflating them is the most expensive mistake available in both cases. The rule I run is the same one payment systems have always used: the delegate cannot hold authority the delegator does not hold, the grant is scoped to the task, and it expires with it.
Verification. A report that work is done is not evidence that work is done. This is the one that does not transfer cleanly, and it is where the analogy breaks in a way worth being careful about. With a person, trust accumulated over time is a legitimate and efficient substitute for re-checking, because a person has continuity, reputation and something at stake. An agent has none of the three. Its record of being right ninety-nine times is not a reason to accept the hundredth, and the instinct that says otherwise is the same automation bias that hollows out an approval gate.
Where the analogy stops
The one I got wrong was verification, and I got it wrong in the most ordinary way available. I let the thing that did the work also tell me the work was done.
It does not look like a mistake while you are making it. The agent builds the thing and runs the checks, and having both come back in one place is simply efficient. What you have actually assembled is a closed loop: the work and the evidence about the work produced by the same party, in the same pass, with nothing independent in between. I have spent a career in systems where that arrangement has a name and a control designed specifically to prevent it. I built it anyway, because it arrived looking like a status update rather than like an audit.
The rule I run now is written into the contract that every session loads: an implementer never certifies its own assured work. It lives in the contract rather than in my head on purpose. A rule I have to remember is not a rule, it is an intention.
That is the sharpest difference from managing people, and the exactness matters. With a person, a long record of being right is a legitimate reason to check less often. That is not laziness, it is how trust is supposed to work, and it is efficient. An agent’s record of being right is not the same object, because there is no continuity behind it and nothing at stake in the next answer. The instinct to relax is identical in both cases. In only one of them is it earned.
The things I would not delegate now cluster, and the cluster has a shape. They are the decisions where being wrong is expensive and being confident is cheap: what to build next, what to stop building, which constraint is real and which is inherited, whether a disagreement between two engineers is technical or is about something else entirely.
None of those are hard because they need more compute. They are hard because they need someone who will still be there in six months to live with the answer.
What this work does and doesn’t show
I want to be careful about the claim, because there is an inflated version of it doing the rounds and it is not the one I am making.
I am not claiming an engineering function needs fewer engineers. I am claiming that if you are going to run one this way, the parts you cannot skip are the parts good engineering managers were already doing and mostly getting away with not writing down. Specification, scoped authority, evidence before acceptance. Agents do not let you get away with it. They fail immediately and visibly on the exact assignment a person would have quietly rescued.
Which makes this a fairly unusual instrument. Anywhere the process was carried by somebody’s judgement rather than by the design, an agent will find it inside a week.
Where I think this goes is not fewer engineers. It is a change in which half of the job is scarce.
The pairing that keeps surfacing in this work has two sides. One person builds the thing inside the real constraints, in the customer’s environment, with all of its inconvenient truths. The other reads the situation — the undocumented workflow, the politics, the requirement nobody wrote down and possibly nobody knows — and turns it into something buildable. Most organisations hire hard for the first and quietly hope the second arrives attached to it.
Agents are getting good at the first half, quickly. They are nowhere near the second, and I do not think that gap closes soon, because the second half is mostly about sitting in a room with people who are not telling you the whole thing and working out what is missing.
So if I were staffing an engineering function now, I would stop hiring hard for the half that is becoming cheap, and start hiring deliberately and expensively for the half that is not.