We got very good at making agents write code. Real code, a lot of it, fast, and that was worth doing.
Writing code is skilled, expensive work. It is a large share of what you pay an engineer for, and handing it to an agent was real progress, not a party trick. But it is one part of the job. It is the part with a clean shape: a task goes in, a diff comes out, and one agent can finish it on its own. That shape is exactly what made it the part you could hand off. Everything else between a diff and a change your customers can use still needs a person, because it needs judgment and it needs someone to carry the work from one step to the next.
So the agent writes the diff and stops. It does not review its own work with any authority. It does not decide the change is safe to merge. It does not know which of your engineers owns that subsystem. It cannot take a review comment, work out what is actually wrong, and walk the change back three steps to fix it. Every one of those is a person's job, and it did not go away when the typing did.
Your engineers are babysitting the agent
Watch what your best engineer actually does with agent output. They read a diff no human wrote, closely enough to put their name on it. They run the tests and read what failed. They decide whether it is safe to ship. They find the one person who knows that subsystem. When review turns something up, they write up what is wrong, hand it back to the agent, wait, and read the whole thing again from the top.
That is not reviewing a colleague's work. It is standing over the agent at every step, because the agent cannot take a step on its own. It does not move from writing to testing. It does not notice it is wrong and back up. A person moves it forward, a person moves it back, one change at a time. Your engineers are not the agent's teammates. They are its hands and its memory, minding it through a loop it cannot run itself.
The faster the agent, the more of this there is. More diffs land, and each one still needs a person to thread it through review, tests, routing, and the loop back. You automated the one step that could run without a human and poured more work into every step that cannot.
If you run the team, you saw it as a number first
You raised the token budget because more code should mean more shipped. The bill climbed. Then you looked at what actually reached customers, and it had not moved. Or it moved once, the quarter everyone turned the agents on, and then it flattened and stayed there.
That gap is the whole story. More spend bought more drafts. It did not buy more of the human attention every draft has to pass through before it ships, and you did not hire a second team to supply it. So the drafts stack up behind the same people, and the number you were watching sits flat while the invoice keeps rising.
The question changed
For three years the question was how fast a model can write code. That is answered, and answering it moved almost nothing your customers can feel. The question that decides whether any of this pays off is who, or what, runs review, tests, routing, and the loop back once the code exists, and how those steps get faster without one of your engineers personally minding every one of them.
Generating code faster does not touch that. Most of the field is still buying the front of the pipeline and wondering why the back of it never moved.
Sign up at wallfacer.ai and work with our team to automate the rest of the job, so your engineers stop babysitting agents and start managing them.