When an agent does the same chore twice, a retro turns it into a script and writes down where it lives. The next agent runs one command instead of ten.
When an agent does the same chore twice, a retro turns it into a script and writes down where it lives. The next agent runs one command instead of ten.
These are field notes from September 2026, from one repo where a fleet of agents builds alongside me. The pattern is simple: every time an agent repeats a chore, the chore becomes a script, and the next agent is told where the script is.
An agent has no memory between runs. Each one arrives fresh and rediscovers how your deploy works, how to check whether a test was already failing, where the logs are. It does this carefully, one tool call at a time, and pays for it every run. A script, and a pointer to it, is what you give it instead.
The retro is the part that makes it work. After each substantial run, a separate agent reads what happened and works through a short list of questions. The one that pays most asks: did it do the same multi-step thing by hand more than once? If yes, that becomes a script in tooling/, with one rule.
Every script ends in a line you can grep for. Something like STATUS_OK=1, printed only when the whole thing worked. Through a pipe, the exit code you see is the last command's, not yours.
The retro also writes a pointer to the script where the next agent looks first. Ask the script, not the agent is about that half.
At least seventeen of these now carry, in their header, the chore they replaced. That header matters: when someone asks why a script exists, the answer is a measured cost, not a preference.
Don't retro the boring runs. Once a pattern has run clean three times in a row with nothing to clean up, repeat runs skip the retro unless they did something new. Otherwise every retro reports "nothing to learn", and that is its own waste.
Measure before you claim. I believe an agent babysitting a deploy burns far more tokens than a script doing the same polling. I have the tool-call counts, about ten calls down to one. I have not measured the tokens yet, so that post comes later.
This is how I work right now, and the scripts will keep changing.
If you do the same multi-step chore twice, turn it into a script in tooling/. The script prints NAME_OK=1 as its last line only when everything worked. Grep for that line; never trust an exit code through a pipe. Add one line to the agent's instructions that names the script.