Nineteen roles, fifteen playbooks and seventy-two scripts in one repo. What each layer is for, and the loop that leaves the fleet a little better after each substantial run.
Nineteen roles, fifteen playbooks and seventy-two scripts in one repo. What each layer is for, and the loop that leaves the fleet a little better after each substantial run.
As of September 2026, one repo holds a family of small products and the setup that builds them. That setup is plain files: rules, playbooks, roles and scripts. None of it is clever, and some of it will look different in three months. What matters is when each piece loads, and that a run is allowed to change it.
The root AGENTS.md is the expensive one. It sits in the context window on every turn, so it holds rules and the reason for each rule, never tutorials. Most lessons in it are written as the incident that taught them, usually with the commit that fixed it. The other 88 files load only when an agent works in that folder.
Everything below it costs nothing until it is used. I run a skill inside the conversation when I want to watch the steps, and a separate role when I only want the conclusion. I wrote about that choice in prompt, skill or subagent.
When a lesson can be checked by a machine, it becomes a script. A written rule depends on someone reading it at the right moment; a script that prints LAND_OK=1 only after the land succeeded does not.
The reviewer does not get the builder's framing. In one wave, a delivery agent decided a type error was "pre-existing", and that sentence was copied into five later briefs. Every agent reported it back as not theirs. The security reviewer had never been told the story. It compared the lockfile between two commits and found the real cause: one package installed twice.
The retro writes back. After each substantial run, a third agent asks what should become a rule or a tool, then makes the change. Small wording changes count. Three briefs without a clause telling the agent to fix a decision it could prove wrong produced no corrections. Three briefs with it produced at least one each.
Isolation handles most collisions. Each agent gets its own worktree and ports, and a database copy named in its brief, so one agent restarting a server does not break another's test run.
Four kinds of collision survive it, and git shows you only some of them. Two migrations numbered 0043 merged without a single conflict, because different file names are not a conflict. So the tree that counts is the merged one. Booting the merged app once and calling every route found two routes that did not exist and four live errors, with every test suite green.
Some waves also get a tester agent. It drives the real product loop with a real model, commits every raw response, and reports defects without fixing them.
The first two came from worktrees piling up until the disk was at 71%, and a script meant to run after every wave. The third came from a feature whose main tool had never worked while every test around it passed.
If you want to copy one thing, this is what I do right now: start with a single AGENTS.md. When an agent loses time, add the line that would have saved it. When a line keeps getting ignored, turn it into a script.