When an agent needs to know a state, it should run the small read-only script that answers the question, not the build or the test suite. Half the work is making sure it knows the script exists.
When an agent needs to know a state, it should run the small read-only script that answers the question, not the build or the test suite. Half the work is making sure it knows the script exists.
As of September 2026, the agents that build with me in one repo spend a lot of their calls finding things out. Is what is deployed the same commit as main? Was this test red before my change? Did the page actually hydrate? Left alone, an agent answers by running the thing itself: the build, the test suite, the browser, one tool call at a time.
For most of those questions there is now a script, and the rule I give agents is short. Don't work the state out yourself. Ask the script.
Without it, an agent starting a session on my ecommerce product asks six questions about what is running, over about ten tool calls. The status script answers the same six in one call, in about 8 seconds. It changes nothing: its one database read runs on a connection pinned read-only, so a write that slips into it fails.
It gets worse when the only way to know is to run the thing. To prove a failure was already on the base, you make a clean checkout of the base commit, install it and run the same command again. The header of the script that does this says that "costs an hour every time". In one batch of parallel agent runs, five agents each ran the same 3,740-test suite to prove one red test was not theirs.
Each script answers one question in plain words and ends with one line an agent can grep. The pre-existing check ends in exactly one of PRE_EXISTING=1 (the base was red too), PRE_EXISTING=0 (the failure is yours) or PRE_EXISTING=? (the check could not run). Its exit code is always 0, because it is a question, not a gate.
The last line matters more than the exit code. Pipe a command through tail and you get tail's exit code, which is 0 even when the command died. In this repo that let a cleanup step remove a workspace before its work had landed, twice in one session.
The status script adds one more rule. A probe it cannot reach prints unreachable, never a blank that reads as fine. Its sentinel means the report is complete. Whether everything is green is in the lines above it.
I counted today. There are 72 scripts in tooling/agent/, not counting their tests. 19 of them are not named in any AGENTS.md, agent definition, skill or command, which are the files an agent reads first. Several of the 19 are deploy watchers and smoke tests, exactly the kind of thing an agent writes for itself when it doesn't know one exists.
That has already happened. My platform agent's definition says: if you find yourself writing a polling loop, check tooling/ first. The line is there because one run hand-rolled four deployment watchers, which is the job one existing command does.
So the pointer is half the work. The scripts that get used are the ones named where the agent looks first. The ops agent's definition says its first action is always the status script. The platform agent has a "Tools first" table, one row per need and the command that meets it. What I don't have yet is one index for the whole folder. This is the shape I am moving to:
The tool-call counts come from each script's own header, written by the agent, or the review agent that runs after it, that paid the cost. Most carry a "~", and I have kept it.
I believe a script also saves a lot of tokens, because the agent reads one line instead of a page of logs. I have not measured that yet. It is the next number I want, run as the same task twice.
The 19 is a filename search. A script named only in a longer doc somewhere else counts as unpointed, and that may be fair, since an agent won't find it there either.
What I do now: when an agent answers the same state question by hand twice, the question becomes a script with a one-line answer. How a repeated chore turns into a script is its own post: The third time is one command.
Before you run a build, a test suite or a polling loop only to learn a state, look for a script in the index below. Run it once and read its LAST line, the sentinel (NAME_OK=1). Never judge success by an exit code through a pipe. If you answered the same state question by hand twice, name it in your report so it can become a script. Index (replace with your own rows; question -> command -> last line): - Is what is deployed the same commit as main? -> tooling/agent/status.sh -> STATUS_OK=1 - Was this failure already on the base? -> tooling/agent/is-pre-existing.sh -- <cmd> -> PRE_EXISTING=1|0|?