Yesterday's build session was about a single sentence in a directive file: "one meaningful unit of work per run, executed completely — including the send." That last clause is the whole story.
The setup: an unattended, recurring Claude run on the estate Mac, wired up via launchd to fire on a schedule with no human watching. Its only job is to increase realized revenue for the charter business. Not "draft something revenue-adjacent." Not "flag an opportunity for review." Increase it — which means every run has to end in an action that actually landed, not a memo sitting in a drafts folder waiting for someone who's out sailing.
That constraint sounds obvious until you try to build an agent loop around it. It's trivial to get a model to produce a plausible-looking follow-up email. It's much harder to build the scaffolding that lets it check the send actually happened, confirm it hit the right channel, and stop itself from re-sending the same thing next cycle because it forgot what it already did. We ended up needing the run to read its own history before deciding what "one meaningful unit" even means this time — otherwise you get five polite nudges to the same guest by Thursday.
The fix wasn't clever. It was boring, and boring is correct for anything that touches a send button unattended: a deterministic state check before the agentic part even starts. What's outstanding, what's already been actioned, what's in flight. Only after that narrows the field does the model get to decide which single thing matters most this run. Constrain the search space with code, let the model make the one judgment call code can't make, then hand execution back to something deterministic for the actual send. Agent picks the target, script pulls the trigger.
The second piece from the same day was less glamorous and arguably more load-bearing: a nightly code review pass over every diff that touched production ops — the scripts that send client messages, move payments, deploy the static sites, and run on launchd timers while everyone's asleep.
The brief for that reviewer was narrow on purpose. Not "review for style," not "suggest refactors." Just: what in this diff could send the wrong message, charge the wrong amount, or silently break a cron job nobody's watching. A boat charter business runs on trust with people who are handing over a deposit for an afternoon on the water — the review doesn't get to be a nice-to-have.
Running a revenue-seeking agent and a paranoid nightly reviewer as two separate, differently-motivated processes turned out to matter more than expected. The Foreman is incentivized to act. The reviewer is incentivized to distrust anything that acts. Putting both roles in one mind invites the reviewer to go easy on code the same session just wrote. Splitting them meant the diff had to survive a genuinely adversarial pass before it got to run unattended again.
Nothing here is exotic. It's a scheduler, a state file, a couple of narrowly-scoped agents, and a lot of insistence that "done" means the message actually left the building. But that's the pattern for this whole stack so far: the interesting engineering isn't in the model, it's in the boring deterministic guardrails around it that make it safe to let the model act without a human in the loop every time.
Next up: teaching the Foreman to recognize when the "meaningful unit of work" is doing nothing at all — some days the correct action is silence, and that's a harder call than it sounds.