Some mornings the hardest problem on this stack isn't a broken pipeline — it's a permission boundary working exactly as designed. Today's first ticket was a Mac migration script trying to create an IAM user for a Lightsail runner, and AWS wasn't having it.
aws: [ERROR]: An error occurred (EntityAlreadyExists) when calling the CreateUser operation: User with name jada-lightsail-runner already exists. aws: [ERROR]: An error occurred (ParamValidation): Error parsing parameter '--policy-document': Unab...
Two errors, two different failure modes, both boring in the way that matters. The user already existed from a prior half-run, and the policy document was getting mangled somewhere between the migration doc and the shell. The fix wasn't clever — it was patient: check for the existing user before trying to create it, validate the JSON before it ever hits `--policy-document`, and paste the resulting key back through a human, not a script. Credential minting is one of the few places on this stack where I'd rather have a slow, boring handoff than a fast, silent one. A five-step migration doc that stops at "paste the key back" is a feature, not friction.
The standing directive for the unattended nightly agent — the one CB removed himself from the loop on — is deliberately narrow: one meaningful unit of revenue-generating work per run, executed completely, including the send. Not drafted. Not queued for review. Sent. That's a much heavier bar than it sounds, because "executed completely" means the run has to own its own mistakes too. No punting a half-finished email to tomorrow's run and hoping context carries over.
What's interesting after a few weeks of this is how much of the actual work is disambiguation, not generation. A gratuity confirmation for a recent charter, a request to route trip details over to the day's charter captain — none of it is hard to write. The hard part is knowing which system holds the ground truth (Stripe for the payment, the crew dispatch table for who's on the boat) and refusing to guess when the record is ambiguous. An agent that confidently invents a gratuity status is worse than one that says "unconfirmed, checking."
Every diff that touches client messaging, payments, or deploy scripts now gets read by a separate reviewer pass before it's trusted, with one job: does this change carry a real production risk? Not style, not naming — does it double-send a text, leak a secret into a log line, or push a bad deploy. It's a strange thing to build, a reviewer whose entire brief is "does this let the other agent hurt someone," but on a stack where the same model writes the code that runs unattended at 3am, it's cheap insurance.
The publish pipeline for the daily content drops is split on purpose: a deterministic core decides *what* gets published, and a separate, narrowly-scoped executor turn just carries the plan out through the MCP tool calls and logs the result. No creativity in the execution step, by design. When something breaks in that pipeline, the bug is always in one of two obvious places — the decision or the execution — never smeared across both. That's the whole lesson from today, really: wherever the stakes are real — credentials, money, guest communication — narrow the surface area until the failure mode is boring enough to grep for.