Method · Sheet 03
Flight-hardware discipline, applied to AI agents.
I don't prompt and hope. I run agents the way a space program runs engineers: written requirements, gated reviews, traceable decisions, and verification that doesn't trust the thing being verified. That's how an RF engineer ships a mains PCB, an instrument-control app and three software products, and trusts all of them.
Eight clauses.
§ 1
Write the brief before anything else
Requirements, constraints, what "done" means and what is out of scope, written down before the first line of code or copper. An agent with a clear brief makes good local decisions; an agent without one makes confident guesses.
If I can't write the success criteria, I'm not ready to start.
In practice · MB-201
"A real, manufacturable product": schematic, PCB, BOM, fab outputs, enclosure and firmware, with no UL listing, so mains stays in a V-0 box.
In practice · MB-301
Six measurable targets, among them 90% of labels identified in under 5 s and 100% of sommelier picks from owned bottles, enforced by the server.
§ 2
Gate every phase
Work runs in phases: architecture, component selection, schematic, placement, routing. At each gate the agent stops, summarizes decisions and open risks, and waits. I approve, or I send it back. Irreversible actions are never delegated.
Placement renders are approved before a single trace is routed. Nothing is ordered or uploaded to a fab house by an agent.
§ 3
Log every decision with its evidence
Every non-obvious call is recorded: the decision, the alternatives considered, the evidence, and who accepted it. The log is how a new session (mine or an agent's) inherits judgment instead of re-arguing it, and how a later phase catches itself contradicting an earlier one.
§ 4
Give each project an operating manual
Every repo has a CLAUDE.md that is a runbook for the agent: hard invariants that must never break, how to build, test and deploy, and every gotcha that has cost real time. It's the difference between onboarding a new engineer every session and working with one who remembers.
When a bug is found, the lesson goes into the manual, not just the fix into the code.
§ 5
Give tools defined roles
Agents work through MCP servers and CLIs, each with one job. One server is the only thing allowed to write design files. Another is locked read-only as a second opinion. The vendor CLI is the source of truth when they disagree.
§ 6
Verify independently
The check can't trust the thing it checks. Math is re-derived from first principles, not by calling the implementation's own helpers. Reviews are done by separate agents with read-only access and a specific brief. AI features are scored against ground truth, not vibes.
A verification that confirms the wrong property is worse than none. It buys false confidence.
In practice · MB-101
107 independent math checks, plus a SCPI-log replay that proves every read happened where the generator was transmitting.
In practice · MB-201
Four parallel reviewer agents ran 281 checks on mains stress, low-voltage stress and pin-outs. A DRC canary plants a violation to prove the rules are live.
In practice · MB-303
Line-by-line security audit: 17 findings, 4 critical, including a real double-refund race. 16 fixed, 1 accepted.
In practice · MB-301
Eval scripts score label reads against 14 real photos with known answers, and the sommelier on latency and cost.
§ 7
Parallelize with a contract
Independent work goes to parallel agents, each with its own brief, its own files and a rule to report every touch of shared code. Reviewers fan out the same way. The merge is mine.
§ 8
Keep memory outside the model
Durable context lives in files: a notes vault with one current note per project, a runbook repo for infrastructure, and per-repo manuals. Secrets never go in; the notes record where a credential lives, not its value.
Where it went wrong, and the rule it produced.
Agents are fast. Gates are what make fast safe. These made it into the manuals.
| NCR | What happened | Corrective action | Ref |
|---|---|---|---|
| 001 | Mains buses routed at 2–3 mm, missing a 12.7 mm pour rule accepted in an earlier phase. | Check every accepted constraint against the finished layout before calling a phase done. | MB-201 |
| 002 | Seven defects at first hardware bring-up that the simulator could not show. | Simulators prove logic, not instruments. Each defect became a regression test. | MB-101 |
| 003 | A dice check that verified the sum rule and passed while the numbering was still wrong. | Write the check for the property you care about, then make it fail once on purpose. | MB-202 |