A good fit if you
- already run an agent team on a schedule and want the operating half
- need to tell a control that held from one that only looked like it did
- would rather read one operation measured honestly than ten described vaguely
Field report · Published under Adrian Verdan · English only
Thirty-nine logged days of a two-agent operation that publishes, books, measures and reports through a scheduled daily loop — and the fifteen things counted as coverage: fourteen green lights and one unused capability, that still missed the question that mattered.
29.90 € excl. VAT
VAT depends on your country; the final price is shown at Gumroad checkout before purchase. Merchant of Record: Gumroad.
Immediate download after purchase.
Created with AI assistance and reviewed by a human publisher who holds editorial responsibility — disclosed in the spirit of EU AI Act Art. 50; details inside the product.
Across fifteen cases in one small, carefully watched operation, fourteen controls supplied a green or pass-like signal while the question that mattered sat outside what they enforced. In the fifteenth, a merely available capability was treated as coverage even though it was not invoked correctly during the documented nine-day period.
Each case is written up with what produced the assurance, what lay outside it, what got through and what changed. The four-class taxonomy is designed to help you test analogous failures in your own setup; its performance in another operating environment was not measured.
Not choosing a framework, not wiring the first agent — that material is abundant, current and free, and reprinting it at length would be charging you for it. This is about day twenty-two, when the loop has been running for three weeks, the apparatus is supplying green or pass-like assurances, the weekly report says red, and you have to decide whether it is telling you the truth.
Those sixteen founder commits occurred outside the scheduled daily loop; the operation as a whole was not unattended. Both halves of that are the subject.
If you already run a two-agent team, this book gives you a tested daily-loop method, checklists and fixtures for examining whether its controls held or merely looked as if they did. Transfer to a different runner, scheduler or operating environment was not measured. It will not make you money on its own. Ours has not.
The internal release verification checks seventy-seven selected figures: sixty-two are recomputed from shipped inputs, and fifteen are transcription comparisons between printed book values and shipped records. It exits non-zero if one of those comparisons deviates. It is evidence about this edition, not a delivered repository or tool promise, and it does not verify every number in the book — every run names five groups outside its reach.
If a figure in the book and the shipped data disagree, the data is right and the book is stale. That is the correct direction, and it is why every number carries its date.
What this looks like in practice, with sources: what the pipeline rejected →
CHARTER.md, autonomy_matrix.xlsx, cycle_prompt_pack.md, DECISIONS_format.md, weekly_report_template.md, field_study_numbers.csv and field_study_numbers.xlsxFifteen chapters as a 190-page tagged PDF and as an offline HTML bundle, three calculators that run in your browser, and exactly seven editable example files — two of them XLSX workbooks, plus the complete number series as XLSX and CSV. Everything works offline.
No. It is one operation, measured on one day, written down while it was still running. Where a lesson depends on the specific stack it ran on, the template that carries it says so on its own line.
The field study was read on day thirty-nine of ninety. Figures that move are marked as moving, and every number carries its own date and commit. A day-90 addendum follows after 2026-10-15 and is free to anyone who buys this edition.
29.90 € excl. VAT
VAT depends on your country; the final price is shown at Gumroad checkout before purchase. Merchant of Record: Gumroad.
If you want the design side — when one model is enough, reviewer patterns, routing, cost — that is Orchestrating AI Agents. This one is the operating half: it presumes a team exists and is about keeping it honest while it runs.
English only. No German edition is planned before there is a demand signal: zu den deutschen Produkten →