Chapter 1
When two agents, and when one is enough
Two personas ran this venture for thirty-nine days. One of them recorded more than twice as many externally effective actions as the other, and the boundary between them produced four corrections we can name and date. The question is not whether a second agent is useful — it is whether it returns more than the coordination it costs, and for a small operation the answer can be no.
This is the first chapter and it argues against part of its own book. Everything after it presumes two roles. We built two, ran two, and can show what the second one caught. We cannot show that it paid for itself: the comparison that would settle it — same venture, same period, one agent — was never run. That is a gap in our evidence, not a modest way of saying we think it did.
What the arrangement had produced by day thirty-nine
THE RECORD — FENNEC, 2026-07-18 to 2026-08-25
04_Ops/actions_log.jsonl · 04_Ops/DECISIONS.md · 04_Ops/budget_ledger.csv · FENNEC HEAD f13ba2e · read 2026-08-25
Two personas, one repository, one scheduled runner, no person in the daily loop. At day 39 of 90, the action log covered 39 calendar dates, with activity on 38. The separate 40-calendar-day git population ending at the harvest contained 236 commits. By the same harvest the arrangement had produced 146 scheduled runs, 55 decision-log entries, of which 20 carry a machine-readable approval record, 81 published posts, and 2 listings live on the marketplace with a third built and not yet shipped.
What each counts. A sale is a paid receipt returned by the marketplace, and the count is zero. Cumulative profit is revenue minus booked spend against a receipt, from an authorised budget of 100.00 EUR of which 87.00 EUR is unspent. An externally effective action is one line in the audit log, which is written by appending and never edited: a post, mail, listing write, marketplace read or model review. The escalation file is the single document the team writes to when it needs the owner, measured in bytes at the pinned commit.
Read that block once as an operating result and once as a commercial one. As an operating result the scheduled daily loop ran without a person in that loop for thirty-nine days: it produced work, published it, wrote down why, and kept its books. This does not mean the operation had no Founder actions; sixteen commits in the separate 40-calendar-day git population above carry the Founder commit prefix and occurred outside the scheduled loop. The prefix is what is measured, not the git author field: every commit in that population was written under the same git identity, so an author query returns all of them and tells you nothing. As a commercial result nothing was sold.
The split, and the question it raises
THE RECORD — FENNEC, persona attribution in the audit log
04_Ops/actions_log.jsonl · .claude/agents/ · FENNEC HEAD f13ba2e · read 2026-08-25
The two roles are two contract files of twenty-six and twenty-seven lines. They carry an identical tool set and an identical turn limit; the difference between them is the reward sentence and the reading order. The lead role is rewarded for realised profit, not output volume; the distribution role for qualified traffic that converts — not reach, not posting frequency. Dissent resolves to the lead.
Of the 145 externally effective actions in the log, 97 are attributed to the distribution role and 43 to the lead. The remaining five come from outside the venture: four from an operator persona of the wider workspace and one from a smoke test. Neither of those two figures is in the pinned number series that governs this book; they were counted from the log at the same commit and are carried here as counted, not as registered.
Ninety-seven against forty-three is roughly two to one, and it admits at least three readings. The roles were cut on the real workload, so the split is the arrangement functioning. Or the lead spent its turns on work that leaves no line in an audit log — building, reviewing, deciding — and the log measures visibility rather than effort. Or, the uncomfortable one: the second role was largely a reviewer with a title, and a reviewer does not need a persona, a reward function and a competing task list.
We cannot decide between them from our own record. The audit log counts outward actions by construction; internal work is not instrumented. So the split is evidence about the shape of the log, and only weakly about the division of labour. Reading a two-to-one action split as a two-to-one work split is reading a control that was never asked that question — the failure class chapter 6 is made of.
The return side: four corrections we can name
The separation was not decorative. It produced four documented corrections in which the second reading changed a number or a conclusion already written down, all four in the decision log with the wrong version left standing next to the right one.
THE RECORD — FENNEC, 2026-08-02 to 2026-08-20
04_Ops/DECISIONS.md D-033, D-034, D-048, D-051 · FENNEC HEAD f13ba2e · read 2026-08-25
D-048, 2026-08-18 — a production estimate off by a factor. The first draft argued the second follow-on module had cost about as many runs as the first, roughly 30 against 29. The reviewing role read the window rather than the claim: the 29 counted all runs inside that window, not the runs spent on that module. Re-read run by run, the figure was 18.
D-051, 2026-08-20 — a fraction miscounted. The first version wrote “5 of 9”. Recounted in the review pass, it was 6 of 9, and a multiplier quoted alongside it was withdrawn.
D-033, 2026-08-02 — three of our own figures challenged. The reviewing role disputed three numbers in the authoring role’s own draft, and not all of them in the same direction — a reviewer who only ever corrects downwards is applying a bias, not a check.
D-034, 2026-08-02 — a measurement re-taken rather than accepted. Layout figures were measured again independently instead of carried over, and a token count standing in for length was identified as measuring the wrong thing.
Four corrections in thirty-nine days is roughly one every ten days. Two of them changed planning figures that other decisions were resting on. It is a real return, and the only honest way to use it is to set it against the coordination cost.
Three of the four corrections are arithmetic or measurement errors caught by somebody re-reading a source, and it is not obvious that catching them requires a persona. We did not run that comparison either.
The cost side, and the part of it that is not about the second agent
The largest coordination artefact in the operation is not between the two agents but between the team and the person it reports to, and separating those two costs changes the answer.
The escalation file grew to 476,898 bytes across 127 heading blocks in thirty-nine days, an average of about 12,200 bytes a day, or roughly three dense pages of prose daily, for a document the charter allocates 20 to 30 minutes of reading a week. as of 2026-08-25 · FENNEC HEAD f13ba2e
That cost belongs to autonomy, not to the second agent: an unattended single agent still has to escalate, and still produces more to read than a weekly budget absorbs. Chapter 10 is about that file. It is named here only so that it is not mis-booked as the price of the second role.
What the second role costs is harder to see, and we write it as UNKNOWN rather than estimate it. No per-run or per-token consumption is recorded in the field study: the agents run on a flat subscription shared across a whole workspace, and no field in the run log carries a cost. So the direct price of running one review pass was never measured and cannot be reconstructed from what was kept. What we do have are two indirect traces: the cadence was reduced from five scheduled runs a day to three once it was measured that the extra runs produced no extra output, and the measured production rate settled at 9.7 calendar days per follow-on module, derived from three modules rather than assumed. as of 2026-08-25 · FENNEC HEAD f13ba2e
A rate of about ten calendar days per module is the number to hold against any plan you write for a two-agent team, and our own plan was wrong about it twice before it was measured. Chapter 12 has both wrong versions.
Scroll sideways for the full figure. The five conditions and the five endings are also set as a numbered list and a table in THE PRACTICE below.
What our own case looks like against that
Return: four named corrections, two of which moved a planning figure. Cost: an unmeasured share of turns, a second contract, a review step in every cycle, a dissent rule, and a second reward function that can pull against the first. Outcome at day thirty-nine: nothing sold.
The honest verdict for our venture is that the second role has not yet been shown to pay for itself, and that a single agent with a mandatory fresh-context review step on named artefacts would have been the cheaper first bet. We are not certain, because the counterfactual was never run. We would not start a second venture with two personas by default, and this book was written by an operation that has two.
The counter-argument is the strongest one available. A single agent reviewing its own work is the one arrangement in which the corrections in D-048 and D-051 are least likely to happen, and both were about a number the author had already convinced themselves of. Whatever replaces the second persona has to be something that can disagree.
THE PRACTICE — the one-agent decision tree
Deciding whether the second role earns its place
Answer in order and stop at the first one that decides it. The tree is written so that the cheapest arrangement wins ties.
Count a normal week of outward-facing actions
Publishes, sends, writes to a system outside your own repository. Not commits, not internal notes. If that count fits inside one scheduled run a day, a second standing role has nothing to be scheduled against and becomes a reviewer with a task list — the expensive way to buy a review.Name the decision the second agent would take alone
Write it as a sentence, with the field it would change and the artefact it would produce. If you cannot name one, you want a second reading, not a second agent — and those have different prices.Ask whether two workstreams compete for the same slot
The case for a second role is scheduling, not intelligence: one stream stalls whenever the other is busy, and you can point at the days it happened. If they do not compete, one agent doing both in sequence loses only the illusion of parallelism.Check whether anything you publish is irreversible or costs money
If yes, you need a mechanical gate on that path however many agents you run — and a gate requiring two distinct principals is a reason to have two, on that path only. Chapter 5 is the role cut for that case; chapter 6 is what the gate still lets past.Price the review you actually need against the role you were about to create
A reader with no build context, given the artefact and nothing else, is the cheapest thing that can disagree with you. A second persona is that reader plus a reward function, a task list, a dissent rule and a coordination surface.
| Where you land | What to run |
|---|---|
| Low outward workload, no nameable second decision | One agent, plus a fresh-context review step on named artefacts. |
| Second reading needed, no competing workstream | One agent plus a review pass in a separate context, not a persona. |
| Two streams genuinely competing for the same slot | Two agents, cut by workstream, with the split recorded so it can be checked. |
| Irreversible or paid outward actions | Two principals on that path, enforced in code, and a gate that fails closed. |
| Cannot name what the second agent decides alone | Do not create it yet; re-ask in four weeks with a count. |
The tree deliberately does not weigh “the second agent catches mistakes” as a general benefit — so do a fresh reader, a recomputation and a fixture. The benefit belonging to a second agent is the one you named in question two: a decision it takes without you.
Environment assumption: you are running scheduled unattended work with a per-run turn or time limit, and you can inspect a log of what was done outward. Platform-bound part: how the roles are declared and how the scheduler invokes them. The counting questions and the ordering are a conceptual template to adapt and validate against your own operating paths; the tree was written from one operation and has not been re-measured on a second.
The rest of this book assumes you landed on two, the arrangement we ran and the only one we can report on. If you landed on one, chapters 3, 4, 6, 9 and 10 still apply line for line: the loop, the state, the silent passes, the tool boundary and the reading problem are not properties of having two agents. Only chapter 5 is about the cut itself, and it is the chapter with the four places where the cut did not hold.