Bottom line
- Revise it. This Claude account has 4% of its week left. The three 5-hour windows since the weekly reset at 21:00 on 6 Oct used 96%, about a third each, so the account gives about three windows a week. That makes Claude, not Codex, the limit. At current settings the 27 open tasks would finish around Wed 18 Nov estimate; with changes 1 to 6, around Wed 28 Oct; with those changes and a second account of the same size, around Wed 21 Oct.
- Claude's 5-hour limit hit at almost the same output every time: 724k to 762k tokens. That held with one builder or four. Four parallel builders reached it in 62 minutes, serial work in about 3 hours. On one account, fan-out buys speed inside a window, not more tasks per week.
- 39% of Claude's output was thinking, with every agent on xhigh effort. The code itself was 32%. Anthropic's default for Opus 5.5 and Sonnet 5.5 is medium.
- No task passed its first review. The three merged tasks needed 10 review rounds between them. Their fix rounds, two per task, used 21% of Claude's output. Each review round cost about 1.7% of the Codex week, at high and at xhigh alike.
- Reviews stop when Claude stops. Only the Fable orchestrator can launch Astra and Sol, so four committed tasks waited from 09:39 while 45% of the Codex week sat unused. At 10:55 a second thread on a separate quota launched all eight reviews, and they were back in 26 minutes.
- The replay shows the two quotas never worked at once. Builders ran for 274 minutes and reviewers for 103, with 0 minutes of overlap. Every merged task spent longer in review and fix rounds than in its first build: 199 minutes against 131.
01The night, replayed
The sections below are totals. This replay puts every agent back on the clock, from the planning chat at 18:15 to the last review at 11:21, so you can watch the loop work and stall. It shows something the totals hide: Claude and Codex took turns all night.
How it all started: fan-out and loops
The replay
What's on screen
- Top left, orchestrator: Fable, what it says it is doing, and the 31-task ledger. Cells turn orange while a task builds, cyan while it is in review and green when it merges.
- Middle, Claude quota: four worktrees. Each builder shows its model, its output so far and the share of that output spent thinking. The 5-hour meter fills toward the ≈730k cap; the week bar under it is an estimate.
- Right, Codex quota: Astra (code review) and Sol (checks and browser tests), their round, and their verdict. The meter is the Codex weekly limit as the rollouts recorded it.
- Bottom, swimlanes: every Fable call, builder, fix round, review and wait on one time axis, revealed as the clock moves.
Who's who. The orange mascots are Claude: the one with a crown is Fable 5.1, the orchestrator. Builders wear a hard hat, yellow for Opus 5.5 and pink for Sonnet 5.5. The white robots run on Codex: the star is GPT-6 Astra, which reviews the diff against the spec, and the sun is GPT-6.1 Sol, which runs the checks and browser tests. The pixel person is you. The Claude mascot is drawn from the Claude Code terminal logo. The GPT robot is an original drawing, not OpenAI's logo.
Before the loop. One Opus 5.5 thread at xhigh ran from 18:15 to 20:44. It researched the platform's page builder with ten sub-agents: nine on Sonnet, plus one Opus agent that built a working Puck trial on a real store in 32 minutes. It shipped a phase 1 editor, then grilled the plan through 31 questions. Out of that came the spec, the 31-task ledger and the start prompt Fable followed from 20:58.
What the replay shows
Claude and Codex never worked at the same time. In the 417 working minutes, builders ran for 274 and reviewers for 103, with 0 minutes of overlap. Fable on its own filled the other 35. Each quota sat idle while the other worked, so the night's wall clock was the sum of the two, not the longer of them.
Review and fix rounds took longer than the build. From first review to merge took 72, 78 and 49 minutes for R1-01, R1-02 and R1-03, against first builds of 27, 75 and 29 minutes: 199 minutes against 131. Of those 199 minutes, 103 were reviews and 81 were fix agents.
Sol set the pace of every round. Astra's median review took 5 minutes, Sol's 11, and 25 on R1-01's last round with browser checks. In the 7 rounds where Astra had already failed the task, Fable waited another 28 minutes for Sol before it sent the fix round.
A second quota unblocked the reviews at once. Launched at 10:55 from a second thread, all eight reviews of R1-04 to R1-07 ran side by side and were back in 26 minutes, while Claude was still out. The night's ten review rounds had run one at a time, 103 minutes in total. The Codex week went from 55% to 61%.
02What the night produced
Three tasks are merged into the integration branch, and four more are built and committed. At 09:52 nothing from this loop was in staging or production git; R1-01 to R1-03 reached staging at 11:19. The new document tables are applied to the staging database.
| Task | What it is | State | Commits |
|---|---|---|---|
| R1‑01 | Phase 1 gaps: hero hash, block insert on http, page create, delete and reorder | Merged 6 Oct 22:57 | afb6f04c0 |
| R1‑02 | Release 1 contract: page and site documents, block registry, schemas | Merged 7 Oct 03:44 | 5a873fb09 |
| R1‑03 | Migration 0119 and the store-document repository, applied to staging | Merged 7 Oct 05:03 | d48747a90 |
| R1‑03b | Test-harness hardening | Deferred | — |
| R1‑04 | Backfill script | Built 09:08, waiting for review | fee885107 800a8fd3e |
| R1‑05 | Public store renders the new format | Built 09:38, waiting for review | 6fd891c02 |
| R1‑06 | Editor on the new format | Built 09:32, waiting for review | 35628f774 9c30b64f5 |
| R1‑07 | MCP on the new format | Built 09:13, waiting for review | 47ac23831 a40c73349 825fe5bb1 |
All four builders started at 08:37 and returned their results between 09:08 and 09:38. Their worktrees are clean. Fable hit the limit at 09:39 while it was taking in those results, so the reviews wait for Claude quota. Twenty-three tasks have not started.
The 15 hours, in three blocks
Table view
| Block | Clock | Minutes | Share |
|---|---|---|---|
| Loop working | 20:58–23:49, 02:00–05:04, 08:37–09:39 | 417 | 46% |
| Waiting for the Claude limit | 23:49–02:00, 05:04–07:00, 09:39–12:00 | 388 | 43% |
| You were out | 07:00–08:37 | 97 | 11% |
| Total | 20:58–12:00 | 902 | 100% |
03The Claude cap is about 730k output tokens per window
Each 5-hour window ended at nearly the same output, while the list-price cost of the same windows rose by 70%. The cap tracks output tokens, which include thinking, much more closely than it tracks dollars.
Claude output since the loop started
Table view
| Window | Work before the limit | Builders at once | Output | Cumulative | Thinking | Cache reads | At list prices | Output per working hour |
|---|---|---|---|---|---|---|---|---|
| 21:00–02:00 | 2h 51m, limit at 23:49 | 1 | 728k | 728k | 292k | 132M | $54 | 255k |
| 02:00–07:00 | 3h 04m, limit at 05:04 | 1 | 724k | 1,452k | 298k | 138M | $63 | 236k |
| 07:00–12:00 | 1h 02m, limit at 09:39 | 4 | 762k | 2,214k | 278k | 214M | $92 | 738k |
Output when the limit hit
Table view
| Window | Output |
|---|---|
| 21:00–02:00 | 728k |
| 02:00–07:00 | 724k |
| 07:00–12:00 | 762k |
List-price cost of the same windows
Table view
| Window | Cache reads | Cache writes | Output | Total |
|---|---|---|---|---|
| 21:00–02:00 | $27 | $11 | $15 | $54 |
| 02:00–07:00 | $30 | $18 | $15 | $63 |
| 07:00–12:00 | $43 | $36 | $13 | $92 |
Same window, spent faster
Fan-out made the batch three times faster, not cheaper. The four parallel builders produced 738k output per working hour, against 236k to 255k for serial work. They hit the same cap, and then the loop sat idle for 2 hours 21 minutes. Their first builds averaged 189k output each (105k to 253k), in line with the serial first builds, so running in parallel did not make each task cost more.
For overnight runs this sets the ceiling. One Claude account gets about two windows a night, which is about 1.5M output tokens. At current settings that is three to four tasks a night estimate, however wide the fan-out. More builders only help if they draw on separate quotas.
04Where the output goes
Thinking is the biggest share at 39%, then the code itself at 32%. Most of the other suspects turned out to be small.
Claude output by what it was spent on
Table view
| Spent on | Output | Share |
|---|---|---|
| Thinking | 868k | 39% |
| Code written | 702k | 32% |
| Shell, git, tests, probes | 148k | 7% |
| Reading files | 141k | 6% |
| Orchestration | 141k | 6% |
| Scratch scripts in /tmp | 125k | 6% |
| Browser checks | 78k | 4% |
| Prose | 10k | 0.4% |
| Total | 2,214k | 100% |
| Who | Output | Share | Of which thinking | At list prices |
|---|---|---|---|---|
| First builds | 1,474k | 67% | 41% | $113.34 |
| Fix rounds | 458k | 21% | 44% | $25.46 |
| Fable orchestrator | 244k | 11% | 27% | $68.59 |
| Relays to the reviewers | 38k | 2% | 1% | $0.92 |
| Total | 2,214k | 100% | 39% | $208.31 |
Fable produced 11% of the output but 33% of the list-price cost. Its context had a median of 422k tokens and peaked at 605k. Three wake-ups after long idles rewrote 1.5M tokens of cache, so cache writes are $41.56 of its $68.59.
Large levers
- Thinking at xhigh: 868k, 39% of output.
- Fix rounds: 458k, 21% of output, plus most of the Codex use.
- Fable's long context: 33% of list-price cost.
Checked and small
- Rewrites: 145 file writes covered 138 files. Only 1% of the characters written were later overwritten.
- Builder cache misses: 3 rebuilds after waits of more than 5 minutes, 968k tokens in total.
- Test and check commands: 0.6% of output.
- Builder browser checks: 3.5% of output across 402 calls.
05Reviews and fix rounds
Each merged task needed three or four review rounds. On Claude, the two fix rounds per task cost 21% of output. On Codex, each round costs about 1.7% of the week, whatever the effort level.
| Task | Review rounds | Fix rounds | Claude output in fix rounds | How it ended |
|---|---|---|---|---|
| R1‑01 | 3 | 2 | 131k | Passed on round 3 |
| R1‑02 | 3 | 2 | 221k | Passed on round 3. Contract frozen. |
| R1‑03 | 4 | 2 | 106k | Round 3 failed on the test harness only. A scoped fourth review passed the production code. The harness work became R1-03b. |
| Total | 10 | 6 | 458k |
Each fix round is a fresh agent. It starts at about 24k tokens of context and reads its way back in. Across the fix rounds, 140k of the 458k output (31%) came before the first code change. A workflow's agent() call has no option to continue an earlier agent's conversation, so a fix round cannot pick up the builder's context.
One task, as it ran
Codex week used per review round
Table view
| Effort | Tasks | Rounds | Meter | Per round |
|---|---|---|---|---|
| high | R1-01 | 3 | 38% → 43% | 1.7 |
| xhigh | R1-02, R1-03 | 7 | 43% → 55% | 1.7 |
Codex tokens by reviewer
Table view
| Reviewer | Sessions | Input | Output | Share |
|---|---|---|---|---|
| Sol | 9 | 39.7M | 181k | 66% |
| Astra | 20 | 20.3M | 95k | 34% |
Raising the reviewers from high to xhigh did not change the cost per round. The number of rounds is what drives Codex use. At 3.3 rounds per task, the four tasks now waiting would take the meter to about 78%, and about four more tasks would empty it before Monday's reset.
06The stall: reviews wait for Claude
The two quotas take turns instead of overlapping. When Claude runs out, Codex has spare capacity but nothing can launch it.
Reviews cannot start while the orchestrator is out of Claude. Astra and Sol run on Codex, but only Fable's thread can launch them, and Fable shares the builders' Claude window. Four committed tasks have waited since 09:39 with 45% of the Codex week unused.
Builders fail at the limit instead of pausing, and nothing retries the resume. The batch launched at 05:04, just as the window ran out, and all four builders failed on their first call. Claude Code pauses a workflow at the limit only in an interactive session. It does not pause under the Agent SDK or in a background session, which matches what happened here. The automatic resume at 7:07am ended with ENOTFOUND, and nothing retried it until you restarted the loop at 8:36am.
A builder cut off by the limit is relaunched as a new agent. R1-02's first build wrote 294k from 22:57 until the 23:49 limit. The relaunched agent spent another 109k, from 02:01 to 02:25, to finish it.
07Verdict: keep the shape, change the dials and the coupling
Keep the ledger, a worktree per task, and the two independent reviewers. Change the effort levels, the review loop, and who launches reviews. That should roughly halve Claude output per task and cut Codex use per task by half or more estimate.
- Drop builds and fixes to medium effort.Thinking was 868k tokens, 39% of all output, with every agent on xhigh. Medium is Anthropic's default for Opus 5.5 and Sonnet 5.5, and its docs report Opus 5.5 at medium matching or beating Opus 5 at high on coding. Keep high for contract tasks such as R2-01. Halving the thinking would cut output by about a fifth estimate.
- Aim for one review round per task.Add the night's review findings to the builder brief as a checklist, and have the builder run Sol's checks before it commits. Allow a second fix round only for a correctness failure. Ten rounds for three tasks cost 17% of the Codex week, and the fix rounds cost 458k Claude output.
- Give each fix agent a map.Have the builder's result list the files it touched, the decisions it made, and the commands that prove the change. Pass that to the fix agent with the findings. Fix agents spent 31% of their output, 140k, before their first code change.
- Scope Sol.Run Sol on tasks with a browser-visible change. On a re-review, have it check only the items that failed. Sol used 66% of the Codex tokens in 9 of the 29 sessions.
- Let reviews run when Claude is out.Review each task as soon as it commits. Launch reviews from a quota the builders don't use: an orchestrator on the second Claude account, or builders launched as T3 threads so the orchestrator is outside their window. Today's wait is 2h 21m with four tasks ready.
- Run a smaller, fresher orchestrator.Start a new orchestrator session for each batch from the ledger, on Opus 5.5 or Sonnet 5.5 at medium with a 200k context. Keep Fable for contracts and the final report. Fable's median context was 422k, and it cost 33% at list prices for 11% of the output.
- Fan out across quotas, and schedule the resume.More builders on one account reach the cap sooner: 62 minutes for four, about 3 hours for one. Spread lanes over separate quotas, and schedule a resume at each reset that retries if it fails. The 7:07am resume failed once and nothing retried it.
- Don't hold a failed review for Sol.When Astra fails a task, start the fix round on its findings and fold in Sol's when they land. In the 7 rounds Astra had already failed, Fable waited 28 more minutes for Sol (RPL-3).
When the remaining 27 tasks could finish
| Setup | Claude per task | Codex per task | Tasks per Claude week | Finish, one account | Finish, two accounts |
|---|---|---|---|---|---|
| As is | ~440k | ~5.7% | ~5 | ~Wed 18 Nov | ~Wed 28 Oct |
| Changes 1 to 6 | ~220k | ~2.5% | ~10 | ~Wed 28 Oct | ~Wed 21 Oct |
There is room in the plan for more parallel work now. R2-01, the Release 2 contract, needs only R1-02, which merged at 03:44. Twelve tasks wait on R2-01 directly, so starting it alongside the R1d reviews is where fan-out across quotas would pay off first.
08Method and caveats
The token figures come from the session logs on this machine, deduplicated per API response. The projections are arithmetic on those figures, not measurements.
- Claude usage. Fable's Claude Code session log plus every workflow agent transcript under it. Responses are deduplicated by
message.id. Usage fields:output_tokens,output_tokens_details.thinking_tokens,cache_read_input_tokensandcache_creationby TTL. Windows are in MYT. - Output split. Thinking is exact. The other output of each response is split by the character size of its text and tool inputs. Builders write files with
cat > file <<EOFand patch scripts, so those count as code written. - List prices. From the pricing table of the earlier sub-agent fanning experiment, in USD per million tokens. Opus 5.5: $4 input, $20 output, $0.20 cache read, $5 and $8 for 5-minute and 1-hour cache writes. Sonnet 5.5: $2, $10, $0.20, $2.50 and $4. Fable 5.1: $10, $50, $0.25, $12.50 and $20. These are equivalents for comparison, not a bill.
- The cap. Three windows. Anthropic does not publish the formula. Use of claude.ai on the same account would count against the same limit and is not visible here. No other Claude Code session on this machine ran in those windows. The owner's usage page showed 4% of the Claude weekly limit left at 10:46 on 7 Oct. The last weekly reset was 21:00 on 6 Oct (a Claude Code limit message from 4 Oct gives that time), and nothing else on this machine used Claude after it, so this run used about 96% in three windows. Use of claude.ai or another device would also count. The next reset date is inferred, not read.
- Codex. 29 rollouts in
~/.codex/sessions/2026/10/0{6,7}/from 21:44 to 05:01. The meter israte_limits.primary.used_percenton the 10,080-minute window, on the Pro Lite plan. It moves in whole points, so per-round figures carry about ±0.5. The high and xhigh rounds were also on different tasks. - Docs. Effort defaults and the medium-versus-high result: model configuration. Workflow cache TTL, the
agent()options and pause-at-limit behaviour: workflows and the 2.1.291 binary. Thinking billed as output and cache reads drawing usage: costs. - Projections. Remaining work is 27 tasks: the 4 built plus 23 not started, with R1-03b left deferred. As-is cost per task is the three merged tasks' average of 405k, plus about 35k of Fable. Revised Claude cost assumes medium halves thinking on first builds (189k → ~150k), one fix round at ~50k, and ~20k of orchestration. Revised Codex cost assumes 1 to 1.5 rounds per task with Sol scoped. Windows are 730k each, at most 4.8 a day. No allowance is made for the Claude weekly limit.
- Other quotas. T3 lists a second Claude account and Cursor-billed Opus 5.5. Whether either is a separate quota from this account was not checked at 09:52. The 10:55 reviews were launched from a thread running through Cursor.
- Replay. Drawn frame by frame on a canvas from the same logs: each Claude agent's model calls (time, output and thinking tokens), each Codex session's start, end and weekly meter, the ledger's merge times, and the orchestrator thread's messages for the captions. A bar runs from an agent's first call to its last. Codex sessions are matched to review rounds by launch order against T3's delegated-task list. The replay's Claude week bar is output divided by 2.31M, an estimate; the 5-hour and Codex meters are read values. The planning prologue and the tree use T3's thread list for sub-agent times. The renderer and its derived timeline are in
docs/reports/puck-orchestration/replay/in this site's repository.
Read at 9:52am on 7 Oct 2026 from the session logs, the Codex rollouts, the ledger and the four R1d worktrees; the replay, RPL findings and update cover the run to 11:21. Rebuild with python3 ~/.claude/skills/report/scripts/build.py docs/reports/puck-orchestration/body.html -o public/reports/puck-orchestration/index.html.