Findings

Revise the loop: the Claude weekly limit, not the 5-hour one, sets the Puck builder's pace

What one overnight agent orchestration spent, where the tokens and the hours went, and what to change so a night produces more. With a replay of the whole night.

Author
Claude Opus 5.5
Published
Oct 7, 2026
Reading time
19 min
Scope
one overnight run · 6 Oct 20:58 to 7 Oct 11:21

Bottom line

  1. Revise it. This Claude account has 4% of its week left. The three 5-hour windows since the weekly reset at 21:00 on 6 Oct used 96%, about a third each, so the account gives about three windows a week. That makes Claude, not Codex, the limit. At current settings the 27 open tasks would finish around Wed 18 Nov estimate; with changes 1 to 6, around Wed 28 Oct; with those changes and a second account of the same size, around Wed 21 Oct.
  2. Claude's 5-hour limit hit at almost the same output every time: 724k to 762k tokens. That held with one builder or four. Four parallel builders reached it in 62 minutes, serial work in about 3 hours. On one account, fan-out buys speed inside a window, not more tasks per week.
  3. 39% of Claude's output was thinking, with every agent on xhigh effort. The code itself was 32%. Anthropic's default for Opus 5.5 and Sonnet 5.5 is medium.
  4. No task passed its first review. The three merged tasks needed 10 review rounds between them. Their fix rounds, two per task, used 21% of Claude's output. Each review round cost about 1.7% of the Codex week, at high and at xhigh alike.
  5. Reviews stop when Claude stops. Only the Fable orchestrator can launch Astra and Sol, so four committed tasks waited from 09:39 while 45% of the Codex week sat unused. At 10:55 a second thread on a separate quota launched all eight reviews, and they were back in 26 minutes.
  6. The replay shows the two quotas never worked at once. Builders ran for 274 minutes and reviewers for 103, with 0 minutes of overlap. Every merged task spent longer in review and fix rounds than in its first build: 199 minutes against 131.
Tasks merged
3 / 31
4 more built, each needs a fix
Claude output
2.21M
39% thinking · 1,922 calls
Output at each limit
724k–762k
3 windows, 1 to 4 builders
At API list prices
$208
equivalent, not billed
Codex week used
38% → 55%
10 review rounds
Review rounds per task
3.3
0 of 3 passed first time
Loop working
6h 57m
46% of 15h 02m
Finish as is
~18 Nov
estimate · one Claude account

01The night, replayed

The sections below are totals. This replay puts every agent back on the clock, from the planning chat at 18:15 to the last review at 11:21, so you can watch the loop work and stall. It shows something the totals hide: Claude and Codex took turns all night.

How it all started: fan-out and loops

One person, three threads. Each branch is a model session, and each loop is a review round. A 17-second loop.
Read it left to right. The planning thread fanned out ten sub-agents, then wrote the spec, the ledger and the start prompt. Fable read them and ran four workflow batches. Each task then looped between its builder and the two GPT reviewers until both passed: three rounds for R1-01 and R1-02, four for R1-03. When Fable ran out of Claude, a second thread on another quota launched the last eight reviews.

The replay

2 min 15 s. While agents work, 1 second is about 8 minutes of the night, and the long waits are skipped. Every bar, counter and verdict comes from the session logs.
What's on screen
  • Top left, orchestrator: Fable, what it says it is doing, and the 31-task ledger. Cells turn orange while a task builds, cyan while it is in review and green when it merges.
  • Middle, Claude quota: four worktrees. Each builder shows its model, its output so far and the share of that output spent thinking. The 5-hour meter fills toward the ≈730k cap; the week bar under it is an estimate.
  • Right, Codex quota: Astra (code review) and Sol (checks and browser tests), their round, and their verdict. The meter is the Codex weekly limit as the rollouts recorded it.
  • Bottom, swimlanes: every Fable call, builder, fix round, review and wait on one time axis, revealed as the clock moves.

Who's who. The orange mascots are Claude: the one with a crown is Fable 5.1, the orchestrator. Builders wear a hard hat, yellow for Opus 5.5 and pink for Sonnet 5.5. The white robots run on Codex: the star is GPT-6 Astra, which reviews the diff against the spec, and the sun is GPT-6.1 Sol, which runs the checks and browser tests. The pixel person is you. The Claude mascot is drawn from the Claude Code terminal logo. The GPT robot is an original drawing, not OpenAI's logo.

The cast: Fable 5.1 with a crown, an Opus 5.5 builder with a yellow hard hat, a Sonnet 5.5 builder with a pink hard hat, GPT-6 Astra with a star and magnifier, and GPT-6.1 Sol with a sun and clipboard.

Before the loop. One Opus 5.5 thread at xhigh ran from 18:15 to 20:44. It researched the platform's page builder with ten sub-agents: nine on Sonnet, plus one Opus agent that built a working Puck trial on a real store in 32 minutes. It shipped a phase 1 editor, then grilled the plan through 31 questions. Out of that came the spec, the 31-task ledger and the start prompt Fable followed from 20:58.

What the replay shows

RPL-1

Claude and Codex never worked at the same time. In the 417 working minutes, builders ran for 274 and reviewers for 103, with 0 minutes of overlap. Fable on its own filled the other 35. Each quota sat idle while the other worked, so the night's wall clock was the sum of the two, not the longer of them.

first-to-last model call per builder and fix agent vs Codex rollout start and end · 20:58–09:39
RPL-2

Review and fix rounds took longer than the build. From first review to merge took 72, 78 and 49 minutes for R1-01, R1-02 and R1-03, against first builds of 27, 75 and 29 minutes: 199 minutes against 131. Of those 199 minutes, 103 were reviews and 81 were fix agents.

merge times from the ledger · review rounds matched to Codex rollouts
RPL-3

Sol set the pace of every round. Astra's median review took 5 minutes, Sol's 11, and 25 on R1-01's last round with browser checks. In the 7 rounds where Astra had already failed the task, Fable waited another 28 minutes for Sol before it sent the fix round.

10 Astra and 9 Sol sessions, 21:44–05:02 · Astra's sub-reviews excluded
RPL-4

A second quota unblocked the reviews at once. Launched at 10:55 from a second thread, all eight reviews of R1-04 to R1-07 ran side by side and were back in 26 minutes, while Claude was still out. The night's ten review rounds had run one at a time, 103 minutes in total. The Codex week went from 55% to 61%.

8 Codex rollouts, 10:55–11:21 · rate_limits.primary.used_percent
Update, 11:21. The rest of this report was written at 09:52, with four tasks waiting for review. At 10:55 a second thread on a separate quota launched those reviews. All four tasks need one fix round: Astra failed three of them, and Sol found a 500 error on one store's About page in the fourth. R1-01 to R1-03 were pushed to staging at 11:19.

02What the night produced

Three tasks are merged into the integration branch, and four more are built and committed. At 09:52 nothing from this loop was in staging or production git; R1-01 to R1-03 reached staging at 11:19. The new document tables are applied to the staging database.

TaskWhat it isStateCommits
R1‑01Phase 1 gaps: hero hash, block insert on http, page create, delete and reorderMerged 6 Oct 22:57afb6f04c0
R1‑02Release 1 contract: page and site documents, block registry, schemasMerged 7 Oct 03:445a873fb09
R1‑03Migration 0119 and the store-document repository, applied to stagingMerged 7 Oct 05:03d48747a90
R1‑03bTest-harness hardeningDeferred—
R1‑04Backfill scriptBuilt 09:08, waiting for reviewfee885107 800a8fd3e
R1‑05Public store renders the new formatBuilt 09:38, waiting for review6fd891c02
R1‑06Editor on the new formatBuilt 09:32, waiting for review35628f774 9c30b64f5
R1‑07MCP on the new formatBuilt 09:13, waiting for review47ac23831 a40c73349 825fe5bb1

All four builders started at 08:37 and returned their results between 09:08 and 09:38. Their worktrees are clean. Fable hit the limit at 09:39 while it was taking in those results, so the reviews wait for Claude quota. Twenty-three tasks have not started.

The 15 hours, in three blocks

From 8:58pm on 6 Oct to the 12:00 reset on 7 Oct. The last wait is still running at 9:52am.
Loop working Waiting for the Claude limit You were out
15h 02m
46% working
The loop worked for less than half of the 15 hours. Most of the rest was spent waiting for Claude's limit to reset.
Table view
BlockClockMinutesShare
Loop working20:58–23:49, 02:00–05:04, 08:37–09:3941746%
Waiting for the Claude limit23:49–02:00, 05:04–07:00, 09:39–12:0038843%
You were out07:00–08:379711%
Total20:58–12:00902100%
Two claims in the earlier draft of this page were wrong. The builders do not run as separate sessions: they run inside Fable's process as workflow agents. And Fable was not spending tokens while it waited: it made no model calls from 08:38 until 09:38, while the builders worked.

03The Claude cap is about 730k output tokens per window

Each 5-hour window ended at nearly the same output, while the list-price cost of the same windows rose by 70%. The cap tracks output tokens, which include thinking, much more closely than it tracks dollars.

Claude output since the loop started

Cumulative output tokens across Fable and every workflow agent, sampled every 10 minutes. The flat stretches are the waits.
Claude output
Three climbs of about 730k each. The third is the four parallel builders, and it is three times steeper than the first two.
Table view
WindowWork before the limitBuilders at onceOutputCumulativeThinkingCache readsAt list pricesOutput per working hour
21:00–02:002h 51m, limit at 23:491728k728k292k132M$54255k
02:00–07:003h 04m, limit at 05:041724k1,452k298k138M$63236k
07:00–12:001h 02m, limit at 09:394762k2,214k278k214M$92738k

Output when the limit hit

Claude output tokens per window.
21:00–02:001 builder
728k
02:00–07:001 builder
724k
07:00–12:004 builders
762k
Table view
WindowOutput
21:00–02:00728k
02:00–07:00724k
07:00–12:00762k

List-price cost of the same windows

Output, cache reads and cache writes at API list prices.
21:00–02:001 builder
$54
02:00–07:001 builder
$63
07:00–12:004 builders
$92
Table view
WindowCache readsCache writesOutputTotal
21:00–02:00$27$11$15$54
02:00–07:00$30$18$15$63
07:00–12:00$43$36$13$92

Same window, spent faster

The 02:00 window with one builder at a time, against the 07:00 window with four at once, on a shared clock. A 12-second loop.
FAN-1

Fan-out made the batch three times faster, not cheaper. The four parallel builders produced 738k output per working hour, against 236k to 255k for serial work. They hit the same cap, and then the loop sat idle for 2 hours 21 minutes. Their first builds averaged 189k output each (105k to 253k), in line with the serial first builds, so running in parallel did not make each task cost more.

wf_a6d7d0b3 journal: 4 results at 09:08, 09:13, 09:32 and 09:38 · limit at 09:39 "resets 12pm"

For overnight runs this sets the ceiling. One Claude account gets about two windows a night, which is about 1.5M output tokens. At current settings that is three to four tasks a night estimate, however wide the fan-out. More builders only help if they draw on separate quotas.

Anthropic does not publish how the limit is calculated. Its cost docs say thinking is billed as output, and that long contexts draw usage on every request through cache reads. Three windows is a small sample, but here cache reads rose 62% from the first window to the third without the cap moving.

04Where the output goes

Thinking is the biggest share at 39%, then the code itself at 32%. Most of the other suspects turned out to be small.

Claude output by what it was spent on

2.21M output tokens, all agents. Thinking is measured exactly. The rest is split by the size of each response's text and tool inputs.
Thinkingall agents on xhigh
868k 39%
Code writtenfile writes and patch scripts
702k 32%
Shell, git, tests, probes
148k 7%
Reading files
141k 6%
OrchestrationT3 calls and agent plumbing
141k 6%
Scratch scripts in /tmp
125k 6%
Browser checks
78k 4%
Prose
10k 0.4%
Code is close to fixed: it is the work. Thinking is the share that the effort setting controls.
Table view
Spent onOutputShare
Thinking868k39%
Code written702k32%
Shell, git, tests, probes148k7%
Reading files141k6%
Orchestration141k6%
Scratch scripts in /tmp125k6%
Browser checks78k4%
Prose10k0.4%
Total2,214k100%
WhoOutputShareOf which thinkingAt list prices
First builds1,474k67%41%$113.34
Fix rounds458k21%44%$25.46
Fable orchestrator244k11%27%$68.59
Relays to the reviewers38k2%1%$0.92
Total2,214k100%39%$208.31

Fable produced 11% of the output but 33% of the list-price cost. Its context had a median of 422k tokens and peaked at 605k. Three wake-ups after long idles rewrote 1.5M tokens of cache, so cache writes are $41.56 of its $68.59.

Large levers

  • Thinking at xhigh: 868k, 39% of output.
  • Fix rounds: 458k, 21% of output, plus most of the Codex use.
  • Fable's long context: 33% of list-price cost.

Checked and small

  • Rewrites: 145 file writes covered 138 files. Only 1% of the characters written were later overwritten.
  • Builder cache misses: 3 rebuilds after waits of more than 5 minutes, 968k tokens in total.
  • Test and check commands: 0.6% of output.
  • Builder browser checks: 3.5% of output across 402 calls.

05Reviews and fix rounds

Each merged task needed three or four review rounds. On Claude, the two fix rounds per task cost 21% of output. On Codex, each round costs about 1.7% of the week, whatever the effort level.

TaskReview roundsFix roundsClaude output in fix roundsHow it ended
R1‑0132131kPassed on round 3
R1‑0232221kPassed on round 3. Contract frozen.
R1‑0342106kRound 3 failed on the test harness only. A scoped fourth review passed the production code. The harness work became R1-03b.
Total106458k

Each fix round is a fresh agent. It starts at about 24k tokens of context and reads its way back in. Across the fix rounds, 140k of the 458k output (31%) came before the first code change. A workflow's agent() call has no option to continue an earlier agent's conversation, so a fix round cannot pick up the builder's context.

One task, as it ran

R1-02, the Release 1 contract, from first build to merge. Review rounds in cyan, fix rounds hatched. A 15-second loop.

Codex week used per review round

Points of the weekly meter, by reviewer effort.
high3 rounds · 38% → 43%
1.7
xhigh7 rounds · 43% → 55%
1.7
Table view
EffortTasksRoundsMeterPer round
highR1-01338% → 43%1.7
xhighR1-02, R1-03743% → 55%1.7

Codex tokens by reviewer

29 sessions from 21:44 to 05:01.
Sol9 sessions · verifier
66%
Astra20 sessions · reviewer
34%
Table view
ReviewerSessionsInputOutputShare
Sol939.7M181k66%
Astra2020.3M95k34%

Raising the reviewers from high to xhigh did not change the cost per round. The number of rounds is what drives Codex use. At 3.3 rounds per task, the four tasks now waiting would take the meter to about 78%, and about four more tasks would empty it before Monday's reset.

06The stall: reviews wait for Claude

The two quotas take turns instead of overlapping. When Claude runs out, Codex has spare capacity but nothing can launch it.

P2ORCH-1Open

Reviews cannot start while the orchestrator is out of Claude. Astra and Sol run on Codex, but only Fable's thread can launch them, and Fable shares the builders' Claude window. Four committed tasks have waited since 09:39 with 45% of the Codex week unused.

builder results 09:08–09:38 · Fable "You've hit your session limit · resets 12pm" at 09:39 · Codex meter 55% at 05:01, no review since
P2ORCH-2Open

Builders fail at the limit instead of pausing, and nothing retries the resume. The batch launched at 05:04, just as the window ran out, and all four builders failed on their first call. Claude Code pauses a workflow at the limit only in an interactive session. It does not pause under the Agent SDK or in a background session, which matches what happened here. The automatic resume at 7:07am ended with ENOTFOUND, and nothing retried it until you restarted the loop at 8:36am.

wf_a6d7d0b3 journal: 4 × started then failed at 05:04 · code.claude.com/docs/en/workflows "When a run hits your usage limit"
P3ORCH-3Open

A builder cut off by the limit is relaunched as a new agent. R1-02's first build wrote 294k from 22:57 until the 23:49 limit. The relaunched agent spent another 109k, from 02:01 to 02:25, to finish it.

wf_b56893a9: R1-02:build:opus × 2 agents, 139 and 153 calls

07Verdict: keep the shape, change the dials and the coupling

Keep the ledger, a worktree per task, and the two independent reviewers. Change the effort levels, the review loop, and who launches reviews. That should roughly halve Claude output per task and cut Codex use per task by half or more estimate.

  1. Drop builds and fixes to medium effort.Thinking was 868k tokens, 39% of all output, with every agent on xhigh. Medium is Anthropic's default for Opus 5.5 and Sonnet 5.5, and its docs report Opus 5.5 at medium matching or beating Opus 5 at high on coding. Keep high for contract tasks such as R2-01. Halving the thinking would cut output by about a fifth estimate.
  2. Aim for one review round per task.Add the night's review findings to the builder brief as a checklist, and have the builder run Sol's checks before it commits. Allow a second fix round only for a correctness failure. Ten rounds for three tasks cost 17% of the Codex week, and the fix rounds cost 458k Claude output.
  3. Give each fix agent a map.Have the builder's result list the files it touched, the decisions it made, and the commands that prove the change. Pass that to the fix agent with the findings. Fix agents spent 31% of their output, 140k, before their first code change.
  4. Scope Sol.Run Sol on tasks with a browser-visible change. On a re-review, have it check only the items that failed. Sol used 66% of the Codex tokens in 9 of the 29 sessions.
  5. Let reviews run when Claude is out.Review each task as soon as it commits. Launch reviews from a quota the builders don't use: an orchestrator on the second Claude account, or builders launched as T3 threads so the orchestrator is outside their window. Today's wait is 2h 21m with four tasks ready.
  6. Run a smaller, fresher orchestrator.Start a new orchestrator session for each batch from the ledger, on Opus 5.5 or Sonnet 5.5 at medium with a 200k context. Keep Fable for contracts and the final report. Fable's median context was 422k, and it cost 33% at list prices for 11% of the output.
  7. Fan out across quotas, and schedule the resume.More builders on one account reach the cap sooner: 62 minutes for four, about 3 hours for one. Spread lanes over separate quotas, and schedule a resume at each reset that retries if it fails. The 7:07am resume failed once and nothing retried it.
  8. Don't hold a failed review for Sol.When Astra fails a task, start the fix round on its findings and fold in Sol's when they land. In the 7 rounds Astra had already failed, Fable waited 28 more minutes for Sol (RPL-3).

When the remaining 27 tasks could finish

All rows are estimates from this night's per-task costs. They assume a Claude week gives about 2.3M output tokens, since 2.21M used 96% of it, and that this account's week resets on Tue 13 Oct at 21:00, seven days after the last reset.
SetupClaude per taskCodex per taskTasks per Claude weekFinish, one accountFinish, two accounts
As is~440k~5.7%~5~Wed 18 Nov~Wed 28 Oct
Changes 1 to 6~220k~2.5%~10~Wed 28 Oct~Wed 21 Oct
Codex no longer sets the pace: at 2.5% per task, a Codex week covers about 40 tasks. The two-account column assumes a second account of the same size with its full week free from 13 Oct; if it has room this week, the finish moves earlier. Builds that move to Codex would spend that spare quota instead of Claude's.
The effort saving has not been measured on this codebase. Run the next batch with changes 1 to 6, measure it with the same scripts, and keep the changes if each task uses at most about 250k Claude output and 2.5% of the Codex week without more review failures. If first-pass failures rise, put builders back on high, not xhigh.

There is room in the plan for more parallel work now. R2-01, the Release 2 contract, needs only R1-02, which merged at 03:44. Twelve tasks wait on R2-01 directly, so starting it alongside the R1d reviews is where fan-out across quotas would pay off first.

08Method and caveats

The token figures come from the session logs on this machine, deduplicated per API response. The projections are arithmetic on those figures, not measurements.

  • Claude usage. Fable's Claude Code session log plus every workflow agent transcript under it. Responses are deduplicated by message.id. Usage fields: output_tokens, output_tokens_details.thinking_tokens, cache_read_input_tokens and cache_creation by TTL. Windows are in MYT.
  • Output split. Thinking is exact. The other output of each response is split by the character size of its text and tool inputs. Builders write files with cat > file <<EOF and patch scripts, so those count as code written.
  • List prices. From the pricing table of the earlier sub-agent fanning experiment, in USD per million tokens. Opus 5.5: $4 input, $20 output, $0.20 cache read, $5 and $8 for 5-minute and 1-hour cache writes. Sonnet 5.5: $2, $10, $0.20, $2.50 and $4. Fable 5.1: $10, $50, $0.25, $12.50 and $20. These are equivalents for comparison, not a bill.
  • The cap. Three windows. Anthropic does not publish the formula. Use of claude.ai on the same account would count against the same limit and is not visible here. No other Claude Code session on this machine ran in those windows. The owner's usage page showed 4% of the Claude weekly limit left at 10:46 on 7 Oct. The last weekly reset was 21:00 on 6 Oct (a Claude Code limit message from 4 Oct gives that time), and nothing else on this machine used Claude after it, so this run used about 96% in three windows. Use of claude.ai or another device would also count. The next reset date is inferred, not read.
  • Codex. 29 rollouts in ~/.codex/sessions/2026/10/0{6,7}/ from 21:44 to 05:01. The meter is rate_limits.primary.used_percent on the 10,080-minute window, on the Pro Lite plan. It moves in whole points, so per-round figures carry about ±0.5. The high and xhigh rounds were also on different tasks.
  • Docs. Effort defaults and the medium-versus-high result: model configuration. Workflow cache TTL, the agent() options and pause-at-limit behaviour: workflows and the 2.1.291 binary. Thinking billed as output and cache reads drawing usage: costs.
  • Projections. Remaining work is 27 tasks: the 4 built plus 23 not started, with R1-03b left deferred. As-is cost per task is the three merged tasks' average of 405k, plus about 35k of Fable. Revised Claude cost assumes medium halves thinking on first builds (189k → ~150k), one fix round at ~50k, and ~20k of orchestration. Revised Codex cost assumes 1 to 1.5 rounds per task with Sol scoped. Windows are 730k each, at most 4.8 a day. No allowance is made for the Claude weekly limit.
  • Other quotas. T3 lists a second Claude account and Cursor-billed Opus 5.5. Whether either is a separate quota from this account was not checked at 09:52. The 10:55 reviews were launched from a thread running through Cursor.
  • Replay. Drawn frame by frame on a canvas from the same logs: each Claude agent's model calls (time, output and thinking tokens), each Codex session's start, end and weekly meter, the ledger's merge times, and the orchestrator thread's messages for the captions. A bar runs from an agent's first call to its last. Codex sessions are matched to review rounds by launch order against T3's delegated-task list. The replay's Claude week bar is output divided by 2.31M, an estimate; the 5-hour and Codex meters are read values. The planning prologue and the tree use T3's thread list for sub-agent times. The renderer and its derived timeline are in docs/reports/puck-orchestration/replay/ in this site's repository.

Read at 9:52am on 7 Oct 2026 from the session logs, the Codex rollouts, the ledger and the four R1d worktrees; the replay, RPL findings and update cover the run to 11:21. Rebuild with python3 ~/.claude/skills/report/scripts/build.py docs/reports/puck-orchestration/body.html -o public/reports/puck-orchestration/index.html.