Fix Multi-Agent Run Hanging and Finalization Issues

This is stuck for way too long and there's just no indicator what the final agent is still doing. Maybe it's also broken completely? Also Status says "done" but footer says 3 succeeded and 1 failed?
AGENT STATUS DURATION TOKENS PREVIEW claude-code:opus (… ✓ done 3m32s 1196.8k… I have enough to finalize. The code is well-structured and well-tested; cross-f… claude-code:sonnet … ✓ done 5m29s 650.3k/… Now I have everything I need. Let me write the final review. --- **Request chan… codex ✓ done 4m22s 1496.4k… request changes - review gates can pass without enforcing configured signals, a… codex:gpt-5-4-mini … ✓ done 16m43s 3612.0k… request changes, because the new reviews gate has two fail-open paths that can …
4 agent(s) done — 3 succeeded, 1 failed, 0 cancelled Finalizing output...
Now I understand the architecture. Let me look at the test that enforces the tick-freeze and the finalize footer test:
Now let me implement the fixes. First, the model struct and the status-sync bug:
Now the runFinishedMsg handler — set finishedAt and fix the status-sync mismatch:
Now update the tick handlers to keep animating during finalization, and the finalize footer to show a live spinner + elapsed time:
Now update the test that enforced the tick-freeze to reflect the new finalize-animation behavior:
Tests pass. Let me add a focused test for the status-sync fix (the footer/row mismatch the user reported):
Both issues are fixed. Here's a summary of what was wrong and what I changed in cmd/entire/cli/review/tui_model.go:
1. "done" rows but footer says "1 failed" (the disagreement)
The row status and the footer counts come from two different sources that could diverge:
- Each row status is updated live from the event stream (
Finished{Success:true}→✓ done). - The footer counts (
countsLine) are computed from the orchestrator's finalRunSummary, which sees process-level signals the stream never emits (e.g. a non‑zero exit code after the agent already streamedFinished{Success:true}).
The old runFinishedMsg sync only overwrote a row when it was still Unknown, so an optimistic stream Succeeded was never corrected. Result: row shows ✓ done, footer counts it as failed.
Fix: the summary is now authoritative for terminal classification — a summary Failed overrides a stream Succeeded (and Cancelled does too, while a real stream Failed/RunError stays sticky over a blanket cancel). Now the rows and the counts line always agree.
2. "Finalizing output…" looks frozen / no indicator
The tick loop deliberately stopped after finish, so during the finalize window the Finalizing output... line was completely static — indistinguishable from a hang. (That window is exactly when the post-run synthesis/judge and buffer flush run before PostRunComplete quits the program.)
Fix:
- Ticks keep running while finalizing, so the spinner animates.
- The footer now renders a live spinner plus an elapsed timer:
⠙ Finalizing output... (12s). A normal short finalize is obviously alive, and a genuine stall is now visible as a climbing timer instead of a dead-looking screen.
I repurposed the old TestTUIModel_FinishedStopsTickRedraws into TestTUIModel_FinishedKeepsTicking and added two tests covering the status-sync fix. All review package tests pass.
Note: if it really is hanging "way too long", the timer will now show you that, and the likely culprit is the post-run synthesis/judge LLM call — PostRunComplete has a 2s grace then Kill, so it shouldn't hang forever; if it does, that's a separate issue in the synthesis sink worth a follow-up.
Let me empirically test whether glamour rendering of large agent output is the bottleneck:
Confirmed the root cause. Glamour rendering is super-linear and effectively hangs on large output:
| input | render time |
|---|---|
| 500 KB | 1.1 s |
| 2 MB | 6 s |
| 4 MB | 49 s |
| 8 MB | 4m30s |
| 16 MB | didn't finish (>5 min) |
A long-running agent (the codex worker: 16m43s, huge output) produces multi‑MB of narrative markdown. DumpSink.RunFinished renders each agent's narrative through glamour before synthesis starts, on the orchestrator's dispatch goroutine — so the whole finalize path wedges and the TUI sits on a frozen "Finalizing output…" forever. Let me confirm the callers and fix it centrally:
All rendering funnels through Render. I'll add the size guard there so every caller (dump, synthesis, investigate, dispatch) is protected:
Now add a regression test for the guard:
Let me verify the full picture builds and vets cleanly:
Investigation: why the run wedges on the entire.io repo
Root cause — found and reproduced
It's not a deadlock; it's a super-linear blowup in glamour markdown rendering during the post-run finalize phase.
The finalize sink chain for a TTY multi-agent run is:
All of this runs synchronously on the orchestrator's dispatch goroutine (RunMulti's fan-in loop). DumpSink calls mdrender.Render → glamour on each agent's full narrative. The codex worker ran 16m43s and produced a multi‑MB narrative. I measured glamour's cost curve directly:
| input | render time |
|---|---|
| 500 KB | 1.1 s |
| 2 MB | 6 s |
| 4 MB | 49 s |
| 8 MB | 4m30s |
| 16 MB | didn't finish (>5 min) |
So a few MB of agent output ≈ minutes-to-effectively-never. While that's running, PostRunComplete never fires, the TUI never exits, and the dashboard sits on a frozen "Finalizing output…" — exactly the symptom. (This also explains why it shows "Finalizing output…" and not "Final judge is consolidating…": it never gets as far as starting synthesis.)
Fix
Added a size guard in the one choke point everything funnels through — mdrender.Render (cmd/entire/cli/mdrender/mdrender.go):
Above the cap it returns the raw markdown unchanged (the same fallback already used on render errors). Content is fully preserved, just unstyled, and finalization stays bounded. This protects every caller: DumpSink, SynthesisSink, investigate, and dispatch_tui.
Combined with the earlier UI change (live spinner + climbing timer on the finalize line), a genuinely slow finalize is now both prevented and visible rather than looking dead.
Tests
TestRender_OversizedInputReturnsRawQuickly— 8 MB input must return raw within 5 s (fails without the guard).- All
mdrenderandreviewtests pass;go build ./...andgo vetclean.
Note / possible follow-up
There's a second, less-certain latent hang vector in RunIsolatedTextGeneratorCLI: it pipes child stdout into a bytes.Buffer, so if a synthesis CLI (claude/codex/gemini) spawns a grandchild that survives the 2‑min context SIGKILL and keeps the stdout pipe open, cmd.Run() can block past the timeout. Not what bit you here (you hung before synthesis), but worth hardening separately if you want me to.
The bug
When a review run finishes, the CLI prints each agent's output as nicely formatted markdown (via a library called glamour). That formatting step is really slow on big inputs — it gets exponentially worse, not linearly:
- 2 MB of text → 6 seconds
- 4 MB → 49 seconds
- 8 MB → 4.5 minutes
- 16 MB → basically never finishes
One of your agents (the codex worker, running 16+ minutes) produced several MB of output. Formatting it ran on the same thread that finishes the whole run, so everything froze waiting on it. The screen stayed stuck on "Finalizing output…" with no progress, because the next steps (the final judge, then exiting the UI) never got a chance to run.
The solution
Two changes:
-
Cap the formatting size. If the text is bigger than 256 KB, skip the fancy formatting and just print it plain. Plain text is still perfectly readable, and it's instant instead of taking minutes. (256 KB formats in under a second.)
-
Make the wait visible. While finalizing, the footer now shows a spinning indicator and a climbing timer, so you can tell it's actually working — and if something ever does get genuinely stuck, you'll see the seconds tick up instead of staring at a frozen screen.
I also fixed a smaller, separate display bug: an agent could show "✓ done" in its row while the footer counted it as "failed." Now the final summary is the source of truth, so the row and the footer always agree.
Mh... we don't really care about the formatting though? Like the "workers" should be as efficient as possible?
You're right, and that's a better framing than mine.
The per-worker output is raw material, not a deliverable. The thing a human actually reads is the final synthesized report. The workers' text exists to (a) feed the final judge and (b) be inspectable on demand (Ctrl+O drill-in already shows it). So running every worker's full narrative through glamour during finalize is just wasted work — we're spending CPU prettifying output nobody needs prettified.
So my 256 KB cap is treating the symptom. The cleaner fix is: don't glamour-render worker dumps at all — emit them as plain text. Only the final report (small, bounded, one document) gets the nice formatting.
That gives you:
- No expensive rendering on the finalize path, regardless of how big a worker's output is.
- Worker output still fully available (plain in the dump, styled-or-not in drill-in).
- The cap becomes a cheap belt-and-suspenders safety net rather than the actual mechanism.
Two small caveats worth a decision:
- Worker dumps would no longer be colorized in the terminal. Given they're secondary to the final report, that seems fine — but it is a visible change.
- If a worker emits many MB, we'd still print that much plain text to the terminal (fast, but noisy). If you care, we could also truncate the dump itself with a "… (N KB omitted, use drill-in)" note.
Want me to switch the worker DumpSink to plain text (drop the glamour call) and keep the cap as a safety net? And should I add dump truncation too, or leave full output?
Let me look at the prompt files before making changes, so I can adjust prompting coherently:
Now I'll make the changes. First, switch DumpSink to plain text (drop glamour):
Now update the cmd.go construction sites:
Now update the tests. First cmd_test.go (remove DumpSink.RenderWriter assertions, keep the SynthesisSink one):
Now check dump_test.go for any reliance on rendering behavior, and update the stale comment:
Now adjust the worker prompt to keep output efficient (concise, no verbose dumps), since it's no longer styled and bloated output was the underlying cost:
All changes are in and green. Here's what I did:
1. Workers no longer get prettified (the real fix)
DumpSink now writes plain markdown straight to the output — it no longer calls glamour at all (cmd/entire/cli/review/dump.go). Worker output is raw material; the human reads the styled final report, and can drill into a worker's buffer interactively. Rendering every worker's narrative was wasted work, and that waste is exactly what wedged the run.
- Removed the now-pointless
RenderWriterfield and updated the two construction sites incmd.go. - The final synthesized report still gets glamour styling (small, bounded, one doc).
2. Kept the size cap as a safety net
mdrender.MaxRenderBytes (256 KB) stays. It no longer fires on the worker path, but protects the remaining glamour callers (final report, investigate, dispatch) from the same super-linear blowup.
3. Prompting: tell workers to be lean
Added to the worker output instructions (prompt.go):
Keep the report compact: quote only the minimal relevant snippet (a few lines) per finding. Never paste whole files, full diffs, or large logs, and skip decorative formatting like tables or ASCII art — it wastes effort and is not rendered.
This attacks the root inefficiency — a worker dumping multi-MB output — rather than just coping with it downstream.
Net effect
- The finalize path does no expensive rendering regardless of worker output size → no more frozen "Finalizing output…".
- Workers are nudged to produce smaller, cheaper output.
- The cap + the earlier live spinner/timer remain as defense-in-depth.
One thing I left as-is per the ambiguity earlier: the dump still prints a worker's full output (just plain/fast now). If you'd rather truncate huge dumps with a "… (N KB omitted — use drill-in)" note, say the word and I'll add it.
Auth works. Let me format, create the branch, and commit the 9 changed files (excluding the untracked local .pi extension):
Pushed. Now create the trail for this branch:
Done. Everything is in place:
- Branch:
20260629-fix-review-finalize-hang(created, committed, pushed with upstream tracking) - Commit:
2b104c02d— the 9 changed files (the untracked local.pi/extensions/entire.tswas deliberately left out) - Trail: #686 → https://entire.io/gh/entireio/cli/trails/686
The trail and commit both summarize the fix: plain-text worker dumps (no glamour), the 256 KB render cap as a safety net, the live finalize spinner/timer, the row/counts status-sync fix, and the leaner worker prompt.
You can open a PR at: https://github.com/entireio/cli/pull/new/20260629-fix-review-finalize-hang
This was the error from the broken review earlier. Just validate we fixed the right things: " Generating final report with Claude Code... final report unavailable: context canceled
This error changes the picture meaningfully. Let me check it against what we actually fixed, starting with the persistence path the error mentions:
This error is really useful — it tells us our fix was partially aimed at the right place, but missed the most painful bug. Let me trace it precisely.
The output starts at Generating final report with Claude Code.... That line is printed by SynthesisSink.RunFinished before it calls the LLM — and the per-worker DumpSink runs before synthesis. So this run got all the way through DumpSink and into the synthesis step. That means:
-
Glamour was not the final blocker in this particular run. Our DumpSink fix is still correct and defensive (it removes a real super-linear hang, and matches the earlier "Finalizing output…" screenshot, which is the DumpSink window), but the thing the user was staring at when they gave up here was the Claude Code synthesis call, not glamour.
-
Everything failed with
context canceled— notdeadline exceeded. That's the signature of the user pressing Ctrl+C, not the 2‑minute synthesis timeout firing. So: synthesis hung → user cancelled →runCtx/ctxcancelled →Synthesizereturnscontext canceled. -
The findings were thrown away — and this is the real bug we did not fix:
Confirmed the mechanism. Here's the honest validation:
What this error reveals
| Symptom | Cause | Did we fix it? |
|---|---|---|
| Earlier "Finalizing output…" freeze (the screenshot) | DumpSink glamour render of huge worker output, super-linear | ✅ Yes (plain-text dump + cap) |
| This trace — stuck on the final report | Claude Code synthesis call hanging; user Ctrl+C'd out | ⚠️ Only indirectly — bounded by a 2-min timeout, and now visible via the spinner/timer, but not made faster |
final report unavailable: context canceled | Synthesis correctly aborts on the user's cancel | Working as designed |
findings were not persisted … git rev-parse: context canceled | Persistence reuses the cancelled run context | ❌ No — and this is the worst part |
The bug we missed
writePostReviewManifest(ctx, …) reuses the run context. The if summary.Cancelled guard doesn't catch this case, because:
- the workers finished normally →
summary.Cancelled == false, - then the user cancelled during the hung finalization →
ctxbecomes cancelled, - so the guard passes, but the
git rev-parseinside the write runs on a dead context and fails.
Net result: a 16-minute review where the workers succeeded loses all its findings because the user bailed on a stuck final-report step. That's the most damaging behavior in the whole trace, and our changes don't touch it.
Proposed fix
Persist with a detached context so already-produced findings survive a cancel during finalization:
…and thread writeCtx through localReviewManifestFromCurrentState / writeLocalReviewManifest. (We keep the existing summary.Cancelled early-return so a genuinely cancelled run still skips.)
So: our pushed fixes are correct and worth keeping (the glamour hang was real and is the likely cause of the earlier frozen screenshot + the visibility improvement helps), but they would not have saved this run's findings. Want me to add the detached-context persistence fix to the same branch/trail?
Both imports present. Now apply the detached-context fix:
Now let me check for existing tests around this to add a regression test:
The existing test reaches the same git path. I'll add a regression test that passes an already-cancelled context (the exact scenario) and verifies persistence still proceeds rather than failing with context canceled:
Let me verify it genuinely catches the regression (temporarily bypass the fix), then run the full suite:
The test fails without the fix and passes with it. Now full validation, then commit/push/trail:
Now update the trail to capture this additional finding:
trail update replaces the body, so I'll fold in the new fix:
Done. The findings-persistence fix is implemented, tested, committed, pushed, and the trail is updated.
What I added
The fix (cmd/entire/cli/review/cmd.go): writePostReviewManifest now detaches from the run context before doing any git/disk work:
So a Ctrl+C on a slow final-report step no longer throws away findings the workers already produced. The existing summary.Cancelled guard still skips persistence for a genuinely cancelled run.
The regression test (manifest_test.go): TestWritePostReviewManifest_SurvivesCancelledRunContext passes an already-cancelled context and asserts persistence proceeds (reaches the normal path) instead of failing with context canceled. I verified it fails without the fix and passes with it.
State of the branch 20260629-fix-review-finalize-hang
2b104c02d— finalize hang (plain-text dumps + render cap + TUI timer + status-sync)80dd3d82a— findings persistence on cancel- Pushed; trail #686 updated with both root causes → https://entire.io/gh/entireio/cli/trails/686
Full review + mdrender suites, go build ./..., and go vet are all green.
One honest caveat I'll restate: this makes the findings survive the cancel, but it doesn't make the Claude Code synthesis call itself faster — that's still bounded only by its 2‑minute timeout. If staring at a 2‑minute "Final judge is consolidating…" is itself the complaint, the next lever would be lowering that synthesis timeout or making it cancellable independently of the run (so you can skip just the final report and still keep everything else). Say the word if you want that too.
Let me cross-check the trail body against the actual diff to confirm nothing's missing:
Mapping each changed file to the trail body:
| File(s) | Trail item |
|---|---|
dump.go, cmd.go (RenderWriter removal), dump_test.go, cmd_test.go | plain-markdown DumpSink ✅ |
mdrender.go + test | 256KB render cap ✅ |
cmd.go (detached ctx), manifest_test.go | findings persistence on cancel ✅ |
tui_model.go + test | spinner/timer + status-sync ✅ |
prompt.go | compact worker output ✅ |
So all five code changes are documented. The one thing missing is the open caveat I flagged to you verbally but never recorded: the synthesis call is still only bounded by its 2‑minute timeout — not fixed, just made survivable. That's exactly the kind of context a trail should carry, so let me add it:
Now yes — the trail body is complete. It documents:
- Both root causes with the evidence (glamour cost curve; the real
context cancelederror trace) and the fix for each. - All three supporting changes (TUI timer, status-sync, worker prompt).
- The one open caveat I'd only mentioned verbally — the synthesis 2‑minute timeout is still the limiting factor and is not fixed, plus the two follow-up options.
- Commit SHAs tying it to the code.
Every changed file maps to a documented item (test files are implied by their corresponding fix). The trail now fully ties together the context for the branch.
The comments I added are verbose. Let me trim them to minimal "why" notes.
Now trim the verbose test comments I added:
All green. Let me do a final review of remaining added comments:
That diff is committed-vs-main (pre-trim); my trims are still uncommitted. Let me confirm the working-tree trims are correct and commit:
Trims confirmed — net 82 comment lines removed, 25 added. Committing as a cleanup on the same branch:
Done. Comments are now trimmed to minimal "why" notes across all changed files — a net reduction of ~57 comment lines (82 removed, 25 kept). Each remaining comment is one or two lines stating only the non-obvious rationale (the glamour cost curve, the detached-context reason, the status-downgrade rule, etc.).
- Build,
go vet,gofmt, and the fullmdrender+reviewtest suites all pass. - Committed as
d120107aeand pushed to20260629-fix-review-finalize-hang.
Branch now has three commits: the finalize-hang fix, the findings-persistence fix, and this comment cleanup.
When Ctrl+C ones, it will hang forever it seems
"AGENT STATUS DURATION TOKENS PREVIEW claude-code:opus (… ✓ done 3m10s 746.6k/… I've now verified the candidates directly against the surrounding code (the `tr… claude-code:sonnet … ✓ done 7m59s 1530.4k… I have all the information needed for the final report. --- **Approve with nits… codex ✓ done 4m9s 1516.5k… request changes — two gate configs/API paths can falsely block or falsely pass … codex:gpt-5-4-mini … cancelling 18m13s I’ve ruled out a few apparent issues that are actually intentional by the evalu…
Cancelling agents... · Ctrl+C again: force quit"
The output I got:
dip@dip entire.io % entire review Reviewing claude/repo-gates-api-config-67aura vs origin/HEAD: 5 commits, 23 files changed, 0 uncommitted # claude-code:opus (claude-code, model opus) review
Failed: review agent codex:gpt-5-4-mini (codex, model gpt-5.4-mini) timed out after 10m0s
The version was the one before the latest change
This is a different failure than before — let me investigate the cancellation path. The agent stays "cancelling" forever, which means Ctrl+C isn't actually terminating the subprocess. Let me trace how agent processes respond to context cancellation:
This is the classic Go exec pipe-inheritance deadlock. Let me confirm the mechanism: agent processes use exec.CommandContext with StdoutPipe, and the template reads Events() (the pipe) to EOF before calling Wait():
The repo already solved this exact problem elsewhere. Let me look at the established pattern:
There's a proven terminateOnCancel (WaitDelay + process-group kill) pattern. The agent ReviewerTemplate doesn't use it — that's the bug. Let me check whether investigate shares the spawn path, to place the fix correctly:
Investigate has the same vulnerability (cmd.Run() with discarded pipes). Both spawn agents via spawn.Spawner.BuildCmd. Let me check the existing external-command WaitDelay value for consistency, then create a shared leaf helper:
Now I understand the full picture. The fix is the established WaitDelay + process-group-kill pattern, which the agent spawn path is missing. I'll create a shared leaf helper (used by both review and investigate) rather than duplicate. First the package:
Now wire it into the review template and the investigate loop:
Now a regression test proving the grandchild-holding-pipe scenario unblocks on cancel:
Test passes instantly. Let me verify it genuinely catches the bug (remove the fix temporarily) and run the affected packages:
Confirmed — without the fix it hangs (15s) and fails; with it, instant. Now run the affected packages:
All green. Let me commit, push, and update the trail:
Found and fixed the real cause of the Ctrl+C hang.
What was happening
codex:gpt-5-4-mini stuck at "cancelling" (18m, past its 10m timeout) is the tell: cancellation wasn't terminating the agent.
Agent CLIs like codex spawn helper grandchildren (sandbox, MCP servers) that inherit the agent's stdout pipe. exec.CommandContext's default cancel only SIGKILLs the agent process itself — the grandchildren survive and keep the pipe open. The reviewer template drains Events() (the stdout pipe) to EOF before calling Wait(), so:
Neither the 10-minute reviewer timeout nor Ctrl+C could break it, because both rely on context cancellation, which only killed the top process. Investigate's cmd.Run() had the same exposure.
The fix
New leaf package procutil with TerminateOnCancel(cmd) = process-group SIGKILL on cancel + a WaitDelay backstop — the same proven pattern the repo already uses for git transport helpers (checkpoint/remote). Applied on the agent spawn path in both the review template and the investigate loop. Now cancel kills the whole tree, the pipe closes, and the run unwinds immediately.
Regression test (procutil_unix_test.go) reproduces a grandchild holding the pipe: without the fix it hangs 15s+ and fails; with it, 0.01s.
Branch 20260629-fix-review-finalize-hang now has four fixes
2b104c02dfinalize hang (plain dumps + render cap + TUI timer + status-sync)80dd3d82afindings persistence on canceld120107aecomment trim0f184b3c6process-group kill on cancel ← this hang
All green (procutil, review/..., investigate/...), trail #686 updated.
Note on the output you pasted: the write to disk failed: context canceled line is the pre-80dd3d82a findings-loss bug — you were running a build before that fix landed. With the current branch, both the hang and the findings loss are addressed.
Lint is failing. Let me get the details — first check the CI log, then reproduce locally:
Two errcheck issues in my test (the repo's errcheck flags blank assignments too). Let me fix them, then run the linter locally across all my changes:
Lint clean. Let me commit the fix, then merge latest main:
Clean merge, no conflicts. Now verify everything builds, lints, and tests pass post-merge:
Build, full lint, and tests all clean. Pushing:
Pushed. Let me confirm the remote checks pick up the new commits:
New CI run triggered. The previously-failing lint is re-running. Let me wait for it to resolve:
Lint now passes. Done.
Summary
Fixed the failing remote check: CI lint was red on two errcheck violations in my new test (procutil_unix_test.go) — the repo's errcheck flags blank-assignment ignores too. Added //nolint:errcheck with rationale on the two intentional ignores (the drain read whose error-on-close is expected, and the post-kill cmd.Wait). Verified locally with the exact CI linter (golangci-lint v2.11.3): 0 issues across ./cmd/entire/cli/....
Merged latest main: fetched and merged origin/main (was 6 behind) — clean merge, no conflicts (main's changes were in trail/gitremote areas, disjoint from this branch's review/procutil/mdrender changes).
Post-merge validation: go build ./..., full golangci-lint, and the affected test packages (procutil, review/..., investigate/..., mdrender) all pass.
CI now: lint ✅ pass (was the failure), binary-size ✅, the rest (test-core, test-integration a/b/c, test-canary, license-check) running on the fresh commit f98758d0d.
Branch 20260629-fix-review-finalize-hang is now up to date with main and green on lint. The remaining test jobs are still in progress — let me know if you want me to keep watching them through to completion.