Fix Codex Subagent Transcript Resolution Path

Claude Code·Opus 5.5·Soph·3d ago·57min·3 Checkpoints·7 file changes·+237/-20·37.9K tokens
3d ago·5m

I fixed this, but the cause is slightly different from the bug report, and it's on a fixed premise: the commit wasn't an ordinary one made after the turn. The parent agent made it itself, mid-turn. I matched timestamps in the two rollouts in ~/.codex/sessions/2026/10/03/:

  • 15:06:58: the child's task_complete.
  • 15:07:06: the parent runs git commit -m 'Add count command'.
  • 15:07:08: the parent's task_complete, which is when its Stop hook runs.

So condensation ran before refreshCodexInventory had ever run. That refresh is the only thing that finds the child's rollout. The rest of the report holds: the task record had no path, and the fallback only knows Claude's directory layout. A commit made after the parent's Stop would normally already have the path.

Fix (cmd/entire/cli/strategy/manual_commit_condensation.go): I added resolveInventoryTaskTranscripts, which materializeTaskRecords calls once per condensation.

  • It collects records that have no declared path and builds lookup entries from the session's SubagentInventory. That inventory is the authoritative list of child agents, so nothing gets guessed from filenames.
  • It then calls the agent's existing agent.ExtractWithSubagentInventory. That reuses the Codex resolver, which accepts a rollout only when its session_meta.id matches the agent ID, with the same scan limits as the turn-end refresh. Going through the agent interface means the strategy package doesn't need to import the cli package.
  • The path it finds is tried after the declared path and before the Claude-layout fallback. Agents without this extractor, such as Claude, do nothing new.
  • Nothing is written back to session state. The record stays in flight, so later condensations still handle it exactly as before.

Test: TestMaterializeTaskRecords_CodexRecordWithoutPathResolvesThroughInventory builds the state the hooks leave before any refresh: an inventory entry and a record, both without a path. It checks two cases:

  • A rollout whose session_meta.id matches the agent is stored as the transcript.
  • A rollout whose session_meta.id names a different agent is rejected and stays transcript path unresolvable.

With the fix taken out, the first case fails with exactly the error from the report.

Verification: mise run fmt, go vet, and go test on the strategy and codex packages pass, and mise run lint reports 0 issues. I haven't run the full mise run check, and nothing is committed.

One cost to know about: when the rollout isn't at a known path, the post-commit hook now does the same limited scan of ~/.codex/sessions that Stop already does. That only happens when a Codex task record has no path.

3d ago·2m

I fixed this finding. A record now goes through the inventory resolver when its declared path is empty or the file no longer exists.

Change (manual_commit_condensation.go, in resolveInventoryTaskTranscripts): the check uses the package's existing fileExists. A record whose declared path still exists is not passed to the resolver, so the usual case reads no rollout and does no scan at commit time. For a stale path, the Codex resolver tries the inventory's paths first. If those fail too, it does its scan by session_meta.id, and that scan already covers archived_sessions, so an archived rollout is found.

Regression test: TestMaterializeTaskRecords_CodexRelocatedRolloutResolvesThroughInventory records a completed task whose declared path (and the inventory's resolved path) point to sessions/…, while the rollout actually sits in archived_sessions/…. It checks that the transcript gets stored. Without the change it fails with transcript unreadable, which is how the stale-path case actually shows up.

One gap left: a declared file that exists but can't be read still skips the inventory lookup. I chose this on purpose. Covering it would mean reading the file twice, or resolving on every condensation even when nothing is wrong.

Verification: the strategy package tests pass after mise run fmt, and mise run lint reports 0 issues. Nothing is committed yet, and I haven't run the full mise run check.

can you open a trail

3d ago·5m

I opened the trail: https://entire.io/gh/entireio/cli/trails/1470 (trail #1470).

  • Branch: I renamed this worktree's branch to soph/codex-subagent-inventory-transcript and pushed it.
  • Commit: one commit with both changes (the inventory lookup for records with no path and for paths that no longer exist) and the two regression tests.
  • Trail description: it covers the cause, using the timestamps from the real run, what the fix does, and the gap left open. That gap is a declared file that exists but can't be read, which still skips the inventory lookup.
  • mise run check: everything passed except TestOpenCodeSeedRepoPlantsDeps in e2e/agents. It fails locally because it needs the opencode bootstrap that a local run doesn't do, and it's unrelated to this change. I said so in the description.
  • Findings: entire trail finding list 1470 shows none yet, but the review may still be running. I'll check again if you'd like.

can update the description of the corresponding PR that was created

3d ago·15.4s

I updated the description of PR #2645 (https://github.com/entireio/cli/pull/2645). It now matches the trail: the problem, the fix, the two tests, and the mise run check note.

I only replaced Entire's placeholder paragraph. These parts are unchanged:

  • the trail link block, which keeps the PR and trail #1470 connected
  • the entire-shadow-pr marker
  • Cursor Bugbot's summary, which already describes the current commit

The PR is still a draft, and I left its title as it was.

can you look at the pr comment

3d ago·36.6s

I looked at the comment, and it's valid. I fixed it locally but haven't committed or pushed yet.

The comment: Copilot left one inline comment, on manual_commit_condensation.go:298 (it says the same applies at :315). docs/development/checkpoint-implementation.md:52 still says task transcripts are resolved "declared path first, agent-layout fallback". That's out of date now that there's a step in between.

Fix: I updated that sentence to give the new order:

  1. the declared path
  2. the inventory lookup, used when the declared path is empty or the file is gone (it accepts a Codex rollout only if its session_meta.id matches the agent ID)
  3. the agent-layout fallback

It also names the two cases this covers: a parent committing mid-turn, and a rollout archived after its path was recorded. The docs-budget test (go test ./docs/development) passes.

Gap Copilot's overview repeats: a declared file that exists but can't be read still skips the inventory lookup. That's the gap I left open on purpose, but it can be closed cheaply. Instead of checking only that the file exists, the code could try to open the file and close it again. That catches permission errors without reading the file twice. It's a few lines plus a test.

Should I close that gap too, then commit and push both changes to PR #2645? Or push just the doc fix?

yes close it, then commit and push

3d ago·8m

I closed the gap and fixed the doc Copilot flagged, then committed and pushed both to PR #2645 (commit 52e51a8ea2).

Closing the gap: a declared transcript that exists but can't be opened now gets the inventory lookup too.

  • I added agent.CheckTranscriptReadable in cmd/entire/cli/agent/transcript_file.go. For an outside path, it opens and closes the file without reading it. A missing file, a permission error or a directory all count as unusable. For files under .entire it only checks the file's metadata, the same way StatTranscriptFile does.
  • resolveInventoryTaskTranscripts now uses that check instead of only checking the file exists.
  • New test: TestMaterializeTaskRecords_CodexUnreadableDeclaredPathResolvesThroughInventory gives a declared file with no read permission (mode 000), placed outside the rollout folders. With the old existence check it fails with transcript unreadable.
  • TestCheckTranscriptReadable_ExternalPaths covers readable, missing, directory and permission-denied paths.

Limit: an unreadable file inside Codex's rollout folders still makes Codex's own search fail. That was already true before, and it's deliberate: if the search can't read every file, it can't prove the match is the only one. So this fix helps only when the unreadable declared file is somewhere else.

Copilot's comment: checkpoint-implementation.md now lists the order as declared path, then the inventory lookup, then the Claude-layout fallback.

Checks:

  • Lint reports 0 issues after the latest formatting, and go vet passes when built for Windows.
  • mise run check has two failures:
    • TestOpenCodeSeedRepoPlantsDeps, the known one that needs the opencode bootstrap.
    • Two claudecode config-probe tests. In one full run they timed out after 5s. They passed three times in a row run alone, and this change doesn't touch that package.

Descriptions: I updated both the PR and the trail. They no longer say the unreadable-file gap is open, and they list the new tests and the timeout note.

entire trail finding list shows no findings so far, but the review of this push may still be running. I haven't replied to Copilot's comment on GitHub or resolved it. Should I?

can you check the trail finding

3d ago·57.9s

I don't think the finding is right. The danger it describes doesn't come from batching, and nothing gets worse than it was before this PR. I haven't dismissed it yet.

The throwaway test (deleted afterwards; the working tree is clean) used three subagents. A's rollout file inside the Codex sessions folder is unreadable. B's rollout path was already known from an earlier refresh. C's rollout is readable but has to be found by scanning the folder.

RunABC
All three in one callunreadablestoredunresolvable
C alone––unresolvable
B + C, A left out–storedunresolvable

Why each claim fails:

  • Batching: C fails the same way whether it's batched or alone, and whether or not A is in the call at all. The scan walks the whole folder to prove a match is unique, so it trips over A's file regardless. Splitting the call per agent would change nothing.
  • "Every other subagent": B survives. A failed scan only clears agents that needed the scan (codex/transcript.go:726-733); agents loaded from a known path keep their result.
  • "Lost for good" due to this PR: on main, condensation runs no inventory lookup at all, so C was already stored as transcript path unresolvable. This PR adds recoveries and doesn't remove any.
  • The test placement: the test puts its unreadable file outside the folders on purpose, because inside them the scan fails closed. I already describe that limit in the PR description.

What is real: Codex's scan stops entirely on any unreadable file under ~/.codex/sessions (or archived_sessions), and only logs that at Debug level. That isn't new. The turn-end refresh in refreshCodexInventory has the same behaviour, so this PR didn't introduce it. Making the scan skip unreadable files would weaken its guarantee that a match is unique, which deserves its own decision.

Suggested dismissal text:

Batching does not change outcomes: Codex's fallback scan walks the whole rollout tree for uniqueness, so an unreadable rollout fails it identically for a single-ID call, and regardless of whether that record is routed here. Children loaded directly from a known path keep their result (codex/transcript.go:726-733). On main, records needing the scan were already stored as "transcript path unresolvable", so this adds recoveries and removes none. The fail-closed scan is pre-existing and shared with refreshCodexInventory; tracking separately.

Should I dismiss it with that text? I can also shelve the scan behaviour as a follow-up.

yes

3d ago·17.3s

Dismissed. Now shelving the fail-closed scan as a follow-up.

Base directory for this skill: /Users/soph/.claude/skills/shelve

Shelve a finding

The store is dumb on purpose: it finds candidates, you decide whether one is the same finding. Never bump or add without the user's answer to the questions below.

0. Preflight

Before anything else, confirm the plugin is reachable:

entire agent-help knows only the Entire CLI's built-in commands, so entire agent-help priority fails with unknown command "priority" even when the plugin is installed. Never use it as the check and never read it as proof the plugin is missing. Usage for this plugin comes from entire priority agent-help [command] or entire priority <command> --help.

If entire priority version itself fails with "unknown command" or "command not found", do not give up silently and do not fall back to another tool. Run these three and report all three outputs to the user:

Then tell them the fix: run mise run dev:publish inside the entire-priority checkout, which relinks the priority plugin into the Entire CLI, or, if entire plugin list itself is not a command, the entire first on PATH is a build without plugin support and a newer Entire CLI must come first on PATH. Stop until the user has fixed it.

1. Gather the finding

From the conversation, collect:

  • name: a short, specific title (under about 70 characters), phrased as the problem, not the fix. "Session cache never expires entries", not "Fix cache".
  • description: two to four sentences with what is wrong, why it matters, and what was observed. Include enough that a future session with no context can act on it.
  • file and lines: the most relevant file and line range, if there is one. Pass the path as it appears in the conversation; the tool stores it repo-relative.

Run from inside the repository the finding belongs to, so repo, commit, and the checkout path are recorded automatically. repo is recorded as gh/<owner>/<repo> for GitHub and et/<project>/<repo> for an Entire-native repo. If the finding belongs to another repository, pass --repo with its key in that form, or with any of its remote URLs.

2. Look for duplicates

Item ids are ULIDs. The JSON carries the full item.id; the short id you show the user is its last 7 characters. Pass the full id to bump and set-priority.

The result has candidates, each with item, tier (exact or similar), score, and reason. exact means an open or in-progress item already has a sighting overlapping this file and line range, or has the same name after normalization. similar comes from full-text search and is often noise.

Read each candidate's item.name and item.description and judge whether it describes the same underlying problem. Same file is not enough; same root cause is.

3a. If one looks like the same item

Ask the user, in one message:

  • "This looks like #<short id> <name> (seen <count> times, priority <priority>). Same item?"
  • "Change its priority? It is currently <priority> (1 highest, 5 lowest)."

If the user confirms it is the same:

If the user says it is different, continue with 3b.

3b. If nothing matches

Ask for a priority (1 highest, 5 lowest; suggest one with a one-line reason, default 3), then:

You do not need to pass --checkpoint: the confirmed add or bump defaults to --checkpoint auto, which asks the Entire CLI which session is running the command (entire session current --json) and, only when that session is identified rather than guessed, records it and runs entire checkpoint create --json in the checkout, which snapshots the current session's transcript into an Entire checkpoint, and links the new sighting to it, so the user can reopen this session's log later with entire priority explain <id>. Never run add or bump before the user has confirmed. If the Entire CLI cannot identify the session, auto silently links nothing; pass --checkpoint create to force the attempt. --checkpoint none opts out. If the checkpoint cannot be created (no entire on PATH, no identifiable agent session, or an older Entire CLI), the sighting is still recorded, the command still exits 0, and one line starting with warning: on stderr says why; sighting.checkpoint_id in the JSON is then null.

Shelving several findings at once

When the user confirms more than one finding in one go, create a single checkpoint for the batch, since a checkpoint snapshots the whole session transcript so far:

Take checkpoint_id from its output and pass --checkpoint <that id> to every add or bump of the batch. Report one checkpoint line for the whole batch instead of one per item, for example "Shelved #7QX2M4B and #K3E95GF, both linked to checkpoint 01M3D8JVVFRC2K3E95GF2WEX97." If entire checkpoint create --json fails, run the adds and bumps without --checkpoint and report the failure once.

Forgetting this costs nothing: plain add and bump calls in the same identified session within two minutes of the first one share a checkpoint automatically, because --checkpoint auto reuses the one linked moments ago instead of creating another.

4. Report

Tell the user the item's short id (the last 7 characters of item.id), its priority, its count, and the linked checkpoint from sighting.checkpoint_id, in the form "Shelved as #<short id> (priority N, seen M times)", for example "Shelved as #7QX2M4B (priority 2, seen 3 times), checkpoint 01M3D8JVVFRC2K3E95GF2WEX97." If the warning: line fired, say instead that no checkpoint could be created and quote the reason from the warning, for example "Shelved as #7QX2M4B (priority 2, seen 3 times); no checkpoint could be created: entire checkpoint create --json: exit status 1: no agent session found." If sighting.checkpoint_id is null and no warning was printed, say that no checkpoint was linked because the Entire CLI could not identify the agent session running the command. Then return to the task at hand.

ARGUMENTS: Codex fallback rollout scan fails closed on any unreadable file under the rollout roots (~/.codex/sessions, archived_sessions): scanFallbackRollouts aborts on the first inspectFallbackCandidate/openScopedRollout error and discards all results, logged only at Debug. Every child that needs the scan (no readable declared/resolved path) then stays unresolved — at turn-end refreshCodexInventory and now at condensation (resolveInventoryTaskTranscripts, PR #2645), where completed task records are dropped after write with "transcript path unresolvable". Pre-existing; skipping unreadable candidates would weaken the uniqueness guarantee (an unread file could be a duplicate session_meta.id), so the fix needs a decision. Measured in a probe: unreadable A in roots → C (needs scan) unresolvable whether batched or alone; B (direct path) unaffected. Code: cmd/entire/cli/agent/codex/codex.go:480-520, transcript.go:726-733.

3d ago·10.0s