Parse Tool Usage from Transcript Files

in compact transcript.jsonl files, we need to parse and get information on tool usage such as entire search. this will be highlighted in the UI. explain the state of the jsonl files now and metadata.json so we can easily parse this on the UI side.
Here's the current state of the compact transcript format and how the UI should parse it. Note that part of this is uncommitted work on the push-compact-for-v1 branch (the committed.go/paths.go/checkpoint.go diff), so the metadata pointer behavior described below ships with that change.
Where the files live
On the entire/checkpoints/v1 branch, each checkpoint is a directory <2-char-shard>/<10-char-rest>/ containing:
How the UI finds the right file: root metadata.json
The root metadata.json (CheckpointSummary) has a sessions array of SessionFilePaths:
With the pending change, transcript points at the compact transcript.jsonl when one was generated, otherwise at full.jsonl. Checkpoints written by older CLI versions always point at full.jsonl. So the UI should follow the pointer and sniff the format rather than hardcode the filename: parse the first line and check for "v": 1 — that field only exists in the compact format. Compact generation is best-effort: it's skipped (with a fallback pointer to full.jsonl) if generation fails, output is empty, or the result exceeds 50MB.
Two distinct metadata.json files exist — don't confuse them. The per-session one (your metadata-example.json) carries agent, token_usage, transcript_lines_at_start, initial_attribution. Importantly, the UI does not need transcript_lines_at_start to slice anything: transcript.jsonl is already truncated to the checkpoint's own slice of the session (the compactor applies checkpoint_transcript_start at write time, per agent format).
Compact line format (Entire Transcript Format, v1)
Every line is self-describing JSON, normalized across all agents (Claude Code, Codex, Gemini, OpenCode, Cursor, Droid, Copilot, Pi):
type: "user"—contentis an array: a text block{"id": "<promptId>", "text": "..."}plus pass-through image blocks. IDE context tags are stripped.type: "assistant"—contentis an array of blocks. Streaming fragments with the same messageidare already merged into one line; thinking blocks are dropped.input_tokens/output_tokensare present when the agent reports usage.
Tool usage — what the UI parses
Tool calls appear only inside assistant content as tool_use blocks, with the tool result already inlined (no separate result lines to correlate):
result.statusis"success"or"error".resultis absent if no tool result followed (e.g. interrupted).- Rich optional metadata:
result.file({"filePath": "...", "numLines": 42}for Read/Edit) andresult.matchCount(Grep). inputis the verbatim tool input — onlytype/id/name/inputare kept from the raw block.
So to highlight entire search usage, scan assistant lines for tool_use blocks and match:
name == "Bash"andinput.commandmatchesentire search(also coversentire checkpoint search)name == "Task"/"Agent"withinput.subagent_type == "entire-search"(the search subagent)name == "Skill"withinput.skill == "search"
Generic tool-usage stats are just (block.name, block.result?.status) over all tool_use blocks — no agent-specific handling needed, that's the point of the format.
Caveats
- The
transcript-output.jsonlin the repo root is stale. It's from an older iteration of the format (separateuser_tool_resultlines, user content as a plain string). The canonical current format is whatcmd/entire/cli/transcript/compact/testdata/*_expected.jsonlshows — use those as UI fixtures, not the repo-root file. - Legacy checkpoints (and the >50MB / generation-failure fallback) give you raw
full.jsonl, which is per-agent format — the UI should treat anything without"v":1on line one as "no structured tool parsing available" rather than trying to parse it. tsis passed through raw from the agent, so treat it as an opaque RFC3339-ish string; it can be absent.- Field order on each line is stable (struct order:
v, agent, cli_version, type, ts, id, input_tokens, output_tokens, content), but parse it as JSON, not positionally.
And about the commit trailer, if we import checkpoints when people enable entire for the first time, we don't want to update all their old commits.. but we want features to still work for entire. what at least in the CLI relies on the trailer message itself to work properly?
Short answer: the Entire-Checkpoint trailer is the only commit→checkpoint link the CLI has. Checkpoint metadata on entire/checkpoints/v1 is keyed by checkpoint ID and does not store the user commit hash, so there is no reverse index — every feature that starts from a commit and needs its checkpoint parses the trailer out of the commit message. If imported checkpoints exist on the metadata branch but old commits carry no trailer, everything keyed by checkpoint ID still works, and everything keyed by commit silently won't see them.
What reads the trailer (and breaks for trailer-less commits)
entire explain / entire checkpoint explain
- Commit → checkpoint:
explain.go:529andexplain.go:2339resolve a commit argument to its checkpoint viaParseCheckpoint. Without a trailer, read-only mode prints a "no Entire-Checkpoint trailer" message;--generateand raw-transcript modes hard-error (explain.go:533). - Checkpoint → commit (the reverse) also goes through trailers:
explain.go:1365/1383scan the commit log and match each commit's trailer against the target checkpoint ID to show which commit(s) contain a checkpoint. So evenexplain <checkpoint-id>on an imported checkpoint would work for metadata/transcript but show no associated commit. explain_export.go:157— same dependency for export.
entire status
status.go:501— HEAD linkage (which checkpoints HEAD carries) feeds the status display and session-divergence warnings.head_checkpoint_flags.go:44— reads HEAD's trailer to surface thehas_review/has_investigationflags. No trailer → no flags, silently.
entire session resume
resume.go:475checks HEAD for trailers;resume.go:544walks back through branch commits looking for the most recent trailer to find the checkpoint/session to resume. Imported history is invisible here — resume from a committed checkpoint only works through the trailer chain.
entire checkpoint rewind
manual_commit_rewind.go:197and thecollectCheckpointclosure atexplain-sidemanual_commit_logspath (trailers.ParseCheckpointat the rewind-points collector aroundmanual_commit_logs.go:2044region) build the committed rewind points by scanning reachable commits for trailers and intersecting with the checkpoints branch. A checkpoint with no trailer-bearing commit never appears as a rewind point. (TheEntire-Sessiontrailer reads inmanual_commit_rewind.go:315/469/532are on shadow-branch commits Entire writes itself, so imports don't affect those.)
entire review
review_context.go:107— collects checkpoint summaries for recent commits viaParseAllCheckpointsto hand context to the reviewing agent. Imported checkpoints contribute nothing.
entire dispatch (local mode)
mode_local.go:256literally runsgit log --grep "Entire-Checkpoint:"to compute which checkpoints are reachable from HEAD, andmode_local.go:398maps checkpoint ID → commit subject. Both are used for candidate selection/dedup.
entire attach
attach.go:486reuses HEAD's existing checkpoint ID if present;attach.go:691skips the amend prompt when the trailer already exists. This is the writing direction — attach works fine on trailer-less commits (it offers to amend the trailer in), which is actually the one escape hatch that exists today.
What reads it but is unaffected by import
These only touch commits created while Entire is enabled, so trailer-less historical commits don't matter:
prepare-commit-msg/post-commithooks (manual_commit_hooks.go:168, 452, 569, 905, 2104) — write/preserve the trailer and use HEAD's trailer to trigger condensation and link sessions. A commit without a trailer just getsBaseCommitupdated and no condensation.- Session reconcile after rebase/amend (
manual_commit_migration.go:72) — matches HEAD's trailer against the active session'sLastCheckpointID.
What doesn't depend on the trailer at all
entire checkpoint list, entire search, reading checkpoint metadata/transcripts/prompts/attribution by checkpoint ID, and pushing/fetching the entire/checkpoints/v1 branch — all operate directly on the metadata branch.
Implication for import
To make the commit-keyed features work for imported checkpoints without rewriting history, you'd need a secondary commit→checkpoint mapping that the ~10 ParseCheckpoint/ParseAllCheckpoints call sites consult as a fallback. Two natural shapes: store the user commit hash in the imported checkpoint's metadata (plus a small index on the metadata branch so commit-keyed lookups don't have to scan every checkpoint), or attach the mapping via git notes (refs/notes/entire), which annotates existing commits without changing their SHAs. Either way, the lookup helper would become "trailer first, mapping fallback" in one shared place rather than touching each call site's logic.