Parse Tool Usage from Transcript Files

Claude Code·Fable 5·computermode·3mo ago·1hr 48min·1 Checkpoint·6 file changes·+396/-29·35.8K tokens

in compact transcript.jsonl files, we need to parse and get information on tool usage such as entire search. this will be highlighted in the UI. explain the state of the jsonl files now and metadata.json so we can easily parse this on the UI side.

3mo ago·3m

Here's the current state of the compact transcript format and how the UI should parse it. Note that part of this is uncommitted work on the push-compact-for-v1 branch (the committed.go/paths.go/checkpoint.go diff), so the metadata pointer behavior described below ships with that change.

Where the files live

On the entire/checkpoints/v1 branch, each checkpoint is a directory <2-char-shard>/<10-char-rest>/ containing:

How the UI finds the right file: root metadata.json

The root metadata.json (CheckpointSummary) has a sessions array of SessionFilePaths:

With the pending change, transcript points at the compact transcript.jsonl when one was generated, otherwise at full.jsonl. Checkpoints written by older CLI versions always point at full.jsonl. So the UI should follow the pointer and sniff the format rather than hardcode the filename: parse the first line and check for "v": 1 — that field only exists in the compact format. Compact generation is best-effort: it's skipped (with a fallback pointer to full.jsonl) if generation fails, output is empty, or the result exceeds 50MB.

Two distinct metadata.json files exist — don't confuse them. The per-session one (your metadata-example.json) carries agent, token_usage, transcript_lines_at_start, initial_attribution. Importantly, the UI does not need transcript_lines_at_start to slice anything: transcript.jsonl is already truncated to the checkpoint's own slice of the session (the compactor applies checkpoint_transcript_start at write time, per agent format).

Compact line format (Entire Transcript Format, v1)

Every line is self-describing JSON, normalized across all agents (Claude Code, Codex, Gemini, OpenCode, Cursor, Droid, Copilot, Pi):

  • type: "user" — content is an array: a text block {"id": "<promptId>", "text": "..."} plus pass-through image blocks. IDE context tags are stripped.
  • type: "assistant" — content is an array of blocks. Streaming fragments with the same message id are already merged into one line; thinking blocks are dropped. input_tokens/output_tokens are present when the agent reports usage.

Tool usage — what the UI parses

Tool calls appear only inside assistant content as tool_use blocks, with the tool result already inlined (no separate result lines to correlate):

  • result.status is "success" or "error". result is absent if no tool result followed (e.g. interrupted).
  • Rich optional metadata: result.file ({"filePath": "...", "numLines": 42} for Read/Edit) and result.matchCount (Grep).
  • input is the verbatim tool input — only type/id/name/input are kept from the raw block.

So to highlight entire search usage, scan assistant lines for tool_use blocks and match:

  1. name == "Bash" and input.command matches entire search (also covers entire checkpoint search)
  2. name == "Task" / "Agent" with input.subagent_type == "entire-search" (the search subagent)
  3. name == "Skill" with input.skill == "search"

Generic tool-usage stats are just (block.name, block.result?.status) over all tool_use blocks — no agent-specific handling needed, that's the point of the format.

Caveats

  • The transcript-output.jsonl in the repo root is stale. It's from an older iteration of the format (separate user_tool_result lines, user content as a plain string). The canonical current format is what cmd/entire/cli/transcript/compact/testdata/*_expected.jsonl shows — use those as UI fixtures, not the repo-root file.
  • Legacy checkpoints (and the >50MB / generation-failure fallback) give you raw full.jsonl, which is per-agent format — the UI should treat anything without "v":1 on line one as "no structured tool parsing available" rather than trying to parse it.
  • ts is passed through raw from the agent, so treat it as an opaque RFC3339-ish string; it can be absent.
  • Field order on each line is stable (struct order: v, agent, cli_version, type, ts, id, input_tokens, output_tokens, content), but parse it as JSON, not positionally.

And about the commit trailer, if we import checkpoints when people enable entire for the first time, we don't want to update all their old commits.. but we want features to still work for entire. what at least in the CLI relies on the trailer message itself to work properly?

3mo ago·2m

Short answer: the Entire-Checkpoint trailer is the only commit→checkpoint link the CLI has. Checkpoint metadata on entire/checkpoints/v1 is keyed by checkpoint ID and does not store the user commit hash, so there is no reverse index — every feature that starts from a commit and needs its checkpoint parses the trailer out of the commit message. If imported checkpoints exist on the metadata branch but old commits carry no trailer, everything keyed by checkpoint ID still works, and everything keyed by commit silently won't see them.

What reads the trailer (and breaks for trailer-less commits)

entire explain / entire checkpoint explain

  • Commit → checkpoint: explain.go:529 and explain.go:2339 resolve a commit argument to its checkpoint via ParseCheckpoint. Without a trailer, read-only mode prints a "no Entire-Checkpoint trailer" message; --generate and raw-transcript modes hard-error (explain.go:533).
  • Checkpoint → commit (the reverse) also goes through trailers: explain.go:1365/1383 scan the commit log and match each commit's trailer against the target checkpoint ID to show which commit(s) contain a checkpoint. So even explain <checkpoint-id> on an imported checkpoint would work for metadata/transcript but show no associated commit.
  • explain_export.go:157 — same dependency for export.

entire status

  • status.go:501 — HEAD linkage (which checkpoints HEAD carries) feeds the status display and session-divergence warnings.
  • head_checkpoint_flags.go:44 — reads HEAD's trailer to surface the has_review / has_investigation flags. No trailer → no flags, silently.

entire session resume

  • resume.go:475 checks HEAD for trailers; resume.go:544 walks back through branch commits looking for the most recent trailer to find the checkpoint/session to resume. Imported history is invisible here — resume from a committed checkpoint only works through the trailer chain.

entire checkpoint rewind

  • manual_commit_rewind.go:197 and the collectCheckpoint closure at explain-side manual_commit_logs path (trailers.ParseCheckpoint at the rewind-points collector around manual_commit_logs.go:2044 region) build the committed rewind points by scanning reachable commits for trailers and intersecting with the checkpoints branch. A checkpoint with no trailer-bearing commit never appears as a rewind point. (The Entire-Session trailer reads in manual_commit_rewind.go:315/469/532 are on shadow-branch commits Entire writes itself, so imports don't affect those.)

entire review

  • review_context.go:107 — collects checkpoint summaries for recent commits via ParseAllCheckpoints to hand context to the reviewing agent. Imported checkpoints contribute nothing.

entire dispatch (local mode)

  • mode_local.go:256 literally runs git log --grep "Entire-Checkpoint:" to compute which checkpoints are reachable from HEAD, and mode_local.go:398 maps checkpoint ID → commit subject. Both are used for candidate selection/dedup.

entire attach

  • attach.go:486 reuses HEAD's existing checkpoint ID if present; attach.go:691 skips the amend prompt when the trailer already exists. This is the writing direction — attach works fine on trailer-less commits (it offers to amend the trailer in), which is actually the one escape hatch that exists today.

What reads it but is unaffected by import

These only touch commits created while Entire is enabled, so trailer-less historical commits don't matter:

  • prepare-commit-msg / post-commit hooks (manual_commit_hooks.go:168, 452, 569, 905, 2104) — write/preserve the trailer and use HEAD's trailer to trigger condensation and link sessions. A commit without a trailer just gets BaseCommit updated and no condensation.
  • Session reconcile after rebase/amend (manual_commit_migration.go:72) — matches HEAD's trailer against the active session's LastCheckpointID.

What doesn't depend on the trailer at all

entire checkpoint list, entire search, reading checkpoint metadata/transcripts/prompts/attribution by checkpoint ID, and pushing/fetching the entire/checkpoints/v1 branch — all operate directly on the metadata branch.

Implication for import

To make the commit-keyed features work for imported checkpoints without rewriting history, you'd need a secondary commit→checkpoint mapping that the ~10 ParseCheckpoint/ParseAllCheckpoints call sites consult as a fallback. Two natural shapes: store the user commit hash in the imported checkpoint's metadata (plus a small index on the metadata branch so commit-keyed lookups don't have to scan every checkpoint), or attach the mapping via git notes (refs/notes/entire), which annotates existing commits without changing their SHAs. Either way, the lookup helper would become "trailer first, mapping fallback" in one shared place rather than touching each call site's logic.