Multi-Root Config Path Agent Support

It appears entire session adopt does not work when you have a setup like mine, where e.g. you have two instances of Claude with each a non-default config directory (see wclaude and pclaude). Brainstorm and plan on how we can support this.
Base directory for this skill: /home/h/.claude-work/plugins/cache/claude-plugins-official/superpowers/b36e0829c6d0/skills/brainstorming
Brainstorming Ideas Into Designs
Help turn ideas into fully formed designs and specs through natural collaborative dialogue.
Start by classifying how much process the request needs, then work through your path: understand the context, refine the idea, present a design, and get your human partner's approval.
<HARD-GATE> Do NOT invoke any implementation skill, write any code, scaffold any project, or take any implementation action until you have told your human partner what you intend and they have approved it. This applies to EVERY task on EVERY path below — the ceremony scales with the task; the approval gate never does. </HARD-GATE>Three Paths
Before your first question, classify the request and say the classification out loud — "this looks bounded, so I'll present a short design here rather than write a spec" — so your human partner can override it:
- Spike — a feasibility question ("can we...", "is it possible...", "quick and dirty is fine") whose output is an answer, not code you keep. Present the question and what you'll try in 2-3 sentences, get a nod, then find out as cheaply as correctness allows. No design doc, no spec file. Report findings as a recommendation; anything you built stays labeled throwaway.
- Bounded — a well-scoped change to code that already exists in this repo: a new flag, a small endpoint, a one-file fix. Understanding the kind of app is not enough — bounded means the flow you are changing is already here to read. If there is no existing flow to change, the task is not bounded. Ask the clarifying questions that matter, present a short design IN CHAT (a few sentences to a few short paragraphs), and STOP. Implementation starts only after your human partner says yes to that design — a bounded task's approval is as hard a gate as an architectural one. No spec file, no implementation plan document.
- Architectural — new projects, new subsystems, changes that restructure how components fit together or alter interfaces others depend on. Follow the full process: questions, approaches, sectioned design, written spec, then the writing-plans skill.
When in doubt between two paths, take the heavier one. The ratchet is one-way: hidden complexity discovered mid-task upgrades the path — stop, say so, and step up. Nothing downgrades mid-task.
Anti-Pattern: "Too Simple To Need Approval"
Every path ends with your human partner approving your intent before implementation. A todo list, a single-function utility, a config change — the design may be two sentences in chat, but you MUST present it and get approval. "Simple" tasks are where unexamined assumptions cause the most wasted work. What scales with simplicity is the artifact, never the approval.
Red Flags
| Thought | Reality |
|---|---|
| "This is too simple to need a design" | Simple means a short design, not no design. Two sentences in chat, then approval. |
| "I'll call it bounded and skip the spec" | Reaching for a label to skip work IS the doubt — take the heavier path. |
| "It's bounded and the design is obvious — I'll start while they read it" | The gate is the approval, not the design's length. Present, then stop until you hear yes. |
| "I understand this kind of app, so it's bounded" | Bounded measures the repo, not your familiarity. A new project has no existing flow — it is architectural. |
| "The spike works, so I'll keep the code" | A spike's output is an answer. Keeping the code is a new request — classify it. |
| "It grew, but I'm almost done — no need to re-classify" | Hidden complexity upgrades the path mid-task. Stop and say so. |
| "They approved the spike, so the follow-up change is approved too" | Each task gets its own classification and its own approval. |
Checklist
Classify first, announce the path, then create a task for each item on your path and complete them in order.
Spike:
- Explore project context — enough to frame the probe
- Present question + probe plan — 2-3 sentences
- Get approval — a nod is enough
- Investigate — as cheaply as correctness allows
- Report findings — a recommendation; label anything built as throwaway
Bounded:
- Explore project context — check files, docs, recent commits
- Ask clarifying questions — one at a time, the ones that matter
- Present short design in chat — approach, files touched, testing
- Get approval — STOP and wait for an explicit yes; presenting the design and starting in the same breath is skipping the gate
- Implement — proceed with the normal development workflow (TDD applies); no plan document
Architectural:
- Explore project context — check files, docs, recent commits
- Offer the visual companion just-in-time — NOT upfront. The first time a question would genuinely be clearer shown than described, offer it then (its own message); on approval its browser tab opens for you. If no visual question ever arises, never offer it. See the Visual Companion section below.
- Ask clarifying questions — one at a time, understand purpose/constraints/success criteria
- Propose 2-3 approaches — with trade-offs and your recommendation
- Present design — in sections scaled to their complexity, get user approval after each section
- Write design doc — save to
docs/superpowers/specs/YYYY-MM-DD-<topic>-design.mdand commit - Spec self-review — quick inline check for placeholders, contradictions, ambiguity, scope (see below)
- User reviews written spec — ask user to review the spec file before proceeding
- Transition to implementation — invoke writing-plans skill to create implementation plan
Process Flow
Terminal states are path-bound. Architectural: the ONLY skill you invoke after brainstorming is writing-plans — never frontend-design, mcp-builder, or any other implementation skill. Bounded: after approval, implementation proceeds directly through the normal development workflow; no plan document. Spike: the terminal state is a reported recommendation.
The Process
The subsections below serve the bounded and architectural paths (a spike stops at "present the probe, get a nod"). Sections from Exploring approaches onward are architectural-path depth — for bounded work, context plus a few questions plus a short in-chat design is the whole process.
Understanding the idea:
- Check out the current project state first (files, docs, recent commits)
- Before asking detailed questions, assess scope: if the request describes multiple independent subsystems (e.g., "build a platform with chat, file storage, billing, and analytics"), flag this immediately. Don't spend questions refining details of a project that needs to be decomposed first.
- If the project is too large for a single spec, help the user decompose into sub-projects: what are the independent pieces, how do they relate, what order should they be built? Then brainstorm the first sub-project through the normal design flow. Each sub-project gets its own spec → plan → implementation cycle.
- For appropriately-scoped projects, ask questions one at a time to refine the idea
- Prefer multiple choice questions when possible, but open-ended is fine too
- Only one question per message - if a topic needs more exploration, break it into multiple questions
- Focus on understanding: purpose, constraints, success criteria
Exploring approaches:
- Propose 2-3 different approaches with trade-offs
- Present options conversationally with your recommendation and reasoning
- Lead with your recommended option and explain why
- YAGNI ruthlessly - remove unnecessary features from every approach and design
Presenting the design:
- Once you believe you understand what you're building, present the design
- Scale each section to its complexity: a few sentences if straightforward, up to 200-300 words if nuanced
- Ask after each section whether it looks right so far
- Cover: architecture, components, data flow, error handling, testing
- Be ready to go back and clarify if something doesn't make sense
Design for isolation and clarity:
- Break the system into smaller units that each have one clear purpose, communicate through well-defined interfaces, and can be understood and tested independently
- For each unit, you should be able to answer: what does it do, how do you use it, and what does it depend on?
- Can someone understand what a unit does without reading its internals? Can you change the internals without breaking consumers? If not, the boundaries need work.
- Smaller, well-bounded units are also easier for you to work with - you reason better about code you can hold in context at once, and your edits are more reliable when files are focused. When a file grows large, that's often a signal that it's doing too much.
Working in existing codebases:
- Explore the current structure before proposing changes. Follow existing patterns.
- Where existing code has problems that affect the work (e.g., a file that's grown too large, unclear boundaries, tangled responsibilities), include targeted improvements as part of the design - the way a good developer improves code they're working in.
- Don't propose unrelated refactoring. Stay focused on what serves the current goal.
After the Design (architectural path)
Documentation:
- Write the validated design (spec) to
docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md- (User preferences for spec location override this default)
- Use elements-of-style:writing-clearly-and-concisely skill if available
- Commit the design document to git
Spec Self-Review: After writing the spec document, look at it with fresh eyes:
- Placeholder scan: Any "TBD", "TODO", incomplete sections, or vague requirements? Fix them.
- Internal consistency: Do any sections contradict each other? Does the architecture match the feature descriptions?
- Scope check: Is this focused enough for a single implementation plan, or does it need decomposition?
- Ambiguity check: Could any requirement be interpreted two different ways? If so, pick one and make it explicit.
Fix any issues inline. No need to re-review — just fix and move on.
User Review Gate: After the spec review loop passes, ask the user to review the written spec before proceeding:
"Spec written and committed to
<path>. Please review it and let me know if you want to make any changes before we start writing out the implementation plan."
Wait for the user's response. If they request changes, make them and re-run the spec review loop. Only proceed once the user approves.
Implementation:
- Invoke the writing-plans skill to create a detailed implementation plan
- Do NOT invoke any other skill. writing-plans is the next step.
Visual Companion
A browser-based companion for showing mockups, diagrams, and visual options during brainstorming. Available as a tool — not a mode. Accepting the companion means it's available for questions that benefit from visual treatment; it does NOT mean every question goes through the browser.
Offering the companion (just-in-time): Do NOT offer it upfront. Wait until a question would genuinely be clearer shown than told — a real mockup / layout / diagram question, not merely a UI topic. The first time that happens, offer it then, as its own message:
"This next part might be easier if I show you — I can put together mockups, diagrams, and comparisons in a browser tab as we go. It's still new and can be token-intensive. Want me to? I'll open it for you."
This offer MUST be its own message. Only the offer — no clarifying question, summary, or other content. Wait for the user's response. If they accept, start the server with --open so their browser opens to the first screen automatically. If they decline, continue text-only and don't offer again unless they raise it.
Per-question decision: Even after the user accepts, decide FOR EACH QUESTION whether to use the browser or the terminal. The test: would the user understand this better by seeing it than reading it?
- Use the browser for content that IS visual — mockups, wireframes, layout comparisons, architecture diagrams, side-by-side visual designs
- Use the terminal for content that is text — requirements questions, conceptual choices, tradeoff lists, A/B/C/D text options, scope decisions
A question about a UI topic is not automatically a visual question. "What does personality mean in this context?" is a conceptual question — use the terminal. "Which wizard layout works better?" is a visual question — use the browser.
If they agree to the companion, read the detailed guide before proceeding:
skills/brainstorming/visual-companion.md
ARGUMENTS: Support entire session adopt when multiple Claude instances run with non-default config directories (wclaude/pclaude)
I have what I need. Here are the shapes this could take — all three honor the two decisions you've made; they differ in how much machinery goes in.
Approach A — targeted resolution (recommended)
Three small pieces, no new agent-interface method:
claudecode.GetSessionDir/GetSessionBaseDirgain aCLAUDE_CONFIG_DIRstep, exactly mirroringcodex.resolveCodexHome(codex.go:558) and theuserClaudeSettingsPathhelper that already lives inclaudecode/generate.go:114. Precedence:ENTIRE_TEST_CLAUDE_PROJECT_DIR→CLAUDE_CONFIG_DIR/projects/…→$HOME/.claude/projects/….session.Stategainsagent_session_dir, written at session start (manual_commit_session.go:694) from entire's ownGetSessionDircall — i.e. computed while the correct env is in scope.- Adopt validates against the recorded dir when present, falling back to derived candidates for pre-existing states. Cold lookups (
searchTranscriptInProjectDirs) try the env-derived base then~/.claude.
Smallest surface, one new state field, no interface churn. Downside: the "which agent owns this path" logic grows a state-aware variant alongside the existing AgentForTranscriptPath.
Approach B — candidate-list capability on Agent
Add an optional SessionDirCandidates(repoPath) ([]string, error) capability (following the agent/capabilities.go opt-in pattern); registry ownership and fallback search iterate it; Codex folds its existing rolloutRoots in. More uniform and sets up every other agent, but it's a documented interface change: docs/architecture/agent-guide.md and the integration checklist both have to move, and every agent gets test coverage for a problem only Claude Code has today.
Approach C — trust recorded state only, drop derivation
Never recompute; always use transcript_path / the recorded dir. Kills cold lookup (entire attach <id> has no state by definition) and entire import. I'd reject this.
I recommend A. B is the right shape if and when a second agent needs multi-root, and A leaves the door open — SessionDirCandidates can be extracted later from the two call sites without redoing the state field or the env handling.
One deviation from your "explicit flag" answer, which I think you'll agree with: no new flag is needed. CLAUDE_CONFIG_DIR=/home/h/.claude-personal entire attach <id> is the explicit override once step 1 lands, and entire import claude-code --path <dir> already exists (import_cmd.go:113). So the escape hatch becomes an error message that names the roots searched and points at the env var, rather than a new --claude-config-dir surface. Say the word if you'd still rather have a real flag.
Does Approach A look right, and is dropping the new flag acceptable?
I think a second agent needs multi-root, as I also have a wcodex and pcodex. I think the primarily issue to solve is: agents can be tied to a custom config path, and we have no proper mechanism to work with this when e.g. adopting.
Two corrections to my table before the design — I checked the constants and got two wrong:
- Cursor:
cursorChatsDirEnvisENTIRE_TEST_CURSOR_CHATS_DIR— entire's own test override, not a user relocation env. - Antigravity:
ENTIRE_ANTIGRAVITY_CONFIG_DIR/_STATUS_DIRare likewise entire's, not the agent's.
So the agents with a real user-facing relocation env are Claude Code (CLAUDE_CONFIG_DIR, unhonored), Codex (CODEX_HOME), Copilot CLI (COPILOT_HOME), Pi (PI_CODING_AGENT_DIR). Cursor, Antigravity, Factory and external agents have no such env and stay single-root.
Design
Architecture. Introduce storage root — the relocatable directory an agent keeps state under — as an optional capability alongside the ~26 existing ones in agent/capabilities.go:
Binding invariant: GetSessionDir(repoPath) must equal SessionDirUnder(roots[0], repoPath). Writes and hook behavior are unchanged by construction, and it's a testable contract.
Components.
agent/capabilities.go— capability + accessor.claudecode— roots$CLAUDE_CONFIG_DIR→~/.claude;SessionDirUnder = root/projects/Sanitize(repo);GetSessionDir/GetSessionBaseDirrewritten on top (test override still wins). This alone fixeswclaude/pclaudederivation.codex— roots$CODEX_HOME→~/.codex;SessionDirUnder = root/sessions; existingrolloutRoots(codex.go:267) folds in soarchived_sessionsis enumerated per root.copilotcli,pi— same conversion from their existing single-root resolvers.- Non-relocatable agents implement nothing; every consumer falls back to today's
GetSessionDir. No external-agent protocol change. session.State— newagent_session_dir, written at session start (manual_commit_session.go:694and the Codex repair branch) from entire's ownGetSessionDir.AgentForTranscriptPath— match against every root's session dir.validateAdoptSourceTranscript— prefer the recorded dir; fall back to multi-root matching.searchTranscriptInProjectDirs— walk every root's base dir.
Data flow. Session starts under wcodex → hook records agent_session_dir = /home/h/.codex-work/sessions. Later entire session adopt --from … running under pcodex loads the source state, sees the recorded dir, confirms transcript_path sits under it, and passes — independently of the adopting process's CODEX_HOME. Cold lookups with no state try each root in turn.
The honest limitation. For sessions that already exist today (no recorded field), cross-instance adopt still can't work: from wclaude the enumerable roots are .claude-work and ~/.claude, and .claude-personal is simply not discoverable. Those fail with an error naming the roots searched and suggesting CLAUDE_CONFIG_DIR=… entire session adopt …. New sessions work with no config. I'd rather ship that than add $HOME scanning.
Error handling. The cross-check stays fail-closed — a transcript is never accepted just because agent_type claims it. Failures name the roots actually searched and point at the env var.
Testing. Per-agent root-precedence tables with t.Setenv; a cross-agent table asserting the GetSessionDir == SessionDirUnder(roots[0], …) invariant; ownership matching a non-default root but rejecting a foreign agent's; adopt validation with recorded / absent / mismatched dirs; an integration test adopting across two config roots. One thing to verify rather than assume: e2e/agents/claude.go:112 already points spawned Claude at a temp CLAUDE_CONFIG_DIR, so after this change entire resolves there instead of the developer's real ~/.claude — an improvement, but the canary needs a run. Docs land in agent-guide.md, agent-integration-checklist.md, and sessions-and-checkpoints.md.
Out of scope: $HOME scanning for undeclared roots, external-agent multi-root protocol, and relocation for agents with no upstream env.
Does this look right? If so I'll write it up as a spec under docs/superpowers/specs/ and then turn it into an implementation plan.
Spec written to docs/superpowers/specs/2026-09-30-agent-storage-roots-design.md (untracked — that tree is gitignored, so there's no commit).
Two things changed during self-review, both worth knowing before you read it:
- Component 6 originally said the recorded directory was resolved "against" the transcript path, which contradicts the rationale. It's now explicit that the value comes from an independent
GetSessionDircall — that independence is the whole reason the later containment check means anything. - Added component 10.
buildAdoptedSessionStateclears a long list of target-local bookkeeping, andagent_session_dirmust not join it. Adoption deliberately keeps the source's liveTranscriptPath; the recorded directory describes that same transcript and has to travel with it. Easy to get wrong by pattern-matching the surrounding resets.
Please review it and tell me if you want changes before I turn it into an implementation plan.
Base directory for this skill: /home/h/.claude-work/plugins/cache/claude-plugins-official/superpowers/b36e0829c6d0/skills/writing-plans
Writing Plans
Overview
Write comprehensive implementation plans assuming the engineer has zero context for our codebase and questionable taste. Document everything they need to know: which files to touch for each task, code, testing, docs they might need to check, how to test it. Give them the whole plan as bite-sized tasks. DRY. YAGNI. TDD. Frequent commits.
Assume they are a skilled developer, but know almost nothing about our toolset or problem domain. Assume they don't know good test design very well.
Announce at start: "I'm using the writing-plans skill to create the implementation plan."
Context: If working in an isolated worktree, it should have been created via the superpowers:using-git-worktrees skill at execution time.
Save plans to: docs/superpowers/plans/YYYY-MM-DD-<feature-name>.md
- (User preferences for plan location override this default)
Scope Check
If the spec covers multiple independent subsystems, it should have been broken into sub-project specs during brainstorming. If it wasn't, suggest breaking this into separate plans — one per subsystem. Each plan should produce working, testable software on its own.
File Structure
Before defining tasks, map out which files will be created or modified and what each one is responsible for. This is where decomposition decisions get locked in.
- Design units with clear boundaries and well-defined interfaces. Each file should have one clear responsibility.
- You reason best about code you can hold in context at once, and your edits are more reliable when files are focused. Prefer smaller, focused files over large ones that do too much.
- Files that change together should live together. Split by responsibility, not by technical layer.
- In existing codebases, follow established patterns. If the codebase uses large files, don't unilaterally restructure - but if a file you're modifying has grown unwieldy, including a split in the plan is reasonable.
This structure informs the task decomposition. Each task should produce self-contained changes that make sense independently.
Task Right-Sizing
A task is the smallest unit that carries its own test cycle and is worth a fresh reviewer's gate. When drawing task boundaries: fold setup, configuration, scaffolding, and documentation steps into the task whose deliverable needs them; split only where a reviewer could meaningfully reject one task while approving its neighbor. Each task ends with an independently testable deliverable.
Bite-Sized Task Granularity
Each step is one action (2-5 minutes):
- "Write the failing test" - step
- "Run it to make sure it fails" - step
- "Implement the minimal code to make the test pass" - step
- "Run the tests and make sure they pass" - step
- "Commit" - step
Plan Document Header
Every plan MUST start with this header:
Task Structure
No Placeholders
Every step must contain the actual content an engineer needs. These are plan failures — never write them:
- "TBD", "TODO", "implement later", "fill in details"
- "Add appropriate error handling" / "add validation" / "handle edge cases"
- "Write tests for the above" (without actual test code)
- "Similar to Task N" (repeat the code — the engineer may be reading tasks out of order)
- Steps that describe what to do without showing how (code blocks required for code steps)
- References to types, functions, or methods not defined in any task
Self-Review
After writing the complete plan, look at the spec with fresh eyes and check the plan against it. This is a checklist you run yourself — not a subagent dispatch.
1. Spec coverage: Skim each section/requirement in the spec. Can you point to a task that implements it? List any gaps.
2. Placeholder scan: Search your plan for red flags — any of the patterns from the "No Placeholders" section above. Fix them.
3. Type consistency: Do the types, method signatures, and property names you used in later tasks match what you defined in earlier tasks? A function called clearLayers() in Task 3 but clearFullLayers() in Task 7 is a bug.
If you find issues, fix them inline. No need to re-review — just fix and move on. If you find a spec requirement with no task, add the task.
Execution Handoff
After saving the plan, offer execution choice:
"Plan complete and saved to docs/superpowers/plans/<filename>.md. Two execution options:
1. Subagent-Driven (recommended) - I dispatch a fresh subagent per task, review between tasks, fast iteration
2. Inline Execution - Execute tasks in this session using executing-plans, batch execution with checkpoints
Which approach?"
If Subagent-Driven chosen:
- REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development
- Fresh subagent per task + two-stage review
If Inline Execution chosen:
- REQUIRED SUB-SKILL: Use superpowers:executing-plans
- Batch execution with checkpoints for review
ARGUMENTS: Implementation plan for docs/superpowers/specs/2026-09-30-agent-storage-roots-design.md (agent storage roots)
Every symbol resolves consistently. Plan saved to docs/superpowers/plans/2026-09-30-agent-storage-roots.md — 10 tasks, 63 checkbox steps, no placeholders (also untracked; that tree is gitignored).
Task sequence:
| # | Task | Gate it passes |
|---|---|---|
| 1 | Declare StorageRootProvider + the three shared helpers | No behaviour change; gives 2–8 something to call |
| 2 | Claude Code declares its roots | Fixes derivation for wclaude/pclaude |
| 3 | Codex declares its roots; rolloutRoots widens per root | Fixes wcodex/pcodex derivation + archived sessions |
| 4 | Copilot CLI + Pi conversions | Mechanical; folded because neither is separately rejectable |
| 5 | Ownership matches every root | Fixes AgentForTranscriptPath and its 4 callers |
| 6 | Record agent_session_dir in session state | The half that makes cross-instance work |
| 7 | Adoption trusts the recorded dir | Kills the reported bug |
| 8 | Cold lookup walks every root | Fixes entire attach by ID |
| 9 | Source guard + cross-agent invariant table + docs | Stops the next agent repeating the mistake |
| 10 | End-to-end adopt across two config roots + full gate | Proves it through the real command |
Three things worth flagging before you read it:
Self-review changed two things. The spec asked for a table asserting GetSessionDir == SessionDirUnder(roots[0]) across all registered agents; I'd only written per-agent versions, so Task 9 gained that table — it's the guard against a future agent declaring roots whose first entry disagrees with where writes actually land. And Task 10 originally invented integration-harness helpers; it now drives the real runAdopt with the helpers that exist (setupAdoptRepo, session.NewStateStoreWithDir), and deliberately avoids claudeAdoptTranscriptPath because that helper sets ENTIRE_TEST_CLAUDE_PROJECT_DIR, which would pre-empt the resolution under test.
I refined the spec's limitation while planning, in your favour. The recorded directory is written at session start and refreshed at every turn start, where TranscriptPath is already refreshed. Turn-start hooks always run inside the owning instance, so any session still taking turns backfills itself. Only already-dormant sessions stay unadoptable.
Almost every test in this plan is non-parallel, because t.Setenv is unavoidable here. That's explicitly called out per task so nobody "fixes" it by adding t.Parallel() and gets flaky cross-test env bleed.
Two execution options:
- Subagent-driven (recommended) — fresh subagent per task, I review between tasks.
- Inline — execute here with batched checkpoints.
Which would you like? (Note the mise run check gate before every commit runs the full CI suite, so either way this is a slow-per-commit plan by the repo's own rule.)
Base directory for this skill: /home/h/.claude-work/plugins/cache/claude-plugins-official/superpowers/b36e0829c6d0/skills/subagent-driven-development
Subagent-Driven Development
Execute plan by dispatching a fresh implementer subagent per task, a task review (spec compliance + code quality) after each, and a broad whole-branch review at the end.
Why subagents: You delegate tasks to specialized agents with isolated context. By precisely crafting their instructions and context, you ensure they stay focused and succeed at their task. They should never inherit your session's context or history — you construct exactly what they need. This also preserves your own context for coordination work.
Core principle: Fresh subagent per task + task review (spec + quality) + broad final review = high quality, fast iteration
Narration: between tool calls, narrate at most one short line — the ledger and the tool results carry the record.
Continuous execution: Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
Rulings, not stalls. A running plan does not wait on a human. Conflicts,
ambiguities, plan defects, a cap you would have asked to exceed — decide
them. The spec is the binding authority, the plan is its argument, and your
judgment settles what neither answers. Record every decision in the ledger as
Ruling: <what you decided> — <why> — <what it costs if wrong>, and keep
going. A wrong ruling costs rework your human partner can see and undo; a
session parked on a question costs their whole day and buys nothing.
Four things stop you, and only these: an irreversible or destructive operation; a security-sensitive action; a side effect outside this worktree that norms say you ask about first (a merge, a push to a shared branch, a publish); and a plan so broken that every path forward is a guess. For those, stop and ask.
When to Use
vs. Executing Plans (parallel session):
- Same session (no context switch)
- Fresh subagent per task (no context pollution)
- Review after each task (spec compliance + code quality), broad review at the end
- Faster iteration (no human-in-loop between tasks)
The Process
Setup
Ensure the work happens in an isolated workspace: use superpowers:using-git-worktrees to create one or verify the existing one. Never start implementation on a main/master branch without your human partner's explicit consent.
Conversation memory does not survive compaction. In real sessions, controllers that lost their place have re-dispatched entire completed task sequences — the single most expensive failure observed. Track progress in a ledger file, not only in todos.
- Each plan owns a workspace: at skill start, run this skill's
scripts/sdd-workspace PLAN_FILE— it prints the plan's git-ignored directory (<repo-root>/.superpowers/sdd/<plan-basename>/), home to every artifact for THIS plan: ledger, briefs, reports, review packages. Another plan's directory is never yours to read or write. - Check for this plan's ledger at
<workspace>/progress.md. If its first line names your plan file, tasks with aTask <N>: completeline are DONE — do not re-dispatch them; resume at the first task without one. A task whose last line is a fix round is mid-loop: resume the loop at the next round. A ledger whose first line names a different plan file — or a stray ledger at the old flat path.superpowers/sdd/progress.md— is another plan's progress: leave it in place and start your own, fresh. - Create the ledger with its identity as the first line:
# SDD ledger — plan: <plan file path>. - The ledger is your recovery map: the commits it names exist in git even
when your context no longer remembers creating them. After compaction,
trust the ledger and
git logover your own recollection. git clean -fdxwill destroy the workspace (it's git-ignored scratch); if that happens, recover fromgit log.
Read the plan once, note its context and Global Constraints, and create a todo per task. If the plan names a Spec, read that too: the spec is the authority the plan argues from, and conflicts inside the plan resolve against it. A plan with no reachable spec gets a ledger note saying so — rulings made without one are provisional.
Before dispatching Task 1, scan the plan once for conflicts, writing down what you checked as you check it:
- tasks that contradict each other or the plan's Global Constraints
- anything the plan explicitly mandates that the review rubric treats as a defect (a test that asserts nothing, verbatim duplication of a logic block)
The scan's output is a table, not a verdict. One row for every pair of tasks that share a file or an interface: the two tasks, what one produces against what the other consumes, and what you found. One row for every task: whether its own text agrees with itself — the tests it specifies against the code it specifies, the files it creates against the files it later touches. "The scan is clean" without those rows is not a scan you ran.
Write the table to the ledger. Rule on everything you find before execution begins — each finding against the plan text that mandates it — and record each ruling in the ledger. If the scan is clean, proceed without comment. Rule on each conflict it surfaces — the spec is the binding authority, the plan is its argument — record the ruling beside its row, and dispatch Task 1. The review loop remains the net for conflicts that only emerge from implementation.
Model Selection
Use the least powerful model that can handle each role to conserve cost and increase speed.
Mechanical implementation tasks (isolated functions, clear specs, 1-2 files): use a fast, cheap model. Most implementation tasks are mechanical when the plan is well-specified.
Integration and judgment tasks (multi-file coordination, pattern matching, debugging): use a standard model.
Architecture and design tasks: use the most capable available model. The final whole-branch review is one of these — dispatch it on the most capable available model, not the session default.
Review tasks: choose the model with the same judgment, scaled to the diff's size, complexity, and risk. A small mechanical diff does not need the most capable model; a subtle concurrency change does. Scoped re-reviews of small fix diffs take a cheap-to-mid tier.
Fix-loop escalation (rounds 4-5): use a model at least one tier above the implementer that got stuck.
Always specify the model explicitly when dispatching a subagent. An omitted model inherits your session's model — often the most capable and most expensive — which silently defeats this section.
Turn count beats token price. Wall-clock and context cost scale with how many turns a subagent takes, and the cheapest models routinely take 2-3× the turns on multi-step work — costing more overall. Use a mid-tier model as the floor for reviewers and for implementers working from prose descriptions. When the task's plan text contains the complete code to write, the implementation is transcription plus testing: use the cheapest tier for that implementer. Single-file mechanical fixes also take the cheapest tier.
Task complexity signals (implementation tasks):
- Touches 1-2 files with a complete spec → cheap model
- Touches multiple files with integration concerns → standard model
- Requires design judgment or broad codebase understanding → most capable model
The Task Loop
Batch small same-shape work. When the plan lists several tasks that are each a small, independent edit of the same kind — the same one-line fix, constant change, or field addition repeated across files — do not dispatch one subagent per task. Compose ONE dispatch brief listing every file and its change, send the whole batch to a single subagent, and review its diff as one unit. Reserve one-dispatch-per-task for work that needs its own judgment, its own tests, or its own review surface.
Everything you paste into a dispatch prompt — and everything a subagent prints back — stays resident in your context for the rest of the session and is re-read on every later turn. Hand artifacts over as files.
Waiting on dispatched subagents: never poll a wait interface with short timeouts, and never sit in one silent, open-ended wait either. While you have local work — ledger updates, packaging the next review, reading reports — keep working; child results arrive on their own. When you are genuinely idle, wait in bounded stretches (five to ten minutes, where your platform allows), and between stretches post one line of status and reconcile your live children: list them, and chase any that finished without reporting. A bounded stretch keeps nearly all of a long wait's efficiency while guaranteeing a stuck or lost child is noticed within minutes, not at the end of the session.
1. Dispatch the implementer
Record BASE (git rev-parse HEAD) before dispatching — the review package
and fix-round diffs need it.
- Task brief: before dispatching an implementer, run this skill's
scripts/task-brief PLAN_FILE N— it extracts the task's full text to a uniquely named file and prints the path. Compose the dispatch so the brief stays the single source of requirements. Your dispatch should contain: (1) one line on where this task fits in the project; (2) the brief path, introduced as "read this first — it is your requirements, with the exact values to use verbatim"; (3) interfaces and decisions from earlier tasks that the brief cannot know; (4) your resolution of any ambiguity you noticed in the brief; (5) the report-file path and report contract. Exact values (numbers, magic strings, signatures, test cases) appear only in the brief. Never make a subagent read the whole plan file. - Report file: name the implementer's report file after the brief
(brief
…/task-N-brief.md→ report…/task-N-report.md) and put it in the dispatch prompt. The implementer writes the full report there and returns only status, commits, a one-line test summary, and concerns. - A dispatch prompt describes one task, not the session's history. Do not paste accumulated prior-task summaries ("state after Tasks 1-3") into later dispatches — a real session's dispatch hit 42k chars of which 99% was pasted history. A fresh subagent needs its task, the interfaces it touches, and the global constraints. Nothing else.
- The dispatch carries the no-subagents contract (it is in the implementer template): the implementer never dispatches subagents — not helpers, and never a reviewer. Review arrives from you, after the report. In real sessions, every reviewer a worker spawned duplicated the task review the controller dispatched anyway — a full extra review seat per task.
- If an earlier task parked a finding in the area this task touches, carry a pointer to that ledger entry in the dispatch.
- Record the implementer's agent identity from the dispatch result — fix-loop rounds 1-3 resume this agent.
- Never dispatch multiple implementation subagents in parallel (conflicts).
Template: implementer-prompt.md
2. Handle the report
Implementer subagents report one of four statuses. Handle each appropriately:
DONE: Generate the review package (scripts/review-package PLAN_FILE BASE HEAD, from this skill's directory — it prints the unique file path it wrote; BASE is the commit you recorded before dispatching the implementer — never HEAD~1, which silently drops all but the last commit of a multi-commit task), then dispatch the task reviewer with the printed path.
DONE_WITH_CONCERNS: The implementer completed the work but flagged doubts. Read the concerns before proceeding. If the concerns are about correctness or scope, address them before review. If they're observations (e.g., "this file is getting large"), note them and proceed to review.
NEEDS_CONTEXT: The implementer needs information that wasn't provided. Provide the missing context and re-dispatch.
BLOCKED: The implementer cannot complete the task. Assess the blocker:
- If it's a context problem, provide more context and re-dispatch with the same model
- If the task requires more reasoning, re-dispatch with a more capable model
- If the task is too large, break it into smaller pieces
- If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
Never ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
If the implementer asks questions — before starting or mid-task — answer clearly and completely, provide additional context if needed, and don't rush it into implementation.
3. Review the task
Per-task reviews are task-scoped gates. The broad review happens once, at the final whole-branch review. Never skip the task review, and never accept a report missing either verdict — spec compliance AND task quality are both required. Implementer self-review never replaces the task review; both are needed.
- Hand the reviewer its diff as a file: run this skill's
scripts/review-package PLAN_FILE BASE HEADand pass the reviewer the file path it prints (or, without bash:git log --oneline,git diff --stat, andgit diff -U10for the range, redirected to one uniquely named file). The output never enters your own context, and the reviewer sees the commit list, stat summary, and full diff with context in one Read call. Use the BASE you recorded before dispatching the implementer — neverHEAD~1, which silently truncates multi-commit tasks. Never dispatch a task reviewer without a diff file. - Reviewer inputs: the task reviewer gets three paths — the same brief file, the report file, and the review package — plus the global constraints that bind the task.
- The global-constraints block you hand the reviewer is its attention lens. Copy the binding requirements verbatim from the plan's Global Constraints section or the spec: exact values, exact formats, and the stated relationships between components ("same layout as X", "matches Y"). The reviewer's template already carries the process rules (YAGNI, test hygiene, review method) — the constraints block is for what THIS project's spec demands.
- Do not add open-ended directives like "check all uses" or "run race tests if useful" without a concrete, task-specific reason
- Do not ask a reviewer to re-run tests the implementer already ran on the same code — the implementer's report carries the test evidence
- Do not pre-judge findings for the reviewer — never instruct a reviewer to ignore or not flag a specific issue. If you believe a finding would be a false positive, let the reviewer raise it and adjudicate it in the review loop. If the prompt you are writing contains "do not flag," "don't treat X as a defect," "at most Minor," or "the plan chose" — stop: you are pre-judging, usually to spare yourself a review loop. The task reviewer may report "⚠️ Cannot verify from diff" items — requirements that live in unchanged code or span tasks. These do not block the rest of the review, but you must resolve each one yourself before marking the task complete: you hold the plan and cross-task context the reviewer lacks. If you confirm an item is a real gap, treat it as a failed spec review — it enters the fix loop with the other findings.
Template: task-reviewer-prompt.md
4. The fix loop
The loop triggers when the review reports spec ❌, any Critical or Important finding, or a ⚠️ item you confirmed as a real gap.
Before the loop starts, two routes leave it immediately:
- Record Minor findings in the progress ledger as you go
(
Task <N>: minor (deferred): <one-liner>), and point the final whole-branch review at that list so it can triage which must be fixed before merge. A roll-up nobody reads is a silent discard. Minor findings never enter the loop. - A finding labeled plan-mandated — or any finding that conflicts with what the plan's text requires — is yours to rule on: weigh the finding against the plan text, decide with the spec as the binding authority, and ledger the ruling before you act on it. Do not dismiss the finding because the plan mandates it, and do not dispatch a fix that contradicts the plan without a recorded ruling. Everything else enters the loop. A fix round is one fix dispatch plus one scoped re-review. Five rounds maximum per task:
Rounds 1-3 — resume the original implementer. Send it the open findings verbatim. Its context is intact: it knows the task, the code, and its own choices. If your harness cannot send another message to a live subagent, dispatch a fresh implementer carrying the brief path, the report-file path, and the findings — the report file is the persistent memory either way.
Rounds 4-5 — dispatch a fresh implementer on a more capable model (per Model Selection), with the brief path, the report-file path, the open findings, and this framing: "A prior implementer attempted this task [N] times; you own it now. Read the report file for what was tried." A loop that survives three resumes usually means the implementer cannot see its own problem — fresh eyes and a capability bump in one move.
Every round, either way: the implementer fixes, re-runs the tests covering the amended code, appends its fix report to the same report file, and returns the short contract. Before re-dispatching the reviewer, confirm the fix report contains the covering tests, the command run, and the output; dispatch the re-review once all three are present. Name the covering test files in the fix message — a one-line fix does not need the whole suite.
The re-review is scoped. Run scripts/review-package PLAN_FILE FIX_BASE HEAD
where FIX_BASE is the head the previous review saw, and dispatch
re-review-prompt.md with the findings list, the
brief, the report file, and the printed diff path. The re-reviewer verdicts
each finding ADDRESSED or NOT ADDRESSED and flags new breakage in the fix
diff only. New Critical/Important breakage in the fix diff joins the open
findings list. Out-of-scope observations go to the ledger as deferred
minors — they never extend the loop.
After each round, append to the ledger:
Task <N>: fix round <R>/5 (<X> addressed, <Y> open — <finding one-liners>; commits <a7>..<b7>)
Never fix findings yourself in the controller session — your context stays clean for coordination, and controller fixes skip review.
The breaker. When round 5's re-review still leaves findings open, stop dispatching. Adjudicate each open finding yourself — you hold the plan and the cross-task context the reviewer lacks:
- The reviewer is wrong, or the point is contestable: park it —
Task <N>: parked — <finding> — Ruling: <why the code stands>. The final review sees both sides. - Real, but nothing downstream builds on it: park it the same way, with a ruling that says it's real and deferred.
- Real and load-bearing — a later task builds on it, or it reveals a
plan defect: rule on the smallest change that unblocks the dependent work,
ledger it as
Task <N>: Ruling: <finding> — <what you decided and why>, and carry it into the next task's dispatch. Parking a structural failure silently lets every dependent task build on it. Stop only when the defect leaves every path forward a guess.
Adjudicate only at the cap. Adjudicating earlier to end a loop is pre-judging with a different name. Every adjudication is a ledger entry — a silent discard is forbidden.
5. Complete the task
When the review comes back clean — or every open finding is parked with a ruling at the cap — append the completion line to the ledger in the same message as your other bookkeeping:
Task <N>: complete (commits <base7>..<head7>, review clean)Task <N>: complete (commits <base7>..<head7>, <K> parked)after a tripped breaker
Then mark the todo complete and move on. Never move to the next task while the review has open Critical/Important issues that are neither fixed nor parked-with-ruling at the cap.
Final Review
The final whole-branch review gets a package too: run
scripts/review-package PLAN_FILE MERGE_BASE HEAD (MERGE_BASE = the commit the
branch started from, e.g. git merge-base main HEAD) and include the
printed path in the final review dispatch, so the final reviewer reads
one file instead of re-deriving the branch diff with git commands. Dispatch
on the most capable available model (see Model Selection), using
superpowers:requesting-code-review's
code-reviewer.md. Point it at
the ledger's deferred-minor and parked lines so it can triage which must be
fixed before merge.
If the final whole-branch review returns findings, dispatch ONE fix subagent
with the complete findings list — not one fixer per finding.
Per-finding fixers each rebuild context and re-run suites; a real
session's final-review fix wave cost more than all its tasks combined.
Then run exactly one scoped re-review of the fix wave
(scripts/review-package PLAN_FILE FIX_BASE HEAD over the fix range,
re-review-prompt.md).
Adjudicate any residual findings as in the task loop's breaker: park with
rulings, or rule on the load-bearing ones and ledger what you decided. Only
the four classes above stop you here. There is no second fix wave —
residual load-bearing findings surface to your human partner when
finishing-a-development-branch presents the options.
Finish
Before you delete anything, collect every ledger line containing Ruling: —
preflight rulings, parked findings, breaker adjudications, all of them — into
your final message under "Rulings I made", in the order you made them, each
with what it costs if wrong. The list is exhaustive: if the ledger holds a
ruling, the list holds it. That list is the only place the decisions you
took on your human partner's behalf reach them — they read it and rework
whatever you got wrong. A ruling that dies with the workspace was a decision
made in secret.
When the final whole-branch review is clean and its fixes are merged,
delete this plan's workspace (rm -rf <workspace>) — the git history is
the record now. Sibling directories belong to other plans; leave them
alone.
Use superpowers:finishing-a-development-branch.
Common Rationalizations
| Excuse | Reality |
|---|---|
| "Close enough on spec compliance" | Reviewer found spec gaps = not done. Fix or hit the cap and adjudicate — those are the only exits. |
| "I'll fix it myself, dispatching is overhead" | Controller fixes pollute your context and skip review. Resume the implementer. |
| "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. |
| "The reviewer will just find something new anyway" | Scoped re-reviews verify fixes; they cannot wander. New findings on untouched code go to the ledger, not the loop. |
| "This finding is obviously wrong, I'll drop it" | You adjudicate only at the cap, and every ruling is a ledger entry. Silent discards are forbidden. |
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
| "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
Example Workflow
ARGUMENTS: Execute docs/superpowers/plans/2026-09-30-agent-storage-roots.md task-by-task
Base directory for this skill: /home/h/.claude-work/plugins/cache/claude-plugins-official/superpowers/b36e0829c6d0/skills/using-git-worktrees
Using Git Worktrees
Overview
Ensure work happens in an isolated workspace. Prefer your platform's native worktree tools. Fall back to manual git worktrees only when no native tool is available.
Core principle: Detect existing isolation first. Then use native tools. Then fall back to git. Never fight the harness.
Announce at start: "I'm using the using-git-worktrees skill to set up an isolated workspace."
Step 0: Detect Existing Isolation
Before creating anything, check if you are already in an isolated workspace.
Submodule guard: GIT_DIR != GIT_COMMON is also true inside git submodules. Before concluding "already in a worktree," verify you are not in a submodule:
If GIT_DIR != GIT_COMMON (and not a submodule): You are already in a linked worktree. Skip to Step 2 (Project Setup). Do NOT create another worktree.
Report with branch state:
- On a branch: "Already in isolated workspace at
<path>on branch<name>." - Detached HEAD: "Already in isolated workspace at
<path>(detached HEAD, externally managed). Branch creation needed at finish time."
If GIT_DIR == GIT_COMMON (or in a submodule): You are in a normal repo checkout.
Has the user already indicated their worktree preference in your instructions? If not, ask for consent before creating a worktree:
"Would you like me to set up an isolated worktree? It protects your current branch from changes."
Honor any existing declared preference without asking. If the user declines consent, work in place and skip to Step 2.
Step 1: Create Isolated Workspace
You have two mechanisms. Try them in this order.
1a. Native Worktree Tools (preferred)
The user has asked for an isolated workspace (Step 0 consent). Do you already have a way to create a worktree? It might be a tool with a name like EnterWorktree, WorktreeCreate, a /worktree command, or a --worktree flag. If you do, use it and skip to Step 2.
Native tools handle directory placement, branch creation, and cleanup automatically. Using git worktree add when you have a native tool creates phantom state your harness can't see or manage.
Only proceed to Step 1b if you have no native worktree tool available.
1b. Git Worktree Fallback
Only use this if Step 1a does not apply — you have no native worktree tool available. Create a worktree manually using git.
Directory Selection
Follow this priority order. Explicit user preference always beats observed filesystem state.
-
Check your instructions for a declared worktree directory preference. If the user has already specified one, use it without asking.
-
Check for an existing project-local worktree directory:
If found, use it. If both exist,
.worktreeswins. -
If there is no other guidance available, default to
.worktrees/at the project root.
Safety Verification (project-local directories only)
MUST verify directory is ignored before creating worktree:
If NOT ignored: Add to .gitignore, commit the change, then proceed.
Why critical: Prevents accidentally committing worktree contents to repository.
Create the Worktree
Sandbox fallback: If git worktree add fails with a permission error (sandbox denial), tell the user the sandbox blocked worktree creation and you're working in the current directory instead. Then run setup and baseline tests in place.
Step 2: Project Setup
Auto-detect and run appropriate setup:
Step 3: Verify Clean Baseline
Run tests to ensure workspace starts clean:
If tests fail: Report failures, ask whether to proceed or investigate.
If tests pass: Report ready.
Report
Quick Reference
| Situation | Action |
|---|---|
| Already in linked worktree | Skip creation (Step 0) |
| In a submodule | Treat as normal repo (Step 0 guard) |
| Native worktree tool available | Use it (Step 1a) |
| No native tool | Git worktree fallback (Step 1b) |
.worktrees/ exists | Use it (verify ignored) |
worktrees/ exists | Use it (verify ignored) |
| Both exist | Use .worktrees/ |
| Neither exists | Check instruction file, then default .worktrees/ |
| Directory not ignored | Add to .gitignore + commit |
| Permission error on create | Sandbox fallback, work in place |
| Tests fail during baseline | Report failures + ask |
| No package.json/Cargo.toml | Skip dependency install |
Common Rationalizations
| Excuse | Reality |
|---|---|
| "I'm obviously not in a worktree — no need to check" | Run Step 0. Harness-created isolation and submodules both fool eyeballing; the detection commands settle it. |
"git worktree add is quicker than hunting for a native tool" | A native tool (e.g. EnterWorktree) owns placement, branching, and cleanup. Bypassing it is the #1 mistake — it creates phantom state your harness can't see or manage. |
| "The worktree directory is surely ignored already" | Run git check-ignore. An unignored worktree directory commits the whole tree into the repo. |
| "Any directory name works" | Explicit instructions beat an existing project-local directory, which beats the .worktrees/ default. |
| "The workspace is fresh — baseline tests can wait" | A dirty baseline makes every later failure ambiguous. Run the tests now; proceeding past failures is your human partner's call. |
ARGUMENTS: Isolated workspace for the agent-storage-roots plan
Setup complete and Task 1 is running. Where things stand:
Isolation: worktree at .claude/worktrees/agent-storage-roots, branch worktree-agent-storage-roots, based on 66c0dcfc6. Baseline: 12,774 tests, one failure — cmd/entire/cli/api TestClient_Request_RespectsCallerContentType, which passes 5/5 in isolation and lives in a package this plan never touches. Recorded as a flake, proceeding.
The pre-flight scan earned its keep. It caught three real defects in my own plan before a single dispatch:
- Task 5 could not have compiled. Its test imported
claudecodefrom insidepackage agent— andclaudecodeimportsagent. Import cycle. Rewritten to extend the file's existingsessionDirAgentmock with a roots-declaring variant, which tests the same behaviour change without the cycle. Real-agent coverage still lands via Task 9's cross-agent table and Task 10. - Task 9's source guard would have failed immediately. I listed 4 files reading the relocation variables; reality has 7.
CODEX_HOMEis also read bycodex/hook_root.goandcodex/trust.go, and bothe2e/agents/codex.goande2e/agents/pi.goread one. Scoped the guard tocmd/**(the e2e files set the variable for a spawned agent — the opposite act), and added Step 3b routing the two Codex files throughresolveCodexHome, since each re-implements the same fallback. - Task 1 named a test double that doesn't exist (
stubAgent; it'smockAgent, pointer receiver).
All three are ruled and recorded in the ledger with what each costs if I got it wrong.
Also amended before execution: your narrowing decision removed SessionBaseDirUnder and SessionBaseDirCandidates entirely — nothing consumed them once discovery stayed single-rooted. The capability is now exactly the two methods you approved, and the classification-vs-discovery rule is written into the spec, the plan's global constraints, and the agent-guide text so a future integration can't quietly widen it.
All 10 briefs are extracted. I'll keep going through the task loop and report when the work is done or something needs your call.
Task 1 committed as 41ba04f68; review is running. Two notes worth keeping:
- The implementer hit the
git stash -utrap (shared stash stack across your worktrees). No damage — stack is empty, tree clean, only the expected commit — but I've barred it from every later dispatch. - A second pre-existing flake surfaced:
cmd/entire/cli TestDialLatencies, timing-based. The implementer verified it against a clean tree. Deferred with theapione.
Task 2 running — that's the one that actually fixes CLAUDE_CONFIG_DIR derivation for wclaude/pclaude.
Tasks 1 and 2 are in, both reviewed clean. Task 2b — the harness isolation fix — is running.
Worth surfacing since it's your machine: mise run check currently fails for you (and only you) because the integration harness inherits your shell's CLAUDE_CONFIG_DIR. The CODEX_HOME / COPILOT_HOME / PI_CODING_AGENT_DIR half of that leak is already on main today — your wcodex/pcodex setup has been feeding real transcript trees into integration tests this whole time. 2b closes all four.
The 2b implementer was right and I was wrong — my brief demanded two mutually exclusive things. Resumed with the ruling:
CLAUDE_CONFIG_DIR gets pinned to the empty string rather than a sandbox path. Task 2 made blank equivalent to unset, so that neutralizes your ambient value while leaving HOME as the lever — which is exactly what attach_test.go sets HOME=fakeHome to exercise. The other three get real sandbox directories, since nothing depends on their fallback.
Three commits in (41ba04f68, 2cc922321, d9cf5bece), 2b under review. Tasks 3–10 to go: Codex/Copilot/Pi roots, multi-root ownership, the recorded session dir, and the adopt fix itself.
Four commits in; Task 3's review is running. The implementer caught a genuine defect in my brief — my rolloutRoots() dropped ENTIRE_TEST_CODEX_SESSION_DIR support and broke a Codex subagent test. They verified it with a temporary worktree rather than git stash, which is what I asked for.
Task 4 running, forewarned about the trap Task 3 hit — my briefs have twice rewritten a helper in terms of StorageRoots() and silently dropped ENTIRE_TEST_* override support. Pi has the same shape.
Five commits in. All four agents with a relocation variable now declare storage roots, and each variable resolves to exactly one file — so Task 9's guard ledger is verified correct before it's written. Task 4's review is running; Task 5 (multi-root ownership) is next.
Task 5 running — the first task that actually uses the declared roots, and the one guarded hardest against widening beyond classification.
Six commits in, Task 5 under review. Next up is Task 6 — recording agent_session_dir in session state, which is the half that actually makes cross-instance adopt work.
Task 6 running — the pivotal one. Everything before it corrected derivation; this is the half that records where a session actually lives, so a wclaude process can reason about a pclaude session at all.
Seven commits in. Task 6 — the pivotal one — is committed and under review; I verified the independence property myself rather than trusting the report, since that was the single way this plan could have quietly become worthless. Task 7 next: the adopt fix itself.
Task 7 running — the one that actually kills the unexpected transcript path … is not owned by a registered agent error you hit.
Task 8's premise holds — searchTranscriptInProjectDirs calls GetSessionBaseDir(), which Task 2 already rewired to roots[0]. So the cold-lookup walk follows CLAUDE_CONFIG_DIR today, and Task 8's real work is just the error message plus pinning both halves with tests.
Eight commits in, and the reported bug's fix is now committed. Task 7's review is running.
That deviation was worth catching: it put "CLAUDE_CONFIG_DIR" into session_adopt.go, which Task 9's guard would have rejected two tasks later. Ledger updated with a reason that distinguishes naming the variable in an error message from resolving paths with it — and flagged for consolidation at the final review, since each agent already owns that constant.
Task 7b running. I chose to fix that Important finding rather than park it — deferring would have shipped a second registry of agent knowledge that silently goes stale, and left Task 9's guard ledger carrying an entry whose only job was to excuse it.
Nine commits in. Task 7b's implementer caught something neat: my brief's suggested doc comment would itself have contained "CLAUDE_CONFIG_DIR" and tripped my brief's own acceptance check. It reworded the comment instead.
Tasks 8, 9, 10 remain — cold-lookup error message, the source guard plus docs, and the end-to-end cross-root adopt test.
First fix round of the run. Worth noting the specific trap: the test had to compare against explicit literal strings, not each agent's own constant — asserting RelocationEnvVar() == claudeConfigDirEnv compares a constant to itself and would pass even if every agent returned the wrong variable.
Ten commits in. Task 8 is the smallest change in the plan — its real value is the negative test that pins cold lookup to a single root, which is the guardrail for the design decision you made earlier.
Task 8 under review. Two tasks left: the source guard plus documentation, and the end-to-end cross-root adopt test that proves your original bug is actually dead through the real command.
Eleven commits in. Task 9 writes the guard that keeps this from being undone — and its ledger is the one thing in the plan I verified empirically after every task rather than trusting.
Task 9 under review — I've asked the reviewer to prove both guards can actually fail, in both directions, rather than just observing that they pass. A guard that can't fail is worse than none.
Twelve commits in; Task 10 is the last. I've asked it to confirm the test actually fails without the fix before trusting the pass — a green test nobody has seen fail proves nothing, and this is the one that's supposed to prove your bug is dead.
All thirteen commits are in and Task 10's review is the last per-task gate. The whole-branch review comes after it.
All twelve tasks are complete — 13 commits, 34 files, +1474/−105, every per-task review clean. The whole-branch review is running on Opus with the ledger's deferred items handed to it for triage.
One fix wave covering all three items: the Codex enumeration bug, the deprecated command form in the guide, and re-framing the design rule itself so the same class of mistake can't slip through again.
I'm parking the reviewer's StorageRoots extraction suggestion — it explicitly graded that acceptable to ship, and refactoring four agent files at the very end trades a real regression risk for a maintainability gain that has no deadline.
Fix is clean and well-commented. Holding for the re-review verdict before I present the whole thing.
Done. 14 commits on worktree-agent-storage-roots, clean tree, mise run check green on the final tree.
The fix, in three pieces
1. Agents declare their relocatable roots — new StorageRootProvider capability (StorageRoots() + SessionDirUnder()). Claude Code never honored CLAUDE_CONFIG_DIR for sessions at all; Codex/Copilot/Pi honored theirs only via the running process.
2. Sessions record where they actually live — the mechanism that makes cross-instance work:
3. Adopt trusts the record — your exact error is gone:
Production diff (tests/docs excluded): 17 files, +584/−107. Full diff at .superpowers/sdd/2026-09-30-agent-storage-roots/review-66c0dcfc6..295cc1862.diff; run git diff 66c0dcfc6..HEAD in the worktree for the live version.
Rulings I made on your behalf
- R1 — shipped past a baseline failure (
api TestClient_Request_RespectsCallerContentType); flaky, passes 5/5 isolated, untouched package. Wrong → a real bug stays masked. - R2/R3/R4 — pre-flight found three defects in my own plan: a test double that didn't exist, a Task 5 test that was an import cycle and could never compile, and a guard ledger listing 4 files when reality had 7. Wrong → the registry test proves multi-root against a mock, not the real agent.
- R5 — inserted Task 2b. Your shell's
CLAUDE_CONFIG_DIRwas leaking into spawned CLIs;CODEX_HOME/COPILOT_HOME/PI_CODING_AGENT_DIRhave been leaking onmainalready. Wrong → pins mask a test wanting the ambient value. - R6 —
CLAUDE_CONFIG_DIRpinned to empty, not a sandbox dir, becauseattach_test.gousesHOMEas its lever. Wrong → a future test needing a sandbox path must set it viaExtraEnv. - R8 (supersedes R7) — inserted Task 7b to fix an Important finding rather than park it, killing a second registry of agent knowledge. Wrong → touched four already-reviewed agent files late.
Two parked, both graded acceptable-to-ship: the four near-identical StorageRoots bodies (extraction recommended, no deadline), and agent-integration-checklist.md still using the old "classify/discover" wording while the guide and code comment carry the corrected "containment, never enumeration" rule — a one-line alignment I left rather than opening a second fix wave.
Two things worth your attention
The whole-branch review caught a real bug none of the per-task reviews could see: Task 3 widened Codex's rolloutRoots(), which is right for its three containment callers but wrong for scanFallbackRollouts, which enumerates. A large legacy ~/.codex could exhaust the shared 500ms scan budget and discard matches already found in the correct root. It also diagnosed why I missed it — I'd framed the rule around who reads the result, but there Entire is the reader, not the agent. The rule now reads "containment, never enumeration."
And the honest limitation: sessions that predate this still fail cross-instance adopt until they take another turn and backfill agent_session_dir. Confirmed live — this session's state has no such field yet.
I haven't pushed, and I've left the worktree and its reports in place. Your branch hidde/agent-config-dir-session-stores is untouched and now fully superseded.
Yes — local only. I verified it rather than reasoning from intent:
Where it lives. <git-common-dir>/entire-sessions/<session-id>.json, i.e. inside .git/. Git cannot track anything there (git ls-files confirms), so it never reaches a commit, a branch, or a push.
Every production reference is local — five of them, all in session-state handling:
| Site | What it does |
|---|---|
manual_commit_session.go:689,722 | writes it at session start |
manual_commit_session.go:771 | Codex partial-state repair |
manual_commit_hooks.go:2956 | turn-start backfill |
session_adopt.go:283 | reads it for the containment check |
Nothing serializes it into checkpoint.TaskStep, and no API or control-plane payload touches session.State.
The second-order check matters more, since that path contains your home directory and the literal string .claude-personal. Persistent checkpoint metadata — the part that can be pushed to a remote — carries no absolute-path fields at all:
Files is repo-relative. The TranscriptPath on checkpoint.TaskStep is an input used to read the transcript off disk (agent.ReadTranscriptFile), not a field that gets persisted into the metadata blob.
So adding this field didn't widen what leaves your machine. Two caveats worth stating plainly: the transcript contents do get captured into checkpoints (that's the product working as intended, subject to the redaction pipeline), and session.State also already stored worktree_path and transcript_path locally before this change — so .git/ was already holding absolute paths, and agent_session_dir is consistent with that rather than a new category of data.
Do a round of review:
- Principal engineer
- Security engineer
- Privacy focused round
- Go idiomatic code and namings, stdlib quality doc blocks, no extensive prose
All four running in parallel — principal and security on Opus, privacy and Go-style on Sonnet.
I've told the Go reviewer to be harsh about comment bloat specifically, and to sort every new comment into keep / cut-to-one-line / delete with replacements. That's the one where I expect to be on the receiving end: I wrote several rationale-heavy blocks that probably belong in commit messages rather than above functions.
First round back. Worth flagging immediately: it drifted outside the branch — SessionEndBudgeter, SkillEvents, TranscriptSanitizer and OutOfBandTokenSource are all pre-existing code I didn't touch, so those cuts aren't mine to make here.
Filtering to what this branch actually added, it kept the two new interface comments and validateAdoptSourceTranscript as earning their length, and wants two cut: SessionDirCandidates (18-line comment on a 16-line function, saying "containment not enumeration" three ways) and fallbackScanRoots. It also independently reached the same conclusion as the whole-branch reviewer on the four duplicated StorageRoots bodies — a shared resolveStorageRoots(envVar, defaultRoot) would cut ~40 lines.
Holding for the other three before I act, so it's one consolidated pass rather than three.
Confirmed: 24 occurrences across 5 test files, including ag.GetSessionDir("/home/h/Projects/entire/cli") — this repo's literal path on your machine. entireio/cli is a real remote, so that would enter shared history along with your username and the .claude-personal profile name. Must-fix.
Holding for the principal and security rounds, then one consolidated fix covering all of it.
Spec Revision 2 is written and F2 is running — the layout predicate, the shared ResolveStorageRoots helper, and the two comment cuts, with negative tests (/etc/passwd under root /, paths outside the agent's layout) that are the security property F4 will depend on.
Five fixes in sequence: F2 capability → F3 record the root → F4 validation → F5 privacy scrub and registry-derived guard. Each gets its own review, same as the original run.
F2's review and F3 are running in parallel on disjoint files. The key thing I asked F3 for is a test that mimics a session started in a subdirectory — transcript under a cwd-derived project dir that differs from the worktree root — and confirmation that it fails against the old behaviour. That's the test whose absence let this ship in the first place.
F2 is in. Waiting on F3's rename to finish, then F4 — the security fix itself — which is the one that actually closes the arbitrary-file-read hole.
F4 is running with its tests written as attacks rather than happy paths — root:"/" plus /etc/passwd, relative and non-clean roots, and a type-confusion case where the state claims Claude Code but the path is a Codex rollout. Each must be rejected. F3's review is running alongside it.
Two of five fixes in. Waiting on F4, then F5 (privacy scrub plus registry-derived guard needles), then a fresh security re-review of the whole corrected branch — since the original security verdict was "do not merge until the HIGH is fixed," that verdict needs re-running against the fix, not just the fix's own review.
F4's review is on Opus, prompted to construct the bypasses itself rather than accept that they're closed — including whether rejecting a non-clean root is right or merely strict, and whether the attack tests would pass against the unfixed code. A rejection test that also rejects before the fix proves nothing, which is the failure mode that let the original bug through.
F5 runs alongside: scrubbing your paths out of the test suite, and making the source guard derive its needles from the registry so a fifth agent can't slip past it.
Holding F6 until F5 lands, since both touch session_adopt_test.go. Then F6 closes the symlink bypass, fixes the trailing-slash regression, and corrects the spec paragraph that overstates the residual risk.
F6 running — symlink resolution, the trailing-slash fix, and correcting the spec paragraph that overstated what's bounded. I've told it to name the residual TOCTOU plainly rather than imply the check is airtight, and not to re-plumb the read path through osroot here; that's a larger change worth its own decision.
Once F6 and its review land, I'll re-run the security and privacy lenses against the corrected branch — the original verdicts were "do not merge until the HIGH is fixed" and "fix the fixtures first", so both need re-running rather than assuming the fixes discharge them.
Three closing reviews running: F6 scoped, plus fresh security and privacy passes against the corrected 19-commit branch. The security one is prompted to attack the current code itself — hard links, /proc, symlinked roots, Windows UNC paths — rather than confirm the fixes exist, and to judge whether my rewritten residual-risk paragraph is now honest or still too generous.
All six fixes are in and individually reviewed. Waiting on the two whole-branch verdicts — security and privacy — which are the ones that were blocking.
Two of three closing verdicts in — F6 clean, privacy mergeable. Waiting on the security re-run, which is the one that was blocking.
Both blocking verdicts are now cleared — security: safe to merge, privacy: mergeable. F7 is the last commit: correcting the two places my documentation overstated the protection, plus the untested filepath.Clean line and a privacy doc note.
Worth saying plainly: the security reviewer cleared the code but caught me writing a security comment that claimed more than the code delivers. That's the kind of error that outlives the fix, because the next person reads the comment and trusts it.
Done. 20 commits on worktree-agent-storage-roots, clean tree, mise run check green, all reviews cleared. Nothing pushed.
Where it ended up
Your wclaude/pclaude and wcodex/pcodex setups now work, via two independent halves:
- Agents declare relocatable roots — Claude Code never honored
CLAUDE_CONFIG_DIRfor sessions at all; Codex, Copilot and Pi honored theirs only through the running process. - Sessions record their owning storage root, resolved inside the owning agent's own hook process, so a
wclaudeprocess can reason about apclaudesession instead of asking the wrong environment.
What the review round changed
The four lenses you asked for found problems the thirteen per-task reviews structurally couldn't, and the branch grew from 14 to 20 commits as a result:
The mechanism barely worked. Claude names projects/ after the agent's cwd, not the worktree root. I was recording GetSessionDir(worktreeRoot), so for any session started in a subdirectory or sibling worktree the recorded path never contained the transcript — and the same mismatch silently defeated the multi-root classification from Task 5. I confirmed it from a real session dumped earlier in this conversation. Neither of my tests could catch it: one asserted only a path prefix, the other set the field by hand. Fixed by recording the storage root instead.
I introduced a security hole. Relaxing adoption to trust a recorded directory removed a containment guarantee: agent_storage_root: "/" matched every absolute path, the agent-type check was skipped, and the adopted path is later read off disk into checkpoint content that syncs to the remote. Then the fix's own review found a symlink escape still open behind it. Both closed; the second review confirmed the gates hold against /proc, hard links, intermediate symlinks and drive-relative paths.
Your username was heading into public history — 24 /home/h/ literals, including this repo's real path on your machine. My briefs seeded them.
And I wrote a security comment that claimed more than the code delivers — "win a race against the read," when the read happens in a later process with no re-validation, so there's no race at all. Corrected, along with the fact that the layout tests only require a path component, making more pre-existing files reachable than I'd documented.
Rulings I made on your behalf
R1 shipped past a flaky baseline test. R2/R3/R4 fixed three defects my own pre-flight found in my plan — including a Task 5 test that was an import cycle and could never have compiled. R5 inserted the harness-isolation task after discovering your shell leaks config dirs into spawned CLIs (the CODEX_HOME half has been leaking on main all along). R6 pinned CLAUDE_CONFIG_DIR to empty rather than a sandbox, because attach_test.go uses HOME as its lever. R8, superseding R7, fixed an Important finding rather than parking it.
Two things I deliberately parked, both graded acceptable-to-ship: agent-integration-checklist.md still uses the old "classify/discover" wording where the guide and code now say "containment, never enumeration"; and reading the adopted transcript through an os.Root would close the remaining swap-at-leisure window, which is a larger change than this branch should carry.
The worktree, ledger and all fix reports are still in place if you want to read any of it. Your old hidde/agent-config-dir-session-stores branch is untouched and now fully superseded.
Nothing changes in how you invoke anything — same commands, same flags, no new surface. Here's the practical shape.
It's inert until you install the build
Your installed binary is 0.11.3-nightly…4e2e839c2 — the old one. Hooks run whatever entire is on PATH, so none of this is live until you install from this branch. Once you do, sessions start recording as they run.
What happens without you doing anything
On session start and again on every turn start, entire writes the agent's storage root into its state file:
One environment read, no I/O. That file lives inside .git/ and is never tracked or pushed. It's also your debugging handle — if adoption misbehaves, that field is the first thing to look at.
The case this fixes
You're in wclaude (CLAUDE_CONFIG_DIR=/home/h/.claude-work) and want to adopt a session that pclaude started in another worktree:
Before: rejected with "unexpected transcript path … is not owned by a registered agent", because entire re-derived the session directory from your CLAUDE_CONFIG_DIR and the transcript was under .claude-personal.
Now: the source session's state says where it actually lives, entire checks the transcript sits inside that root and matches Claude Code's layout, confirms the agent type, and adopts. Identical for wcodex/pcodex via CODEX_HOME.
Where you still do something manually
Cold lookup by ID stays deliberately single-rooted — that was your call, and it's right: attach reads what it finds, so surfacing a session from a tree your running agent can't open would hand you something unresumable. So from wclaude, reaching a pclaude session by ID means naming it:
You won't have to guess that. The error now names the directory it searched and the variable to set, instead of a bare "not found".
The one gap
A session that was already running when you install has no recorded root. It self-heals on its next turn — the turn-start hook backfills it. Only sessions that are already dormant stay unadoptable cross-instance, and for those the explicit env-var prefix is the workaround.
What else quietly got better
entire session attach and entire resume now resolve under your actual config dir rather than ~/.claude, which they never did before — that was broken for you independently of adopt. And integration tests no longer inherit your shell's CLAUDE_CONFIG_DIR/CODEX_HOME, so mise run check stops reading your real transcript trees.
Can we think of a way to make entire session attach <id> work magically without having to redefine CLAUDE_CONFIG_DIR?
Yes — and checking the premise first changed my view of it.
entire session attach never spawns the agent. The only subprocess in attach.go is a git commit --amend; attach reads the transcript and records linkage. So the argument we used to keep cold lookup single-rooted — "you'd get a session your own agent can't open" — is really an argument about entire resume, which does launch the agent. I over-applied it to attach.
That reopens the door. Here's the mechanism I'd build.
Self-populating root registry
Entire already resolves the agent's storage root on every session start and turn start. Have that same code path upsert it into a small per-user file under userdirs.ConfigRoot:
Cold lookup then searches active root → default → other known roots for that agent.
Why this rather than the alternatives:
- It learns from what you actually do. No config to write, no flag to remember. After you've run one session under each profile, it knows both — which for you is already true today.
- It only ever contains roots entire itself resolved from a real agent's environment. That's the crucial difference from scanning
$HOME: not a guess, not arbitrary input, and it never picks up a tree entire has never seen. Scanning was the option you rejected earlier, correctly. - It's per-user, not repo-controlled, so it sits in
userdirsand doesn't touch the repo-settings trust boundary.
Keeping it honest rather than silent
The one real hazard survives: attaching a session found in another tree means entire resume on it later will fail. Since attach is user-attended, that's a message, not a reason to refuse:
Magical when it works, and it tells you the one thing you'd otherwise discover later as a confusing failure.
What I'd keep single-rooted
GetSessionDir for writes, resume, and restore. Those genuinely must match the running agent — a write landing in a tree your agent never reads is silently useless, and that reasoning does hold there.
Cost and caveats
Roughly one focused task: the registry file (write path in the two hook sites, read path in searchTranscriptInProjectDirs), plus concurrency-safe upsert and lazily skipping roots that no longer exist. It adds per-user state, which is new — small, but it's a thing that can go stale, and it's empty on a fresh machine until the first session under each profile.
Want me to build it on this branch, or land it as a follow-up after you've merged what's there? I'd lean follow-up: the current branch is reviewed, clean and fixes the reported bug, and this is a distinct feature that deserves its own review rather than riding along at the end of a pass that already grew twice.
Memory write is blocked by your permission settings, so the spec is the record. One caveat worth flagging: docs/superpowers/ is gitignored, so that design exists only on disk in this worktree — if you want it to survive, it needs to move somewhere tracked (a Linear issue, or docs/architecture/).
Where things stand
Branch: worktree-agent-storage-roots, 20 commits off 66c0dcfc6, clean tree, mise run check green, nothing pushed. Security and privacy both cleared on the corrected branch.
Not merged, not pushed — both remain your call.
Left in place: the worktree, the ledger at .superpowers/sdd/2026-09-30-agent-storage-roots/progress.md, and every implementer and review report. The skill I was following says to delete that workspace once the final review is clean; I've kept it because the review rounds turned up enough that you may want to read the reasoning yourself.
Your old branch hidde/agent-config-dir-session-stores is untouched and fully superseded.
The follow-up, recorded
entire session attach <id> finding sessions across instances without the env-var prefix, via a self-populating per-user root registry — with the premise correction that makes it defensible: attach spawns no agent, so the single-root rule I applied to it actually belongs only to resume.
That correction is the part worth carrying forward even if the feature never gets built. It's the second time on this branch that a rule stated in terms of who reads the result turned out to be the wrong framing — the first was the Codex enumeration bug, where the reader was Entire rather than the agent. The durable version is the one the whole-branch reviewer gave: multi-root for containment tests over a path you already hold; single-root for anything that enumerates a directory.
The implementation work is finished — 20 commits, clean tree, all reviews cleared, nothing pushed. There's no in-flight task to resume.
Four things are genuinely open, and all of them are decisions rather than work:
1. Merge and push. Yours by default, and I won't do either without you asking. The branch is worktree-agent-storage-roots.
2. The follow-up design will evaporate. It's in docs/superpowers/, which is gitignored, so it exists only on disk in this worktree. If you want it to survive, it should become a Linear issue or a tracked doc. I can draft either.
3. Two parked minors. agent-integration-checklist.md still says "the extra roots classify, they never discover" while the guide and code now say "containment, never enumeration" — a one-line fix, and leaving the two inconsistent is exactly how the framing error propagates. Separately, reading the adopted transcript through an os.Root would close the remaining swap-at-leisure window; that one is a real change deserving its own review.
4. The SDD workspace. Ledger and ~20 implementer/review reports are still at .superpowers/sdd/2026-09-30-agent-storage-roots/. I kept them deliberately since the review rounds found enough that you might want the reasoning. Say the word and I'll delete it — git history is the record otherwise.
If you meant "continue" as in build the attach registry after all, I'm happy to — I'd just want that stated explicitly, since you'd agreed to defer it and it's a new feature rather than a continuation.
My suggestion, in order: the checklist one-liner (trivial, and it's the wording that caused a real bug), then capture the follow-up somewhere tracked. The rest can wait until you've decided about merging.
G1 running — the registry itself, no callers yet.
One design choice I made that's worth flagging now rather than after: I'm having it record at session start only, not on every turn. The root can't change mid-session, so per-turn writes would add hook latency for nothing. Trade-off is that if the registry file is deleted, a long-running session won't re-learn its root until the next session starts. Cheap and self-healing, but say if you'd rather it refresh per turn.
Both running. One thing G2 forced that's worth knowing: mapping an arbitrary root to that agent's base directory needs SessionBaseDirUnder, the method we deleted earlier in the branch when nothing consumed it. It now has a consumer, so it comes back — YAGNI cutting the other way. GetSessionBaseDir keeps its contract as the active root's base dir, so the cross-agent invariant test still holds.
The regression guard I care about most: with no recorded roots, behaviour must be byte-identical to today. That's every existing user who doesn't run multiple profiles.
Queued for a cleanup pass once G2 lands — can't run two implementers at once without conflict risk. Waiting on G2.
Both running. G2's review is on Opus with one instruction I care about: verify the attach-does-not-resume premise independently. That claim is the entire justification for reversing a narrowing decision made earlier in this same branch — if it's wrong, the widening is unsound and has to come back out.
The other thing G2 surfaced is worth remembering independently of this feature: adding a method to StorageRootProvider silently broke two test doubles' structural satisfaction of the interface. No compile error — they just quietly stopped being providers, which would have disabled whatever they were testing. Go's implicit interfaces make that failure mode invisible, and the repo has no var _ Interface = (*T)(nil) assertions anywhere to catch it.
G4 running. Worth being blunt about what happened there: the "honest notice" I was pleased with in the design was built on a wrong model of what attach does. I assumed it recorded linkage and left the transcript where it lay, so resume would need the original tree. It doesn't — it captures the redacted transcript into the checkpoint, so resume restores from there into whatever root is active. The notice was telling users to do something unnecessary, using a command that doesn't accept a session ID.
The implementer wrote exactly what I specified, and the test pinned the defect rather than catching it. That's twice now on this feature that my reasoning about behaviour was the failure, not the code — the first being the same premise applied to cold lookup.
G4 also adds compile-time interface assertions, which is the structural fix for the silent-breakage hazard rather than just repairing the two doubles that happened to break.
Scoped re-review running across both fix commits, verdicting all seven findings.
One detail I asked it to guard specifically: making the invariant test fail instead of skip when an expected provider isn't one is only correct if Cursor, Antigravity and Factory AI Droid still skip properly — they genuinely aren't providers, so a test that now fails for them would be a regression rather than a tightening.
Done. 24 commits total, clean tree, mise run check green, nothing pushed.
What entire session attach <id> does now
From wclaude, attaching a session that pclaude started:
No env-var prefix. It searches the active root first, then any root Entire has previously recorded for that agent, learned automatically at session start — never a $HOME scan, never config you write. With no recorded roots the behaviour is byte-identical to before, which the regression test pins.
Two things the reviews caught that were mine, not the implementers'
The notice I designed was wrong twice over. It printed CLAUDE_CONFIG_DIR=<root> entire session resume <id>. But entire resume takes a branch, not a session ID — the command couldn't run. And the prefix was pointless anyway: attach writes the redacted transcript into the checkpoint, and resume restores from the checkpoint into whatever root is active. I'd modelled attach as recording linkage and leaving the transcript where it lay. It doesn't. The whole note was removed; the surviving line states only where it was found.
My justification for the widening was false as written. I told the reviewer attach's only subprocess is git commit --amend. OpenCode's FetchTranscript runs opencode export from inside attach's transcript resolution. The conclusion holds — read-only, runs before the widened search, OpenCode isn't a StorageRootProvider — but that reasoning was in the docs where someone would rely on it.
Also worth keeping: adding a method to StorageRootProvider silently unhooked two test doubles with no compile error. There are now var _ StorageRootProvider = (*T)(nil) assertions on all four agents and both doubles, and the invariant test hard-fails instead of skipping. They compiled clean on the first try, so nothing else was already broken.
Open, all yours
Merge and push. The spec still lives in gitignored docs/superpowers/ — its follow-up section now reads as implemented, but it won't survive unless moved somewhere tracked. And the two earlier parked items stand: the checklist wording, and reading the adopted transcript through an os.Root to close the remaining swap-at-leisure window.
That's the hinge, and it holds. Here's what actually happens.
In a plain terminal, it works — and this is the case it helps most
Your shell doesn't export CLAUDE_CONFIG_DIR; only the wclaude/pclaude wrappers do. So:
StorageRoots()returns just[~/.claude]— the default, which for you is a tree you never use.- The active search walks
~/.claude/projects. It's missing or empty, andtranscript.go:150-151swallows the walk error for a non-existent base dir, so that's a clean miss rather than an abort. That swallow is load-bearing here — if it errored, the search would stop before reaching anything useful. - It then iterates the recorded roots, skips
~/.claude(the active one), and searches.claude-workand.claude-personal. - Hit →
Attached session 019ed5fe from /home/h/.claude-personal.
So the plain-terminal case is a stronger motivation than the cross-instance one I pitched this on. Running entire outside any agent is exactly when the environment tells it nothing, and before this change that meant searching a directory you've never written to.
The precondition
The registry is populated by hooks at session start, so:
- Before you install the new binary: empty. Attach in a plain terminal still finds nothing — same as today.
- After: populated once a session has run under each profile. For you that's one
wclaudesession and onepclaudesession, then it's permanent.
It also means a fresh machine, or a profile you haven't used since installing, won't be found until it's been used once. The explicit CLAUDE_CONFIG_DIR=… entire session attach <id> still works as the escape hatch and is unaffected — it just becomes the rare path rather than the required one.
One thing I'd check before trusting it
I haven't exercised this end-to-end with a real populated registry — the tests construct the registry directly rather than letting hooks fill it. The chain from "hook writes at session start" through "plain terminal reads it" is verified per-link but not as a whole. If you want, the cheap confirmation after installing is: run one session under each wrapper, then cat ~/.config/entire/<registry file> and check both roots are there before relying on attach finding them.
Review the branch using sub-agents:
- Principal engineer
- Security engineer
- Right solution for problem
- Idiomatic Go, proper (interface) namings and contracts, stdlib quality doc blocks and not overly prose
All four running in parallel against the full 24-commit diff, each with a grounding gate: no finding survives without a concrete failure — specific inputs, specific wrong outcome. Style preferences and speculative hardening get dropped.
The one I'm most interested in is the third lens, because it's the one nobody has run yet. It asks whether this should exist in this shape at all — specifically whether a far smaller change would have fixed your reported bug, given transcript_path was already in session state, and whether the mid-branch reversal on cold lookup means the rule was wrong from the start rather than correctly excepted.
I've also told the Go reviewer to stay strictly inside the diff — the last one spent most of its output critiquing pre-existing code this branch never touched.
Go lens back. Two concrete cuts, both fair:
The adoptOwnerFromRecordedRoot doc comment is an essay — ~80 lines, including two narrative attack walkthroughs that retell what the regression tests right below already demonstrate. So the root:"/" attack is explained three times: twice in the comment, once in the tests. It wants those two paragraphs cut to one line, keeping the bullet contract and the TOCTOU limitation (which the security review added deliberately and appears nowhere else). That's the right call — I wrote those paragraphs defensively after the security findings, and overcorrected.
The single-rootedness rationale appears verbatim three times — on the interface method and again on both GetSessionBaseDir implementations. Keep it on the interface where the contract lives, one-line the implementations.
Naming, contracts and idiom came back clean: the 4-method grouping is defensible since all four are always implemented and used together, splitting RelocationEnvVar out was right because it keeps the "pure arithmetic, no I/O" contract clean, and the compile-time assertions are justified by a documented real incident. One minor: resolveAgentStorageRoot (strategy) and agent.ResolveStorageRoots are easy to confuse on a grep.
Holding for the other three before I act, so it's one consolidated pass.
The problem-validation lens is the sharpest of the four. Verdict: right solution, overscoped — and the reported bug does work, traced end to end.
Two findings land hard:
The registry has one consumer and doesn't help the command that actually failed. KnownStorageRoots is read only by attach's cold lookup. Adoption's fallback never consults it. So a pre-upgrade session sitting in ~/.claude-work, adopted from a pclaude process, still fails with "re-run with CLAUDE_CONFIG_DIR set" — even though ~/.claude-work is right there in the registry file. Entire has the answer and declines to use it. That also means the "sessions predating this change can't be adopted" limitation I documented as unavoidable is largely dissolvable.
The Codex rolloutRoots widening changes what lands in checkpoints. It gates readRegularRolloutContext and openScopedRollout, not just classification — so a rollout reference pointing into ~/.codex is now read and its content enters a .codex-work checkpoint. Nothing in your reported problem required that.
It also caught that my governing rule was mis-stated from the start: the real axis isn't containment-vs-enumeration, it's does a launch or a write follow? That's the third correction to this framing and the first one that actually predicts the exceptions. And it flagged that claudecode/pi GetSessionBaseDir still carry the old doctrine in comments while their sole caller now deliberately widens — the next maintainer reads the opposite of what the code does.
Two lenses still out. Holding.
Confirmed, and it's my regression. This is the most important finding of the entire session.
main had the guard at registry.go:119:
My Task 5 refactor dropped it:
And pathHasDirPrefix(path, "") appends a separator to the empty string, so it reduces to strings.HasPrefix(path, "/") — true for every absolute path. external.getSessionDir returns resp.SessionDir, nil, so a plugin emitting {"session_dir":""} or omitting the field yields ("", nil) with no error.
Consequence: that plugin claims every transcript. List() is alphabetical, so an early-named one wins. Then shouldSkipForwardedHook sees owner.Name() != ag.Name() for the real agent and skips every hook — no checkpoints written at all, silently. Plus resolveSessionAgentType stamps the wrong agent type on every session.
What's sobering is the miss profile: Task 5's own review, the whole-branch review, and three separate security passes all looked at this function. Every one of them scrutinised the new multi-root path I added. None checked what the refactor removed from the existing path. Reviews anchored on the change, and the deletion was invisible.
The same lens also found two more Importants worth acting on: the adoption comparison isn't genuinely two-factor (root, transcript path, and the agent used for the layout test all come from the same state file, and the code's own comment concedes the type check is "necessarily a no-op" on that branch), and the turn-start refresh comment's reasoning is false — it depends on process env, not worktree, so a hook firing under the active instance overwrites a root that attach just recorded, breaking the adopt that recording was meant to enable.
Security lens is still out. I'll hold for it and then do one consolidated fix wave rather than two passes over registry.go.
Critical fixed — d267ef654, test confirmed failing pre-fix, plus a defence-in-depth guard inside pathHasDirPrefix.
Now wave two with your decisions. Starting with the security one, since it's the most consequential and everything else edits around it.
Stopping before H1 — this changes the plan materially. Two decisive facts:
Their registry.go keeps the dir == "" guard. They split the error and empty checks but preserved both (and added a debug log for the refused-variable case). So the Critical I found an hour ago was purely my branch's regression, in a file they also modified. Rebasing will conflict in exactly that spot.
They did not touch session_adopt.go, session/state.go, or transcript.go at all — those diffs are empty. So your reported bug is still broken on merged main: AgentForTranscriptPath now resolves GetSessionDir under the adopting process's CLAUDE_CONFIG_DIR (.claude-work), and a pclaude transcript under .claude-personal still doesn't match.
What that means for the branch
Duplicated, should be dropped: per-agent relocation resolution (their agent.ResolveHome vs my ResolveStorageRoots), harness env scrubbing (their RelocationEnvVars() scrub vs my Task 2b), the source guard (home_guard_test.go vs my storage_root_guard_test.go), and relocation-variable naming (RelocationEnvVars() vs my RelocationEnvVarNamer). Theirs is better in two respects: it covers Gemini and Factory, and it refuses relative values via userdirs.RequireAbsoluteOverride — which is precisely the robustness nit the security lens raised against mine.
Still unique and still needed: recording AgentStorageRoot in session state, the two-factor adoption validation, SessionPathUnder layout classification (which fixes the cwd-derived project-dir mismatch their GetSessionDir(repoPath) still has), and the known-roots registry for attach.
Roughly half my 25 commits are now redundant, and the surviving half should be rebuilt on their abstractions rather than mine.
My recommendation
Rebase onto merged main, drop the duplicated half, and rebuild the adoption/state/registry work on agent.ResolveHome. That's real rework, but the result is maybe 600–800 lines instead of 1,900, sits on the shared abstraction, and avoids two competing resolvers in one package.
The alternative — finishing the four fixes I was about to start, then rebasing — means polishing code I'd then delete.
I can't git fetch from here (no credentials), so I'm reading their branch ref, not the post-merge main. Before committing to a rebase I'd want you to run ! git fetch origin main so I'm working against what actually landed.
How do you want to play it?
I have their design now, and it's good — in places better than mine. ResolveHome(envVar, defaultRel) with a deliberate no-trim rule (trimming "/tmp/x " would resolve to a directory the agent never wrote to), ~ expansion for the agents that do it themselves, and relative values refused through userdirs.RequireAbsoluteOverride.
It also simplifies my port. Their ResolveHome returns one home per agent, not a list — so my StorageRoots multi-root concept goes away entirely, and the known-roots registry becomes the single, more precise source of "other roots the user actually has" rather than guessing [env, default].
The plan
New branch off merged main, porting only what survives:
AgentStorageRootin session state, recorded from theirresolveClaudeConfigDir()/ResolveHomeresult at session start.SessionPathUnderlayout classification — fixes the cwd-derived project-dir mismatch theirGetSessionDir(repoPath)still has, and which breaks adoption for worktree and subdirectory starts.- Two-factor adoption validation on the recorded root, with the
os.Root-scoped read you asked for. - Known-roots registry, wired into both attach's cold lookup and adoption's fallback.
Dropped as redundant: ResolveStorageRoots, RelocationEnvVarNamer, the storage-root guard, the harness scrubbing, and all four per-agent StorageRoots implementations. Their versions cover Gemini and Factory too, which mine never did.
The 25-commit branch stays on disk as reference — nothing is lost.
One thing I need from you
I can't git fetch (no credentials in this sandbox), so my origin/main is stale at exactly my old base. I'm reading their branch ref, not what actually landed on main. Please run:
Then I'll branch from the real merged main. If you'd rather not, say so and I'll base off origin/fix-hardcoded-claude-home instead — it's the merged content, but any commit that landed alongside it would be missed.
P1 running against the new base.
Two things carried forward from the review rounds that would otherwise have been lost in the rebase: the turn-start refresh now has a guard from the outset (only overwrite when the newly-resolved home still contains the transcript) rather than the unguarded version that a principal reviewer caught clobbering records; and the independence requirement between AgentHome and TranscriptPath is stated up front, since deriving one from the other silently turns the whole adoption check into a tautology.
The port is also materially smaller than what it replaces. My old StorageRootProvider carried four methods largely to re-derive what agent.ResolveHome now does centrally — including the probe for config dirs set inside Claude's settings files, which my version couldn't see at all. The new capability is two methods, and SessionHome() just forwards to each agent's existing resolver rather than reimplementing the policy.
P2 is the one that fixes your bug, and it's carrying every lesson the review rounds bought — rather than rediscovering them.
Baked in from the start this time: the clean-compare (so a trailing slash doesn't silently reject a legitimate home), symlink resolution on both sides (resolving only the transcript breaks legitimately symlinked config dirs), the layout test as a real gate rather than a prefix check, the agent-type check honestly labelled tautological on the recorded-home branch instead of pretending it isn't, and the os.Root-confined read you chose — which closes the swap-at-leisure window the validation alone can't, since the read happens in a later process.
Also persisting the resolved path, which the last security pass caught: validation was resolving symlinks and then storing the unresolved form, so every later reader re-followed the link anyway.
Two down, one to go — the registry for attach, layered after main's probe.
Worth noting one thing P2 surfaced that I'd have got wrong: it tried to confine the read with a bare os.OpenRoot, and the repo's own TestRootBasesAreTrusted anchor-trust guard rejected it on the first mise run check. It routes through SessionStore instead. That's the codebase enforcing a filesystem contract I'd specified loosely in the brief — the guard caught it, not a reviewer.
Waiting on the security review of P2, then P3 (the attach registry) is the last piece.
P2b running. The pattern here is worth naming, because it's now happened twice in the same area: a partial mitigation reads as a complete one. P2 confined the reads it could see from where it was working — the condensation paths — and the finalize path, which is the one whose bytes actually reach the pushed branch, sat one file away. I asked for "confine the read" and didn't ask for an enumeration of every read site, so nobody counted them.
P2b now has to account for every ReadTranscriptFile call — confined, or a stated reason it can't carry an adopted path.
Two of the LOWs are the same shape as things caught earlier in the project: a helper that fails open while its comment claims it refuses to proceed, and a guard satisfied by a file-level grep rather than by the contract actually holding. Both are cases where the code and its description had drifted apart, which has been the recurring failure mode of this whole effort far more than logic errors have.
P2c running, with the scope bounded rather than taken at face value. Two facts should shrink it a long way from "ten implementations": AgentHome is only ever set for the five agents implementing AgentHomeProvider, so the other six can't reach the confined path at all; and bytes-based extraction variants already exist, which may let the caller do one confined read and pass bytes — no interface change.
I've told it that widening the core Agent interface for a property only five agents can have is probably the wrong shape, and to stop and say so if the right mechanism turns out materially bigger than the brief assumes.
Current state: three commits on agent-session-root-recording off merged main, build and vet clean. P3 — the attach registry — is the one remaining piece after this.
Both running — P2c's review and P3, the final piece.
P3 inherits every mistake the abandoned branch made, pre-corrected: no osroot.Shared for the existence probe (the repo documents, with measurements, why that exhausts RLIMIT_NOFILE for every other anchor), the idempotence test asserts on mtime rather than content so it actually proves no write happened, and the "found elsewhere" notice states a fact instead of printing the command that was wrong twice over last time.
The review has the harder job. I've told it explicitly not to accept the implementer's enumeration and to build its own, because both prior confinement rounds missed exactly one site — P2 missed finalize, P2b missed the per-agent readers. The claim under test is that "last" is finally true.
P4 running on the three holes. P5 — making confinement structural rather than opt-in — follows once it lands, since both touch the same call sites.
Where this stands: five commits on agent-session-root-recording off merged main, build clean, nothing pushed. Your original bug is fixed and covered by a test that fails without the fix. The remaining work is all security hardening that the review rounds uncovered.
The honest summary of this session is that the feature was the easy part. What's actually consumed it is that entire session adopt reads a session-state file from a directory the user merely names with --from, and treats every path in it as trustworthy. Three rounds of confinement each closed what they could see, and a fourth enumeration found a path needing no symlink at all — TaskRecords[].DeclaredTranscriptPath, read whole into a pushed checkpoint, on main today.
That one is worth separating in your head from the rest: it isn't mine, and it affects anyone using adopt, not just the multi-profile case. You chose to fix it here, which I agree with, but it's also the thing I'd want someone upstream to know about independently of whether this branch ever merges.
P5 running. Two instructions in it matter more than the rest.
The cross-agent test must fail, not skip. An earlier test in this project used if !ok { continue }, which silently passed over exactly the agents it was supposed to police — that's how the opt-in gap survived a dedicated invariant test.
Every ledger entry must be verified, not copied. It's seeded from P4's enumeration, but a ledger whose stated reasons are false is worse than none: it converts "nobody checked" into "someone checked and signed off", which is harder to revisit.
And it has to prove the guard works by planting a bare read and showing it caught — the same standard the regression tests have been held to. A guard nobody has seen fail is indistinguishable from a guard that can't.
While that runs — this is a natural stopping point, so here's where things actually stand.
Branch: agent-session-root-recording, 7 commits off merged main (ba41caf7b), clean tree, nothing pushed. The abandoned 25-commit branch is still on disk as reference.
Your bug is fixed, with a test that fails without the fix, and verified end to end through the real runAdopt.
Roughly half the original work was thrown away when trail 1308 merged — and that was the right call. Their ResolveHome is better than my ResolveStorageRoots: it covers Gemini and Factory, refuses relative values, and has a probe that reads a config dir set inside Claude's settings files, which my version couldn't see at all.
What actually consumed this session wasn't the feature. It was discovering that entire session adopt treats every path in a --from worktree's session state as trustworthy. Five rounds of security work followed, and the pattern each time was the same: a fix closed the sites it could see, and a fresh enumeration found more. The last round found a path needing no symlink at all.
Two things from that are worth separating from this branch:
The CRITICAL is pre-existing and affects everyone using adopt, not just multi-profile users. TaskRecords[].DeclaredTranscriptPath is read whole into a pushed checkpoint with no validation, on main today. I verified it at ba41caf7b. We're fixing it here, but someone upstream should know independently of whether this branch merges.
The structural fix immediately paid for itself — binding the capabilities revealed that Claude Code's confinement was a hardcoded special case, so four rounds of hardening had been narrower than they looked for the most-used agent.
Once P5's review returns I'd suggest stopping and letting you decide on merge. If more findings come back, they'll be in the same family, and at some point the honest move is to ship the bug fix and track the adopt trust model as its own piece of work rather than chasing it indefinitely inside a feature branch.
P6 running. Three of its findings are the same failure wearing different clothes, and it's worth naming because it's been the through-line of this entire session:
A check that looks like it covers something, and doesn't. The guard's pathspec excluded the packages that keep breaking. The binding test's registry is hand-maintained, so a sixth agent passes silently — the exact "skips what it should police" flaw that test was written to correct. And two ledger entries state reasons that aren't true, which is worse than no entry because it converts nobody checked into someone signed off.
The planted-read proof is the sharpest instance: it was genuinely run, genuinely passed, and tested nothing — because it planted the read inside the scope that was already covered.
One finding is a straightforward regression I'd want fixed regardless of security: confining GetTranscriptPosition made a Pi mid-file parse error return an error where it previously returned a valid count, so hasNewTranscriptWork reports no work — silently skipped checkpointing, on legacy sessions that never needed confinement at all.
After P6 I'd stop and hand this to you. The remaining items are documented, the bug is fixed, and the adopt trust model is a bigger piece of work than a feature branch should absorb.
Resumed. I've flagged one thing specifically: the modified-file list is confined to cmd/entire/cli/agent/ — no strategy/, no review/. If widening the pathspec to cmd/**/*.go genuinely surfaced nothing new in the packages where all four prior bugs lived, that's a surprising result, not a clean one. It would mean either the needles are too narrow to match what's there, or the widening didn't take effect.
I've asked for that stated explicitly with evidence rather than passed over, and for the planted-read proof to be re-run in strategy/manual_commit_hooks.go specifically — planting it in agent/ is what made the original proof vacuous.