Resume Sessions Without Branch Checkpoints

I'd like to make "entire resume" work without a branch. This way I can share a session and then someone can resume it, even if it was just investigative
Done. entire session resume now works without a branch, and entire session share creates the thing to resume.
What changed
session resume <target> (resume.go) routes its positional argument through the same auto-detect ladder checkpoint resume already used, so a checkpoint ID resolves with no branch: the session log is restored in place, and the ref is fetched on demand if it isn't local. Branch targets behave exactly as before.
session share [session-id] (session_share.go) condenses now and pushes — your "checkpoint create plus the extra push". It resolves the caller's session, condenses via the new strategy.CondenseForShare, pushes through strategy.PushSharedCheckpoints, and prints the entire session resume <id> line to hand over.
Three findings that shaped it
Branchless resume already existed, hidden inside checkpoint resume. And an investigative session already gets a persistent checkpoint — CondenseAndMarkFullyCondensed runs at session end with no commit needed. The real gap was discovery and transport: the checkpoint ID lives only in your local session state, and the ref only leaves the machine when pre-push drains the queue.
CondenseForShare is a third eager-condense variant because neither existing one is safe against a live session — CondenseSessionByID sets Phase = IDLE and can delete the state file, CondenseAndMarkFullyCondensed marks FullyCondensed and deletes the shadow branch. Share changes none of those and keeps the shadow branch, since the session keeps writing to it.
The push reuses the existing machinery rather than calling git directly, because manual_commit_opf_refs.go runs OPF re-redaction before anything is sent. pushSharedCheckpoint calls EnsureRedactionConfigured first — without it the OPF gate reads as "off" and would ship un-redacted content.
Two judgment calls, both flagged for you
A session with pending FilesTouched is refused with "commit first". Those files belong to the checkpoint your next commit writes, and prepareEagerCondensation already forbids condensing them. Investigative sessions touch no files, so this never bites the case you asked for — but if you'd rather share steal the carry-forward, that's a one-line change.
Share refuses to guess which session is yours: it requires a positively identified caller or an explicit ID, rather than falling back to "most recent session around" the way read-only commands do. Sharing publishes a transcript, so a wrong guess publishes someone else's. Relatedly I classified session share as user-owned in agentHelpClassification — agents won't reach for it. That's the product choice CLAUDE.md asks to have reviewed.
Verification
mise run lint clean, 677 integration tests pass, canary passes, go test ./docs/development passes, GOOS=windows go vet clean. 20 new tests across four files. I confirmed the three new resume tests genuinely fail against the pre-change code, and the delivery test pushes to a real bare remote where another clone could fetch it.
Two unit failures are pre-existing and unrelated: TestDialLatencies (network-flaky under parallel load — passes in isolation and passed on the re-run) and TestOpenCodeSeedRepoPlantsDeps (the known local red that needs e2e/bootstrap).
Nothing is committed. One gap worth naming: sharing mid-session advances the checkpoint window, so a later commit's checkpoint carries only what came after the share — the same tradeoff the session-end path already makes, but mid-session it's more visible.
Found it — soph/checkpoint-create, singular, which is why my earlier search came up empty. I was wrong to tell you the command didn't exist; it just wasn't on main.
Answer to your question: it handles a files-touched session by never touching session state. CreateSnapshotCheckpoint loads the state fresh, condenses that in-memory copy, and never saves it back — so FilesTouched, StepCount, CheckpointTranscriptStart and any pending reservation are untouched. The next commit condenses the same range again and the two checkpoints overlap, deliberately.
The part that's specifically about touched files is noCommitAttribution: attribution diffs the shadow tree against HEAD assuming HEAD holds the work just committed. With no commit, HEAD predates the uncommitted changes, so every one would score as a human removal. So it skips attribution rather than refusing. It also skips skill telemetry, since the unsaved state means the next commit's condensation emits those events.
What I changed as a result
Merged origin/soph/checkpoint-create (clean, one small auto-merge in agent_help_cmd.go), deleted CondenseForShare and its tests, and rebuilt session share on CreateSnapshotCheckpoint — which is what you asked for at the start.
My design was wrong on four counts, and the snapshot approach fixes all of them. It advanced the checkpoint window, so the next commit's checkpoint lost the pre-share conversation — that tradeoff I flagged is now gone. It refused files-touched sessions — unnecessary. It used the reservation machinery, so an interrupted write could wedge a session. And it computed attribution with no commit backing it, which is the mis-attribution bug the branch's comment describes.
Share now also reuses resolveCheckpointCreateSession, so the caller-identification rule is one implementation rather than my parallel copy.
What's still mine: resolveShareRemote (fail-closed, resolved before the write so you're never left with a checkpoint and no explanation), PushSharedCheckpoints (backend dispatch), and printShareResult. Plus the session resume routing, unchanged.
Verification: lint clean, 677 integration tests pass, canary 4/4, docs tests pass. Test count is down from 20 to 13 because I dropped the eight CondenseForShare tests and didn't duplicate what checkpoint_create_test.go and manual_commit_snapshot_test.go already cover. TestDialLatencies still reds under full-suite parallel load and passes 3/3 alone — pre-existing, in placement_latency.go, untouched by either branch.
One thing to decide: I made a WIP commit on worktree/lucky-cloud-358c so the merge could proceed, so this work now depends on soph/checkpoint-create landing first. If you'd rather keep them independent I can rebase the share half out, but shipping them in order is simpler.
PR #2632 hasn't actually merged — I checked the API directly rather than trusting the cached listing:
merged: false,mergedAt: null, state open, basemainrefs/heads/mainon the remote is89c2616087and contains neithercheckpoint_create.gonormanual_commit_snapshot.gomergeable_state: "blocked"— one failing check, Entire Gates, plus 1 unresolved review comment (run)
So something else may have merged, or the merge didn't land. Worth a look before I assume anything.
What I did anyway, because it's right either way: the branch had moved since I merged it — I had b7a1493, its head is now f3b1fcd. It was rebased onto a newer main and gained a commit: "Fail snapshot checkpoints on redaction errors instead of dropping the transcript." That adds condenseOpts.failOnRedactionError, set only by CreateSnapshotCheckpoint, so a redaction failure errors instead of writing a transcript-less checkpoint.
That one matters more for share than for checkpoint create, since share publishes: without it you could push a checkpoint whose transcript was silently dropped, and the recipient would resume into nothing. share inherits it for free — no change needed on my side.
I merged the current head in. Three conflicts, all in checkpoint-create's own files (manual_commit_condensation.go, manual_commit_snapshot.go, and its test) — artifacts of the rebase, not real disagreements. I verified none of my commits had ever touched those files, took the branch's version for all three, and confirmed the result is byte-identical to the branch head.
Verification on the updated base: lint 0 issues, go vet clean, both changed packages pass in full, 677 integration tests, canary 56/56 and 4/4, docs tests pass. All 13 of my tests pass, as do the branch's 8. TestDialLatencies passed this time too, confirming it's load flake.
Where things stand: my worktree branch is now soph/checkpoint-create@f3b1fcd + my work, so it's ready the moment #2632 lands. Two things for you:
The dependency is real — session share calls CreateSnapshotCheckpoint, so this can't go in before #2632 unless I split the resume half out. The session resume change is genuinely independent and could ship on its own today if you'd rather not wait on a blocked PR.
And my branch carries two WIP commits plus two merge commits, the first of which merged the pre-rebase head. Before this becomes a PR I'd squash the WIP commits and drop the redundant merge — say the word and I'll tidy it.
Rebased onto the merged main and reconciled with what actually shipped.
#2632 changed in review before merging, and one of the changes reversed the answer I gave you earlier. The merged CreateSnapshotCheckpoint now refuses a session with pending file changes (ErrPendingFileChanges) rather than ignoring them. The rationale in the merged code is better than either of our earlier positions: the next commit checkpoints that work with attribution, so a snapshot would just publish a weaker duplicate. Detection uses resolveFilesTouched, not state.FilesTouched, so an edit made earlier in the current turn — visible only in the live transcript — still counts, while paths outside the worktree are dropped so planning sessions aren't caught.
Two other changes landed that I hadn't seen: it now runs under the session state lock (the per-session redaction prefix cache is written in two steps, and an unlocked writer racing a commit's condensation corrupts every later checkpoint of that session), and failOnRedactionError makes a redaction failure an error instead of a silently transcript-less checkpoint.
What I changed as a result. My help text said "uncommitted file changes are fine — they are simply left unattributed," which was true of the pre-review branch and is now wrong; same in cli-conventions.md. Both now describe the refusal. I added describeShareCheckpointError so share turns the sentinel into "commit first, then share" with the attribution reason, instead of surfacing a bare condition — plus a test pinning it. The strategy-level refusal is already covered by #2632's own tests, so I didn't duplicate that.
History is clean now. I cherry-picked my two WIP commits onto 3299a0f9d3 and soft-reset, so the branch is exactly main plus my 12 files, uncommitted — no WIP commits, no redundant merges. The two files main had moved independently (resume.go, agent_help_cmd.go) three-way merged without conflict. Backup branch deleted.
Verification on merged main: lint 0 issues, go vet clean, GOOS=windows go vet clean, both changed packages pass in full, 692 integration tests, canary 56/56 and 4/4, docs tests pass.
One consequence worth flagging: with the refusal in place, share is now strictly for sessions that changed nothing — which is the investigative case you started from, so it serves the original ask. But if you edit a file mid-investigation you must commit before you can hand it off. If that turns out to be too strict in practice, the place to revisit it is CreateSnapshotCheckpoint, not share.
No — nothing is committed or pushed. All 12 files are still working-tree changes.
That's deliberate: you haven't asked me to commit, and when I did make WIP commits earlier it was only so I could merge checkpoint-create. I unwound those in the rebase, so the branch is now exactly 3299a0f9d3 (merged main) with my changes uncommitted on top.
Current state:
- Branch:
worktree/lucky-cloud-358c, sitting on merged main, no commits of mine - Upstream: none configured; the branch doesn't exist on origin
- Changes: 5 new files, 7 modified
It's all verified and ready to go — lint clean, full test suites green, rebased on current main. Want me to commit and push it, and open a PR? If so, tell me whether you want it as one commit or split into the two independent halves: the session resume change stands alone and doesn't depend on session share.
Done — everything is pushed now.
Trail 1478 — https://entire.io/gh/entireio/cli/trails/1478 Draft PR #2654 — https://github.com/entireio/cli/pull/2654
Branch soph/session-share off main at 3299a0f9d3, one commit (072df213), 12 files. I named the branch to match the repo convention (soph/*, like soph/checkpoint-create) rather than pushing the herdr worktree branch name; worktree/lucky-cloud-358c still exists locally pointing at main. The post-commit hook captured the session automatically — Entire-Checkpoint: 01M45JZKSGA1GYVY54X961BB70 — and the push sent 4 checkpoint refs.
The trail body calls out the three decisions worth a look: caller-only session resolution, resolving the remote before the write, and refusing pending file changes. It also flags that the two halves are independent if you'd rather split them, and that classifying share as user-owned in agentHelpClassification is a product call.
Not yet clear: findings. entire trail finding list 1478 says "No findings found," but that's because the agent review hasn't run yet — the trail's findings gate reads "pending — waiting for the agent review of the current head." An empty list at this point means not-yet-reviewed, not clean, so I'm not claiming it's clear. Gates right now:
Want me to watch for the review to land and work through any findings, or leave it with you?
What would you like to clarify about the --remote question?
For context on where I'm coming from: the flag currently lets you push the shared checkpoint anywhere, but a recipient only ever looks in two places — the elected checkpoint sync remote, then origin (CheckpointReadRemotesWithElection, checkpoint_read_remotes.go:77-93). So --remote upstream can succeed and still print a resume command that resolves for nobody.
I added the flag by analogy with doctor migrate --remote, but that one is pushing refs for their own sake, not minting a handoff command that has to work on someone else's machine — so the analogy may just be wrong.
Happy to answer anything about how the read path picks candidates, who realistically hits a non-origin topology, or whether there's a use case for the override I haven't considered.
That answers it, and it argues against my flag more strongly than the bot did.
--remote exists in exactly two places: mine (session_share.go:60) and doctor migrate (doctor_migrate.go:106). That's the whole precedent — and I copied it without checking that the two commands have different obligations.
The supported way to direct checkpoint traffic is the checkpoint_push_remote setting, and the comment on ResolveCheckpointSyncRemote spells out why that matters here:
"Either way the result is readable from the same clone: the read paths (resume, explain, discovery) consult the elected remote first and fall back to origin as a read-only legacy tier."
The setting moves push and read together, because CheckpointReadRemotesWithElection resolves its candidates through that same election. A CLI --remote can only move the push. So the flag is structurally incapable of keeping the handoff readable — it's not a matter of warning harder.
And the fork case the flag might seem to serve is already handled: the comment says that topology "is served automatically by capture, or explicitly by checkpoint_push_remote."
doctor migrate --remote is defensible by comparison — it's a one-off backfill of refs you already own, and the only place it's referenced is its own error strings telling you to pass it when election fails (doctor_migrate.go:167,170). It's an escape hatch for a broken election, not a routing choice. session share mints a command that has to work on someone else's machine, so the same hatch becomes a footgun.
So I'd drop the flag: push to the elected sync remote, and if the election fails, report that rather than offering an override that desynchronises push from read. Anyone who genuinely needs a different target sets checkpoint_push_remote, which fixes both sides.
Want me to go ahead on that basis and fix all four comments? The other three are unambiguous: strict delivery result for the git-branch path, EnableHomeProbes() in runCheckpointResume (which also closes the pre-existing gap in checkpoint resume), and NewSilentError on the OPF cancel.
yeah there should be no --remote here, you "share" through the same way as all other checkpoints (and the UI allows you to share the session with different controls then)
I looked at both. Trail findings were clean and CI green (19/19) — but the PR had 5 inline comments (4 from Copilot, 1 from Cursor Bugbot duplicating the first). All four distinct issues were real; I verified each against the code and fixed them.
1. Git-branch share reported false success — the serious one, and on the default backend (PrimaryIsRefs is true only when explicitly configured). prePush's own comment says "err... is fail-soft and nil even when the remote refused the ref", yet it returns nil regardless of deliveredCount/anyFailed, and also when the sync gate skips. My code read that nil as success and printed a resume command for a checkpoint that never left. I extracted opfRewriteV1IfEnabled and deliverV1Refs out of prePush and added PushCheckpointBranch — the git-branch analogue of PushQueuedCheckpointRefs: same gates, same helpers, errors when nothing lands. Both backends now report a confirmed count, so Counted is gone.
2. EnableHomeProbes() bypassed — worse than reported. It lives in runResume, which my change skipped for checkpoint and commit targets, and runCheckpointResume never called it either. So entire checkpoint resume <id> has had this gap since it shipped; I just extended it. Fixed in runCheckpointResume, which closes both.
3. --remote dropped, per your call. Your reasoning also turned out to be what the code already says: ResolveCheckpointSyncRemote's comment notes the fork case "is served automatically by capture, or explicitly by checkpoint_push_remote... the read paths consult the elected remote first and fall back to origin." The setting moves push and read together; a flag could only ever move one.
4. OPF cancel printed twice — now NewSilentError, matching the convention I was already following three lines above for context.Canceled.
The refactor in #1 touches the pre-push hot path, so I checked the extraction is byte-faithful and leaned on the existing suites: lint 0 issues, GOOS=windows go vet clean, both packages pass in full, 692 integration tests, canary 56/56 and 4/4.
Pushed as 1cd760fc. I didn't reply to or resolve the PR comments — you haven't asked me to, and I'd rather not post on your behalf.
Not yet confirmed: the re-review. The trail is back to … findings pending — waiting for the agent review of the current head with 13 checks running, so "No findings found" right now means not-yet-reviewed, not clean. Want me to watch for it to land?