Fix OPF Byte Cap Scope Per Checkpoint Ref

can you test this https://entire.io/gh/entireio/cli/trails/1380/fix-strategy-dedup-opf-byte-cap-scope-it-per-checkpoint-ref in new binary build to test hte fix
Base directory for this skill: /Users/peytonmontei/dev/entire/devenv/entireio/cli/.worktrees/opf-investigation/.claude/skills/test-repo
Test Repository Skill
This skill validates the CLI's session management and checkpoint creation by running an end-to-end test against a fresh temporary repository.
When to Use
- User asks to "test against a test repo"
- User wants to validate strategy changes (manual-commit)
- User asks to verify session hooks, commits, or checkpoint creation
- After making changes to strategy code
Testing Approaches
Automated Testing (recommended for validation):
Run the comprehensive integration test suite. Best for verifying correctness after code changes.
Manual Testing (this skill): Use the test harness for:
- Debugging specific strategy behaviors
- Interactive exploration of the checkpoint workflow
- Manual verification of edge cases
- Understanding how the system works step-by-step
Test Procedure
Setup
Step 1: Build the CLI
Step 2: Approve the test harness (one-time)
Add this pattern to your Claude Code approved commands, or approve it once when prompted:
Optional: Set strategy (defaults to manual-commit):
Test Steps
Execute these steps in order:
1. Setup Test Environment
2. Simulate Session
3. Verify Results
Expected results:
| Check | Result |
|---|---|
| Active branch | Optional Entire-Checkpoint: trailer |
| Session state | ✓ Exists |
| Shadow branch | ✓ entire/{hash} |
| Metadata branch | ✓ entire/checkpoints/v1 |
| Pending checkpoints | ✓ At least 1 |
4. Check Listing After Further Changes
Expected Behavior:
- The checkpoint from step 2 is still listed; the new uncommitted changes do not
disturb it. There is no restore step: the CLI has no
rewindcommand, and nothing writes a checkpoint back over the worktree.
5. Cleanup
Quick Commands
Show environment info:
Run full test in one go:
Expected Results by Strategy
Manual-Commit Strategy (default)
- Active branch commits: NO modifications (no commits created by Entire)
- Shadow branches:
entire/<commit-hash[:7]>created for checkpoints - Metadata: stored on both shadow branches and
entire/checkpoints/v1branch (condensed on user commits) - AllowsMainBranch: true (safe on main/master)
Additional Testing (Optional)
Test Subagent Checkpoints
For testing task checkpoints (subagent execution):
Test User Commits (Condensation)
For manual-commit, test log condensation:
Available Claude Code Hooks
All hooks use the command: entire hooks claude-code <hook-name>
user-prompt-submit- Called when user submits a prompt (before session starts)session-start- Called when session startsstop- Called when session stops (creates checkpoint)pre-task- Called before Task tool executionpost-task- Called after Task tool executionpost-todo- Called after TodoWrite tool execution (for incremental checkpoints)
Report Format
After running the test, report:
I tested PR #2533 (trail 1380, head 373171411a) end to end with the real opf binary. The per-ref cap, pinned delivery and queue cleanup all work, but I found a blocker: on a normal first push, the background worker is killed before it redacts anything.
Setup: isolated repo on the git-refs backend with a local bare remote, OPF on for private_person, and two checkpoints mentioning "Alice Johnson". One is small (~3.3 KB) and one is big (~14.9 KB). Most pushes used ENTIRE_OPF_BATCH_LIMIT=4000, so the small ref fits and the big one doesn't.
Blocker: the detached worker dies on its first progress line
- What happens: with the PR binary, the first push always spawns
__opf_flush(it's logged), but nothing gets redacted and nothing is logged, even at debug level. This happened 0/4 times correctly on fresh repos. - Cause:
execx.SpawnDetachedsetscmd.Stdout/Stderr = io.Discard, so Go gives the child a pipe that the parent reads. Aftergit pushreturns, the parent exits and nobody reads the pipe. The worker's first stderr write (→ OpenAI Privacy Filter: scanning checkpoints…) then raises SIGPIPE, which kills it. - Why it sometimes works: it only survives when it fails before printing anything. My second push was like that, because the only ref left was over the cap.
- Fix, confirmed: setting the child's stdout/stderr to nil (Go then sends them to
/dev/null) gave 3/3 good runs. - Why tests missed it:
SpawnDetacheddoes nothing undertesting.Testing(). - Scope: the helper is older (July 2026, used for the session sweep, telemetry and trail refresh), but this PR's worker is the first caller that writes to stderr.
Everything else passed (patched binary)
| Check | Result |
|---|---|
git push time | 0–1s every push; OPF never ran inline |
| Per-ref cap | Small ref rewritten, big ref left untouched and still queued. Under the old shared cap, both would have been blocked |
| Pinned delivery | The next push delivered only the small ref, at the exact trailered hash (5dd1c7e); the big ref was held back |
| Remote content | 0 "Alice Johnson", 4 [REDACTED_PERSON] in each delivered ref |
| Cap lifted | Worker rewrote the big ref and the next push delivered it; queue file empty afterwards |
entire status --json | checkpoint_opf_pending went 1 → gone (the field is omitted at 0) |
Minor
- The raw-cap error still says "would buffer … across all unpushed commits", which is wrong for git-refs now that it counts per ref.
- The same
git-refs pass failedWARN is logged twice per worker run.
I haven't changed the PR. The fix is a two-line change in cmd/entire/cli/execx/spawn_detached.go, plus the error wording. Want me to push that to peyton/opf-batch-cap-fix?
The bug from T-139 is still fixed, including under PR #2533's new background worker. The thread itself has two loose ends with the reporter (Chetan).
T-139's original bug: with OPF enabled and either categories: {} or {"private_person": false}, a push on the PR binary withholds the checkpoint ref with the "no detection category is enabled" error. Nothing got the OPF trailer and nothing reached the remote.
Loose ends in the thread:
- His Aug 23 Copilot report. He reported that summarization runs Copilot with tools enabled on untrusted transcript content. Stefan agreed. It was fixed in
dbe0d142c1on Sep 8, which removes--allow-all-toolsand denies shell, write and url tools; it shipped in v0.11.0. Paulo's Sep 9 email still says "we'll update you once the review is complete", and no follow-up was sent. The thread is TODO, with his "thanks" as the latest message.- There's a residual risk the commit accepts on purpose: read-only tools are still available, so injected transcript text could steer file reads into a summary. The Claude generator has the same gap. Closing it would mean
--available-toolswith an empty set, which isn't done yet. Worth deciding before replying whether to call this fully fixed or partially mitigated.
- There's a residual risk the commit accepts on purpose: read-only tools are still available, so injected transcript text could steer file reads into a summary. The Claude generator has the same gap. Closing it would mean
- Repair of already-tagged commits. On Aug 13 you told him these "won't be retroactively repaired, we're tracking that separately." I haven't found anything showing that was done.
Want me to draft a reply to him covering the Copilot fix? I won't send anything from Plain myself.
PR #2533 fix: mise run check is running in the background. Once it passes, I'll commit the stdio fix, its regression test and the error wording, then push to peyton/opf-batch-cap-fix.
Yes. I built two binaries and tested both with the real opf model against a throwaway repo and a local bare remote:
| Binary | Built from | Result |
|---|---|---|
entire-fix | PR head 373171411a, unmodified | Found the bug: 0/4 first pushes got redacted, because the background worker died silently |
entire-final | PR head plus my uncommitted patch | 3/3 first pushes redacted; the full flow passed (per-ref cap, delivery, [REDACTED_PERSON] on the remote, big ref once the cap was lifted); both no-category configs from T-139 still withhold the ref |
On size: the PR is large at +3,211/−456 across 35 files. My patch adds +22/−13 in 2 files (the stdio fix and the error wording), plus a new ~25-line test file.
mise run check is still running: lint is clean and the unit tests so far pass. I'll commit and push to peyton/opf-batch-cap-fix once it finishes green.
I pushed the fix to PR #2533 as c8d3ca72c1. mise run check passed before the commit (lint clean; unit, integration and canary tests green), and so did mise run lint after it. CI is running on the new head, with 5 checks still pending.
The commit adds +47/−13 in 3 files:
- The fix:
execx.SpawnDetachednow hands the child the null device instead ofio.Discard. Previously the background worker died on its first "scanning checkpoints…" line aftergit pushreturned. The command setup is split into a smalldetachedCommandhelper so a test can check it. - Regression test: asserts the child's stdout/stderr aren't a pipe. A unit test can't reproduce the real crash, because the helper does nothing under
go test, so the evidence it works is the rebuilt-binary runs (3/3 good, versus 0/4 before). - Error wording: the raw-byte cap error now says "in one unit of work" instead of "across all unpushed commits".
I didn't merge. Still open from before: whether to draft the T-139 reply to Chetan about the Copilot fix.
Almost. It's ready once CI finishes, and the PR description should mention what I found and tested.
- CI: 5 checks pass, 10 are still running on
c8d3ca72c1, and none have failed. - Trail findings: none open. I dismissed the only one because it was a false positive (reason recorded on the finding).
- Local verification:
mise run checkpassed. The realopfend-to-end test passed on a binary rebuilt with the fix, and the no-category fail-closed behavior from T-139 still holds. - PR state: not a draft, no merge conflicts, waiting on a required review.
- Description is out of date: it doesn't mention the detached-worker SIGPIPE fix or the rebuilt-binary testing. Reviewers should know that the unmodified PR head silently failed to redact on first push.
- Size: at +3,258/−469 across 37 files, it's a big review. The description should say where to start, probably
manual_commit_opf_flush.goandpush_common.go.
Want me to update the description with the new fix and the validation results?
I've updated the description on PR #2533: https://github.com/entireio/cli/pull/2533
- Result now covers the fix for the background worker crash (
c8d3ca72c1), including why unit tests couldn't catch it. - Where to start reviewing is new and lists four starting points: the worker, delivery and recovery, per-ref collection, and queue cleanup.
- Validation now includes the real-binary
opfruns: 0/4 before the fix, 3/3 after, the full delivery and redaction flow, and the T-139 no-category check.
The old description is saved at scratchpad/pr-body-old.md if you want to compare. Once CI finishes green, the PR is ready to go to a human reviewer.
can you give me short/simple slack sentence, like "fixes cap and jam from opf" or something. Also, no description in trail
Trail 1380 now has a description: the PR description without the trail link and the Claude Code footer.
For Slack:
OPF fix is up for review: no more redaction stalling your
git push, the size cap now applies per checkpoint instead of one big session blocking everything, and the background redaction actually runs now. https://github.com/entireio/cli/pull/2533
Shorter:
Fixes OPF's cap jam and push stall, and makes background redaction actually run. https://github.com/entireio/cli/pull/2533
All three findings are real. Two are straightforward fixes; the third (#3) is a product call.
1. High: one huge ref starves every ref queued behind it. This is valid, but the threshold is higher than the finding says.
- The 3h deadline is computed with a 2× margin, but the actual run takes 1×. So a ref only times out above roughly 9.5 MB of prose, not 4.6 MB.
- Your earlier measurement had a median of about 9.5 MB raw per ref's full ancestry, so real refs can plausibly reach that.
- When it happens, the timeout trips the process-wide breaker and the loop stops. The same ref is first in the queue every run, so each run spends 3h on it and nothing behind it ever gets redacted or delivered.
- Even below the timeout, a big first ref delays all its siblings by hours in every run.
- Proposed fix: before calling the model, skip any ref the deadline can't finish (prose bytes × 1.14 s/KB > 3h) as a per-ref cap error, the same way oversized refs are skipped now. Also process queued refs smallest first. About +60 lines including tests.
2. Medium: two workers can run on the same backlog at once. Valid.
- The only guard is the 5-minute spawn throttle, and a worker can legitimately run for hours.
- A second push after 5 minutes starts a second worker that repeats the same expensive model call. Whichever finishes later loses the ref update and its work is thrown away.
- Proposed fix: the worker takes a non-blocking lock for its whole lifetime using the existing
internal/flock, and exits immediately if another worker holds it. The OS releases the lock if the worker dies. About +50 lines including tests.
3. Medium: git-branch git push can now block for hours. The facts are right: on main the cap was 2 MiB and the timeout a fixed 30s, and git-branch still rewrites inline during the push.
- But reverting brings back the original bug on git-branch: every push carrying a real session would abort at the cap, so OPF would be unusable there.
- The docs already say that on git-branch a large session means a long push.
- So the choice is between "fails fast, OPF never works" (main) and "works, but slowly" (this PR). A middle option is a separate, smaller budget for the inline path, for example refusing chains above roughly 30 minutes of estimated time with an error pointing at
ENTIRE_OPF_BATCH_LIMIT.
My recommendation: fix #1 and #2 in this PR (about +110 lines in total). For #3, either keep the current behavior and dismiss it with that rationale, or add the inline budget, which is about +40 lines more.
Should I fix #1 and #2? And for #3, keep the current behavior or add the inline budget?
Yes. An inline budget is the cap again, and it would end up lower than main's.
- At the measured 1.14 s/KB, a 30-minute budget works out to about 1.5 MB of prose. Main's cap was 2 MiB.
- A single real session is about 3.3 MB of prose, so every real git-branch push would be refused. That's the original bug again.
- Any time limit short enough to keep
git pushfast will refuse real sessions, because the model takes minutes to hours on them.
The only real fix for git-branch is getting the scan off the push path. That's hard on git-branch because the checkpoint branch travels with the user's push, so deferring the scan means holding back that branch.
My recommendation: dismiss #3 with this rationale. As a possible follow-up, git-branch could print an estimated wait before the inline scan starts, so a long push isn't a surprise.
Want me to dismiss #3 and fix #1 and #2?
also wouldnt want someone sitting for that long on the wait, but i dont want to add a manual cap. how can we make sure there isn't a hang instead
You can do it without a cap: move git-branch to the same deferred model git-refs now uses, so the push never waits on the model. I said earlier that was hard on git-branch, but the code shows it isn't blocked. Entire pushes entire/checkpoints/v1 itself from the pre-push hook (pushRefIfNeeded, manual_commit_push.go:189), separately from the user's own push. So it can hold that branch back and let the user's push go through.
What it would do on git-branch:
- Pre-push: if v1 has commits without the OPF trailer, don't push v1 this time. Start the detached worker and let the user's push complete immediately.
- The worker rewrites v1 in the background, however long that takes.
- The next push delivers v1 once its tip carries the trailer.
- If the user names v1 in their own
git pushwhile it's unscanned, pre-push still refuses it. That's the one case that has to stay blocking.
The hard part: v1 is a single branch, and every commit writes new checkpoint data to it. During an hour-long scan, an active user will almost always move v1 before the worker finishes. With today's rule (abort if the branch moved), the worker would rarely land. It needs to install the rewrite and then replay any newer commits on top, leaving them unscanned for the next pass. git-refs doesn't have this problem because each checkpoint has its own ref.
Size: roughly +400–600 lines including tests. That belongs in its own follow-up PR, not in #2533, which is already +3.2k lines.
Plan:
- In #2533: fix #1 and #2, and dismiss #3, linking a follow-up issue for moving git-branch off the push path.
- Follow-up PR: git-branch deferred rewrite with the replay-on-conflict install.
Should I go ahead with that, including filing the follow-up issue?
what other impacts would moving git-branch to the same deferred model git-refs now uses - besides the opf?
Beyond OPF itself, the main cost is that checkpoint data reaches the remote one push late. That's already true on git-refs with OPF on.
Checkpoints arrive one push late. The worker only rewrites and never pushes; a later git push delivers the result.
- In the common "push once, open a PR" flow, the PR and trail on entire.io show the commit's
Entire-Checkpointtrailer, but the transcript isn't there until you push again. - If you never push again, it never arrives.
- Anything that reads checkpoints from the remote sees stale data in the meantime: reviews and trail monitors, teammates running
entire resumeor explain on another machine, and the web UI.
The worker would need network access in the background. The git-branch rewrite fetches the remote's v1 tip live so it has a base to rebuild on; git-refs doesn't need this.
- In a detached process that means credential helpers, SSH agent and VPN all have to work with no terminal. A prompt it can't answer would hang the worker invisibly.
- It needs non-interactive git and its own timeout. Otherwise we've moved the hang from the push to the background.
Ephemeral environments never deliver. In CI or a throwaway container, the process usually exits right after the push and takes the worker with it. Checkpoints from those runs stay local and are lost. git-refs has the same problem today.
Push capture and first-push handling get more complicated. The pre-push flow only records the checkpoint remote once v1 actually lands there. Deferring v1 postpones that. It also stacks a second deferral on top of the existing "don't push v1 to an empty remote" rule, so there are more states to get right.
Commits keep competing with the worker. Every commit writes new checkpoint data to v1. That's the replay-on-conflict work I mentioned. Without it, active users would almost never get a rewrite to land.
Visibility. entire status currently reports pending OPF work only for git-refs, so it would need to cover v1 as well. Otherwise git-branch users have no sign that their data is being held.
What doesn't change: local reads (explain, resume on the same machine), commit latency (OPF never runs at commit time), and v1 history rewriting. The push already rewrites v1 today; this only changes when it happens.
Worth considering instead: with OPF on, entire enable could recommend git-refs, or the checkpoint migration command could offer to move the repo over, rather than building a second deferred pipeline for the legacy backend. git-branch users would then either choose a long push or switch, and the PR only needs to document that. That's cheaper, and it avoids the network-in-background and conflict-replay problems entirely.
can we fix this the main cost is that checkpoint data reaches the remote one push late. That's already true on git-refs with OPF on. - like there shoyuldnt be a delayed push with the filter on
Yes. The worker can push the result itself as soon as the redaction finishes. The only remaining wait is the scan time, and nothing can be shipped before it's redacted anyway.
Why it doesn't push today: the worker's design comment (manual_commit_opf_flush.go:10) says "a background process must not be able to ship anything", so only a user-initiated push can send data off the machine. That rule is stricter than it needs to be. The user's push already decided to run OPF, and which remote, for exactly those checkpoints. The worker can carry out that decision without making a new one.
How it would work:
- Pre-push passes the push target (the remote the user pushed to) to the worker when it spawns it. That also fixes the stated reason the worker can't handle git-branch: it has no remote to compute the unpushed range against.
- Once each ref is rewritten, the worker delivers it right away, using the same delivery pinned to the exact trailered hash and the same exact-generation queue cleanup as a normal push.
- It only delivers what that push authorized: refs that were queued at spawn time, to that remote, and only if they carry the trailer. Checkpoints created after the push wait for the next push, as they do now.
- Network access in the background has to be non-interactive: no credential prompts (
GIT_TERMINAL_PROMPT=0, SSHBatchMode) and a timeout. If auth or the network fails, the ref stays queued and the next push delivers it, which is today's behavior as the fallback. - The worker also records the checkpoint remote when its delivery lands, the same way a normal push does.
What still doesn't fix itself: CI and throwaway containers that kill the worker when the job ends. There, OPF could run inline when there's no terminal, so the job waits. That's a separate call.
Size: roughly +250–350 lines including tests, for git-refs. The git-branch version is the follow-up discussed above, and it would reuse this.
It changes a safety rule this PR states, so I'd rather not add it to #2533 unannounced. Options:
- (a) Do it in #2533 and update the description, since it's the same feature and reviewers would see the whole design at once.
- (b) Do it as a follow-up PR on top of #2533. Smaller reviews, but #2533 ships with the one-push-late behavior in the meantime.
Which do you want? Either way I'd do findings #1 and #2 first.
okay do findings 1 and 2, and do we do this here or somerhwere ele? like new pr or on this one
Findings #1 and #2 are fixed, pushed to PR #2533 (head 9c609c554b), and resolved on the trail. Together they add +225/−26 across 5 files, and mise run check passes.
- #1, one huge ref starving the queue (
6b0f326da9): OPF input is now sent to the model in calls of at most 1 MiB each, and each call gets its own time limit. A very large ref now just takes longer instead of hitting the 3h limit and blocking the refs behind it. There's no size cap, and a typical checkpoint still goes out in one call. - #2, two workers at once (
9c609c554b): the worker holds a lock for its whole run, and a second one exits immediately. I checked this with a real binary: with one worker mid-scan, a second exited in under a second, and the first finished its 34s scan and redacted the ref. - Docs: the OPF docs now describe the 1 MiB calls and the one-worker-per-repo rule, and no longer mention the 3h clamp as a limit.
Still open:
- Finding #3 (git-branch push can block for hours). I left it open because you hadn't decided. The options are to dismiss it with a pointer to a follow-up, or to handle it in the "no delayed push" PR.
- The "no delayed push" work (the worker delivers its result as soon as redaction finishes) goes in a new PR stacked on #2533, as agreed. I haven't started it yet.
- Your network:
git pushthroughgit-remote-entirefailed with "no such host" on this phone hotspot. Go's built-in DNS resolver can't use the hotspot's IPv6 DNS server, while the system resolver can. It worked withGODEBUG=netdns=cgo. Other users on similar networks would hit the same thing, so it's worth filing separately.
It'll definitely be an issue: git-branch is the default backend, so anyone who turns on OPF without changing backends hits #3. My recommendation is to give git-branch the same model as git-refs. The push never waits, a background worker does the scan and pushes the result itself, and there's no cap.
How it works on git-branch:
- Pre-push: if v1 has commits that haven't been scanned, Entire skips pushing v1 this time. It starts the worker with the remote you pushed to, and your
git pushcompletes immediately. - Worker:
- Fetches the remote's v1 tip with no prompts (
GIT_TERMINAL_PROMPT=0, SSHBatchMode) and its own timeout. - Rewrites the unscanned commits and installs the result. Any commits that landed on v1 during the scan are replayed on top.
- Pushes the exact scanned hash to
refs/heads/entire/checkpoints/v1. Newer commits on top aren't scanned yet, so they wait for the next pass. - Pinned-hash delivery is what makes git-branch workable: the worker can publish the scanned part of v1 while newer commits keep landing on top.
- Fetches the remote's v1 tip with no prompts (
- If you name v1 in your own
git pushwhile it's unscanned, pre-push refuses that push. This is the one case that has to stay blocking. - Fallback: if the worker can't reach the remote (auth or network), v1 stays local and the next push delivers it.
Where to do it: the stacked PR I planned for worker self-delivery. It becomes "OPF never blocks a push, and the worker delivers" for both backends, since they share the worker, the delivery pinned to the exact hash, and the non-interactive network setup. Estimate: roughly +700–900 lines including tests.
Merge order matters: as it stands, #2533 makes git-branch pushes wait for the full scan, which can be hours on large sessions. Main fails fast instead. So #2533 shouldn't ship on its own. Either merge the two together, or hold #2533 until the stacked PR is ready. Finding #3 can stay open until then, marked as fixed by the follow-up.
Alternative I'd avoid for now: making git-refs the default whenever OPF is enabled. It's smaller, but it changes storage for existing repos and needs migration. That's a bigger product decision than this bug.
Should I start the stacked PR on this design?
so basically for git refs, we delay pushing the checkpoint? im worried aout this getting messy, but how do we handle it now anyway bceause we can learn from the git ref storage implementation now, it should theoretically be more efficient than v1 branch
Yes. On git-refs with OPF on, a checkpoint ref is held back until the worker has scanned it. Today it then goes out on your next push; with the stacked self-delivery PR, it would go out as soon as the scan finishes. So the delay would be the scan time, not "wait for the next push".
Why git-refs handles this cleanly:
- Each checkpoint is its own ref. The unit of work is one checkpoint. A slow or broken one only holds itself back, and the others ship.
- Nothing is shared between checkpoints. A new commit creates a new ref instead of moving a shared tip, so a long scan rarely conflicts with ongoing writes. The per-ref compare-and-swap covers the rare case.
- "Already scanned" is decided locally. The rewrite stops at the first commit carrying the OPF trailer, so it doesn't need to fetch the remote first.
- Delivery is simple. The queue holds
(ref, hash)pairs, delivery pushes the exact scanned hash, and cleanup removes only that pair.
Why porting this to v1 gets messy: v1 has none of those properties. It's one shared chain that every commit moves. The rewrite has to fetch the remote tip to know what's unpushed, a long scan will almost always conflict with new commits, and one problem holds back everything. Every piece of the design I outlined (background fetch, replay on conflict, pushing a scanned hash that isn't the tip) exists only to work around that. You're right to be wary of it.
Better answer to #3: don't rebuild this on v1, use git-refs for OPF.
- Enabling OPF on a git-branch repo (at
entire enableor when settings change) offers to migrate to git-refs. The migration command (entire doctor migrate-checkpoints) already exists. The docs note that migrated refs each get scanned on the first push, one ref at a time, so the cost is model time, not a rejection. - git-branch repos that don't migrate keep today's inline scan in
git push. Before the scan starts, Entire prints an estimated wait and suggests migrating, so a long push is never a surprise. There's no cap and no new v1 machinery. - Finding #3 then becomes: "git-branch stays inline by design; OPF steers users to git-refs." That's a documented tradeoff rather than an open bug.
Size: the migration prompt plus the estimate message is roughly +150–250 lines, against +700–900 for porting the deferred model to v1.
One thing to check before committing to this: whether git-refs is ready to be the recommended backend in general. Right now it's opt-in and git-branch is the default. If it is, this is the clean path. Want me to go this way?
it is the recommended backend in general, but should also be able to handle people on v1 branch, can make suggestion to move if someone is on v1 branch and has opf enabled, but again it should handle both storage methods
Here's a design that handles v1 properly without the mess. It takes the lesson from git-refs one step further: the background worker never touches the v1 branch. It only scans content and caches the redacted results. The v1 rewrite stays inline in pre-push, as it is today, but becomes fast because the model has already run.
Why this avoids the mess: all of v1's pain comes from rewriting a shared, moving branch in the background. The slow part, though, is only the model call, and that works on blob contents, which never change. So only the model call moves to the background, and the branch rewrite stays inline where it already works.
Flow on v1:
- Worker:
- Walks local v1 back to the first commit with the OPF trailer, the same watermark git-refs uses. That's all local, so no fetch or remote target is needed.
- Runs OPF on the blobs it finds and stores the redacted output in a cache keyed by blob hash plus category set.
- Writes no refs, so it can't conflict with new commits.
- Pre-push, everything cached: the rewrite reads from the cache, needs no model call, and finishes in seconds. v1 is pushed with this same push, exactly as today, including the existing compare-and-swap and the "moved during rewrite" handling.
- Pre-push, something not yet cached: v1 isn't pushed this time, and your
git pushcompletes immediately. Entire starts or keeps the worker going and prints something like "checkpoints held while OPF scans (~N min); your next push sends them." - Suggest migrating: OPF on git-branch prints a one-time hint recommending git-refs, pointing at
entire doctor migrate-checkpoints.
What you get:
- No cap, no push that hangs, and no background history rewriting.
- The worker's scan-and-cache step and the 1 MiB chunking are shared with git-refs.
- There's a precedent: a redaction cache already exists in the git common directory (
checkpoint/redact_cache.go, for incremental regex redaction). This would be a sibling cache for OPF results.
The tradeoff: on v1, checkpoints arrive one push late, the same as git-refs today. The self-delivery follow-up could later close that for both backends. For v1 it would be a "push only when fully cached" step, which is much simpler than a background rewrite.
Cache details to get right:
- The cache key has to include the category set and the
opfversion, so changing either forces a rescan. - The cache holds only redacted output, never cleartext beyond what's already in
.git. - Its size needs bounding, and entries pruned when their commit ships.
Size: roughly +700–800 lines including tests, as a separate PR on top of #2533. It's about the same size as the background-rewrite design, but with far less concurrency risk, because nothing in the background writes refs or talks to the network.
It could also later replace git-refs' background rewrite, so both backends share one model. I wouldn't do that in this PR.
Should I start on this?
Yes. If the cache design is where both backends end up, #2533 is carrying machinery that the next PR would delete. I'd cut it down before review rather than ship code we plan to rewrite.
What changes in the cache design: the background worker only scans and never writes refs, and the rewrite is inline and fast in pre-push for both backends. So the parts of #2533 that exist only to make background ref rewriting safe go away:
- the worker rewriting refs (
RunOPFFlushrewrite passes and the "retry when the ref changed" loop) - the worker safely requeueing a ref it has rewritten
- recovery races between the worker and delivery
- the worker lock from #2 (it moves to the next PR in simpler form, since a scan-only worker is safe to overlap)
What stays in #2533, useful under either design:
- delivery pinned to the exact scanned hash, and queue cleanup that removes only that generation (concurrent checkpoint writes still happen)
- the byte cap scoped per checkpoint ref
- model calls split into ≤1 MiB chunks, and the size-scaled timeout
- the
SpawnDetachednull-device fix (a general bug, and tiny) - the corrected error wording and docs
Resulting stack:
- #2533, reduced: OPF correctness and capacity fixes, with no worker. git-refs still scans inline, so a large push still waits. That's the same as main today, so nothing gets worse. Roughly half its current size.
- PR 2, OPF scan-only worker plus result cache, both backends: the push never waits on the model. If anything isn't cached yet, the checkpoints are held for that push and go out on the next one. This fixes #3 for v1 and also covers git-refs. It includes the one-time hint for v1 users to migrate.
Costs:
- Some of #2533's already-tested worker code is thrown away.
- #2533 ships without the non-blocking push, so the two should land close together.
- Reshaping the branch needs a force-push, which I'll prep and hand you as one
!command. - The real-binary validation in the description has to be redone for the reduced version.
The alternative is to ship #2533 as it is and migrate git-refs to the cache model later. That's less work right now, but it means reviewing a background rewriter we intend to replace, and supporting two models for a while. I'd do the reduction.
Should I go ahead and reshape #2533 that way?
i think we should rework 2533, reove anything we'd delete and we should do the work now for storage methods now with the filter instead of waiting. Is it best to create anotehr pr for this or add to 2533 after rework
I've rebuilt #2533 locally at faca8396c9, but it isn't pushed yet: that takes a force-push, which you need to run (command below).
It's two commits on the latest main, +1,609/−319 across 17 files, down from +3,479/−490:
ae9b18fdfb: the prose-leaf byte count counts identical leaves once.faca8396c9: delivery is pinned to the exact scanned hash, queue cleanup removes only that generation, the caps apply per checkpoint ref, and blob sizes are checked before reading.
What left #2533: the background worker and everything around it (retries, requeueing, the worker lock), the higher caps, the size-scaled timeout, chunked model calls, the SpawnDetached fix, the pending count in entire status, and the two investigation docs. They're preserved on the old head (9c609c554b) for PR 2. Default caps are back to main's (2 MiB prose, 200 MiB raw), and OPF runs inline on both backends as it does on main.
Main moved underneath us: #2522 added a bounded flush that moves failed refs to the back of the queue. I merged it with this PR's (ref, hash) queue entries so rotation moves exact generations, and added a test that rotating an old generation leaves a newer one of the same ref in place.
Verification: mise run check passed, and so did mise run lint on the final tree. The real binary with the real opf also passed:
- git-refs: the ref under the cap was scanned and shipped in the same push, and the oversized one was withheld with a per-ref error. The next push shipped it. Both refs are redacted on the remote.
- No categories enabled: the ref is withheld and never trailered.
- git-branch: v1 is pushed with the OPF trailer and redacted content.
Force-push: this overwrites the PR head, and the lease makes it refuse if the branch moved since 9c609c554b. The GODEBUG setting works around the hotspot DNS problem from earlier.
After you push, I'll:
- Replace the PR and trail descriptions with the new version (saved at
scratchpad/pr-body-v2.md). It explains the cut-down scope, says the follow-up must merge together with this PR, and lists where to start reviewing. - Resolve finding #3 on the trail: with main's cap and timeout back, git-branch can no longer block for hours, and the follow-up covers the rest.
- Start PR 2 from
9c609c554b's parts: the scan-only worker plus the OPF result cache for git-refs and v1, the capacity changes, the migration hint, and a real-binary test on both backends.
The force-push landed: #2533's head is now faca8396c9.
- PR description: replaced with the new one (cut-down scope, merge together with the stacked PR, where to start reviewing, real-binary results).
- Finding #3: resolved. With main's cap and timeout back, git-branch can't block for hours; the stacked PR covers the rest. The trail has no open findings.
- CI: 2 checks pass and 13 are still running. None have failed.
Not done: the trail description. The safety check blocked my update (it flagged --overwrite), so it still shows the old version. To update it yourself:
If it complains that the description changed since it was read, add --overwrite.
Next is PR 2, stacked on #2533: the scan-only worker plus the OPF result cache for git-refs and v1, the capacity changes, and the migration hint. Should I start it now, or wait for #2533's CI?
The file exists now (38 lines, the PR description without the trail link and the Claude Code footer). Run the same command again and it will fill in the trail description:
Nothing else was affected: the PR description is already updated and the finding is resolved.
The trail description is filled in, and every test check on #2533 passes (16 pass, 1 skipped). The one failure is Entire Gates, which isn't a test.
- Why it fails: the trail reports merge as blocked with
merge_gates_required, and approval as blocked withself_approval_not_allowed. So it most likely needs approval from someone other than you. I can't see the exact gate rule from here, so the high risk score could also count. - Monitor scores:
| Monitor | Score | Change |
|---|---|---|
| Confidence | 82 | −3 |
| Drift | 5 | −11 |
| Risk | 72 | −8 |
| Security | 18 | −4 |
- Why risk is high: the rationale is that the PR "modifies the privacy-critical egress path". That's inherent to what the PR changes, not something a code fix would lower.
- Findings: none open.
The next step is a human review. You can post the Slack message from earlier to find a reviewer; the PR description now says to merge it together with the stacked follow-up.
Should I start PR 2 now?
Found the root cause, and it's a real privacy bug that also exists on main. It's the same class as T-139: a commit is stamped Entire-OPF-Applied: true while PII still ships in cleartext.
What happens:
BatchBytesWithPrivacyFilterjoins all unique leaves into oneopfinput, separated by\x1e.- The model sometimes returns a span that crosses a separator. Here it returned
private_person "Code<SEP>Alice Johnson": the previous leaf ended in "…Claude Code", and the model read "Code Alice Johnson" as one name. partitionIndex(redact/opf.go:583) treats any span crossing a boundary as invalid and drops the whole span. Neither leaf gets redacted.- The rewrite then stamps the commit as OPF-applied, and the name ships.
How I confirmed it:
- Calling
opfdirectly on the note alone redacts the name. - Batching the commit's actual blobs misses it 3 times out of 3, and debug logging in a scratch copy of the source shows the dropped span.
- Main's binary does the same thing, so neither #2533 nor divergence recovery causes it. My test file just happened to trigger it.
Real-world exposure: any leaf that starts with a name, right after a leaf that ends in a name-like word. Transcripts are full of those, like "Claude Code" followed by a prompt starting "Alice…". It's content-dependent, not an edge case.
Fix: split a crossing span at each separator and give each piece to the input it overlaps, instead of dropping it. That fails safe, because redacting "Code" too costs nothing. It's roughly +40 lines plus tests: a unit test for the split logic and a regression test with a fake runtime that returns a crossing span. I can also test it against this exact real batch.
Where it goes: its own small PR off main, not #2533 or PR 2. It's a fail-open security fix that should ship on its own schedule and be easy to review. Given the T-139 precedent, you may also want to flag it through the security process, since trailered commits already pushed with OPF on may contain unredacted names.
Other results from the three targeted tests on #2533:
- Upgrade from main's queue format: passed.
- Rejection plus moving failed refs to the back of the queue: passed.
- Divergence recovery: worked. The only problem was this pre-existing redaction miss.
Should I start the separate fix PR now, before PR 2?
yes lets start this fix, adn write plan, but dont create an issue since we're addressing now. and did you test 2533?
I've committed the fix as 758fac0e73 and pushed it to peyton/opf-cross-separator-spans, but haven't opened the PR yet. It's +85/−45 across 2 files.
Verified:
- New tests: a table test for the split (7 shapes) and a regression that feeds
RedactBatcha span crossing a separator. Both failed before the fix and pass after. - On the exact batch that leaked, the note is redacted 3 out of 3 (main misses it every time).
- A real-binary push lands trailered with no cleartext name anywhere.
mise run checkandmise run lintpass.
Before I open the PR, one decision is yours: this is a fail-open privacy bug in a public repo. A regular PR discloses it before a release ships the fix. Note that the pushed branch and its commit message are already visible.
Options:
- Regular PR with a neutral description. Recommended. OPF is opt-in, and T-139 was handled the same way (fixed publicly, credited in the changelog).
- Private advisory fix (GHSA). Hold the PR and route it through the security process, as the ACE issue (GHSA-hqjp-v5g5-vxvf) was.
Tell me which, and I'll open it (without merging). The PR will list the "Claude [REDACTED_PERSON]" over-redaction and the follow-up to keep structural metadata fields out of OPF.
I opened the fix as PR #2624: https://github.com/entireio/cli/pull/2624. Its 15 CI checks are still running and nothing has been merged. The description covers:
- the fail-open: names shipping under the OPF trailer
- the fix
- the "Claude [REDACTED_PERSON]" over-redaction tradeoff
- the dead end with
\n\nas the separator - the follow-up to keep structural fields like agent, model and branch out of OPF
- the validation, including 3 of 3 misses on main and 3 of 3 redactions on this branch against the real model
Does it cover both storage options? Yes. The fix is in the redact layer that both backends call, git-refs through manual_commit_opf_refs.go and git-branch through manual_commit_opf_rewrite.go. So both are fixed, along with the non-batched redaction paths.
What's still open:
- #2533: tests pass. It's waiting on a human review for the Entire Gates check, and should merge together with PR 2.
- #2624: CI is running. It's independent of #2533 and can merge on its own, before either.
- PR 2: not started. It's the scan-only worker plus the OPF result cache, so OPF never blocks a push on either backend. It also brings back the higher caps and chunked model calls, and adds the migration hint for v1 users.
Should I start PR 2?
Keep them separate:
- #2624 is a small, self-contained security fix (+85/−45 in 2 files). It fixes a leak that exists on main today, so it should merge as soon as it's reviewed. Folding PR 2 in would hold that fix back behind a large new subsystem and a much longer review.
- PR 2 depends on a different base. It builds on #2533's pinned delivery and exact-generation queue, not on #2624. It stacks on #2533, and the two merge together.
- They barely overlap. #2624 changes span mapping in
redact/opf.go. PR 2's chunking changesredact/batch.goand the timeout code inredact/opf.go. Whichever merges second needs a small rebase, nothing more.
Order: #2624 whenever it's approved, then #2533 and PR 2 together. Should I start PR 2 on top of #2533?
Neither failure is caused by the fix.
test-codex-hooks: both failing tests (TestCodexAppServerHooksList_BareWorktreeUsesLayoutRootand…_LinkedWorktreeUsesPrimaryCheckout) fail before running any Entire code. The Codex CLI on the CI runner is broken:Error: Missing optional dependency @openai/codex-linux-x64. Reinstall Codex: npm install -g @openai/codex@latest. #2624 only changes span mapping inredact/opf.go, and this same job passed on #2533's run a few hours ago. So the likely cause is a new@openai/codexnpm release missing its Linux binary. I couldn't confirm this against other branches because I hit GitHub's API rate limit while checking, so I've stopped querying.- Entire Gates: the same approval gate as on #2533, which needs a reviewer.
- Still running:
test-coreandtest-integration(a, b, c). None have failed.
Once the rate limit resets, I can re-run the test-codex-hooks job. If it still fails, the fix is to pin the Codex version the CI workflow installs (a separate PR). Want me to do that, and start PR 2 in the meantime?
The test-codex-hooks job is re-running on #2624 (GitHub API rate limit reset, back to 5,000 calls). I'll report the result when it finishes.
The re-run passed: test-codex-hooks is green now, so the earlier failure was a flaky Codex install on the CI runner, not the fix. There's nothing to pin.
#2624 now has 17 checks passing and 1 skipped. The only failure is Entire Gates, the reviewer-approval gate, same as on #2533. So both PRs are waiting only on a reviewer.
Should I start PR 2 next?
Yes, I tested #2624 (branch head 758fac0e73) at every level:
| Test | Result |
|---|---|
| New unit tests | TestSplitBatchSpan (7 span shapes) and the RedactBatch regression both failed before the fix and pass with it |
The exact batch that leaked, real opf model | Main leaves "Alice Johnson" in cleartext 3/3; the fix redacts it 3/3 |
Real-binary git push, git-refs | The remote checkpoint is trailered with 0 cleartext names; the note became [REDACTED_PERSON] approved… |
| Normal checkpoint, no merged span | "agent": "Claude Code" stays intact, the same as main |
mise run check locally | Pass: lint 0, unit, integration, canaries 56/56 and 4/4 |
| CI | 17 pass; only Entire Gates (the approval gate) is red |
The binaries and probes were built from the branch contents, and the commit contains exactly those two files.
Not run with this binary:
- git-branch end to end. It calls the same
BatchBytesWithPrivacyFilter, and I checked that call path in code, but I haven't pushed through it. I can run it now: about 2 minutes, same harness. - Re-running every earlier #2533 scenario on top of this fix. They don't touch span mapping, so I'd skip that.
Want me to run the git-branch one?
The fix works on git-branch too, so #2624 is now tested end to end on both storage backends.
| Binary | note.txt on the remote v1 branch | Cleartext "Alice Johnson" |
|---|---|---|
| main | Alice Johnson approved… | 1, pushed with the OPF-applied trailer |
| #2624 | [REDACTED_PERSON] approved… | 0 |
I added this result to the PR's validation section. #2624 is waiting only on a reviewer.
Should I start PR 2?
is it ready for human review? 2624? if so give me the short, slack sentence then we can start on pr 2
Yes, #2624 is ready for human review: CI passes except the approval gate, the trail has no findings, there are no merge conflicts, and it's tested on both storage backends. Confidence is 88, security 12. Risk is 72 because the change sits in the OPF redaction path.
For Slack:
Small OPF privacy fix up for review: names could leak past the filter while the checkpoint was still marked as scanned. Fixed and tested on both storage backends. https://github.com/entireio/cli/pull/2624
Starting PR 2 next.
Fixes an OPF leak where names slipped past the filter. https://github.com/entireio/cli/pull/2624
The 72 is about where the change lives, not about something wrong in it. The monitor's rationale:
Modifies the OPF batch redaction offset-mapping that decides what transcript content is redacted before irreversible push to the remote; a bug in the new span-splitting math could under- or over-redact, though the change actually fixes an under-redaction leak and is small, contained, and covered by targeted regression tests.
- Why it's high: anything that decides what gets redacted before an irreversible push scores high. The offset math is the one place a mistake would matter.
- The monitor already rates the change as safe: Security is 12. It calls the change a strengthening that closes a bypass, with "correct offset clamping, no out-of-bounds/panic risk".
- This is the first score on this trail, and these scores can move ±20 between runs, so one data point isn't a trend.
The one concern it names, a bug in the split math, is also the one thing worth hardening. I can add a Go fuzz test on splitBatchSpan that checks, for random inputs and spans:
- every piece stays inside its input's bounds
- pieces never cover separator bytes
- the pieces together cover exactly the span's overlap with the inputs
Pieces never overlapping falls out of those checks. That's about +40 lines and directly answers the stated risk. The area will still score high, but reviewers get a stronger guarantee than seven table cases.
Want me to add the fuzz test before review, or send it to review as it is?
mise run check is still running in the background, and the commit and push happen after it passes. I'll confirm when that's done.
The fuzz test is committed and pushed to #2624, which is now at fd473e859d.
- Fuzz test:
FuzzSplitBatchSpancompares every result against a simple byte-by-byte model of what should be redacted. About 425,000 random cases over 90 seconds found no mismatches. Its seed cases also run as part of normalgo test. - Checks:
mise run checkpassed (lint clean, all tests green, 56/56 and 4/4 on the end-to-end checks), and so didmise run lintbefore the push. - PR description: now describes the fuzz test and its result.
CI is re-running on the new commit, and the risk score may be re-evaluated too. After that, #2624 is ready for review.
Should I start PR 2 now?
After the fuzz test commit, the risk score dropped from 72 to 38.
| Monitor | Before (758fac0e73) | After (fd473e859d) |
|---|---|---|
| Risk (lower is better) | 72 | 38 |
| Confidence (higher is better) | 88 | 92 |
| Security (lower is better) | 12 | 15 |
| Drift (lower is better) | 6 | 8 |
The new risk rationale: the change "makes it more conservative… reducing PII leak risk with comprehensive test coverage including fuzzing." The concern it had before, a possible bug in the span-splitting math, is gone from the rationale. The +3 on security and +2 on drift are within normal run-to-run variation, and both rationales are still positive.
#2624 is ready for review. The Slack line still works:
Fixes an OPF leak where names slipped past the filter. https://github.com/entireio/cli/pull/2624
Should I start PR 2?
PR 2 is open as #2631 (https://github.com/entireio/cli/pull/2631), stacked on #2533. With it, git push never waits on the privacy filter, on either storage backend, and the background worker pushes the checkpoints itself once they're scanned. Nothing is merged, and CI has just started (12 checks pending).
How it works:
git pushholds back any checkpoint that hasn't been scanned yet, prints a notice, and returns immediately.- A background worker (one per repo) scans the held checkpoints, then pushes them to the same remote.
- Results are kept in a cache in the git directory, so a later push can send already-scanned checkpoints in about a second. The cache holds only hashes and offsets, never text.
Tested on a real binary with the real opf model:
| Scenario | Result |
|---|---|
| git-refs | Push took 0s; the worker delivered both refs about 20s later, redacted |
| git-branch (v1) | Push took 1s; the worker pushed v1 about 15s later, redacted |
| No categories enabled | Held back with the config error; no worker started |
| Remote rejects the worker's push | Checkpoints stayed queued; the next push delivered them in 1s from the cache |
| Second push during a ~45s scan | Returned instantly; no second worker; both checkpoints delivered |
mise run check passes.
A test-isolation bug I introduced and fixed along the way: my first version found the cache via the process's working directory. During testing that wrote 33 entries into this checkout's real .git/entire-opf-cache. The entries only contained offsets, all came from that day's test run, and I deleted them. The cache is now found from the repository object being rewritten, and a regression test covers that.
Known gaps, listed in the PR:
- An explicit
git push origin entire/checkpoints/v1isn't blocked while v1 is unscanned. That's the same on main. - In CI the worker usually dies with the job, so those checkpoints stay local rather than shipping unscanned.
entire cleandoesn't reclaim the new cache yet; entries older than 30 days are pruned automatically.
PR status:
- #2624 (span fix): ready for review; merges on its own.
- #2533: ready for review; its description now names #2631 as the PR to merge with.
- #2631: CI running.
Slack line for #2631:
OPF no longer makes
git pushwait: scanning happens in the background and checkpoints ship when it's done (both storage backends). https://github.com/entireio/cli/pull/2631
Yes. I've now tested both storage backends end to end with a binary built from #2631's pushed head (05e1e64683); the binary's version string matches it.
| Backend | git push | entire status while held | Worker delivers | Remote content | Status after |
|---|---|---|---|---|---|
| git-refs | 0s, both refs held | 2 pending | ~15s later, no second push | trailered, 0 cleartext names, [REDACTED_PERSON] ×4 per ref | cleared |
| git-branch (v1) | 0s, v1 held, tip shown | 3 pending | ~12s later | trailered, 0 cleartext names, [REDACTED_PERSON] ×8 | cleared |
The other three scenarios also ran on this build: no categories enabled, a remote that rejects the worker's push, and a second push during a scan.
Almost: #2631 is ready for review once CI passes on the new head (524201d4c9).
What changed since my last message:
- Trail finding fixed and resolved. If the scan cache can't be opened (bad permissions, full disk), pre-push used to show the "scanning in the background" notice forever. It now withholds the checkpoints with an error that names the cause and the fix, and starts no worker. A regression test covers it.
- Fuzz test added. It targets the one concrete worry in the trail's risk score: that under-redacted content could be stamped as scanned. For random inputs, it checks that applying cached results matches a one-pass scan, and that a single missing leaf makes the apply refuse. About 610,000 inputs over 90 seconds, with no failures.
- Checks:
mise run checkandmise run lintpass, and the PR description is updated.
What a reviewer should know:
- Risk is 83. The rationale is the scale of the change: it rebuilds the redaction path, the worker pushes on its own, and the caps went up. The fuzz test answers the one concrete concern, but expect the score to stay high. Security is 24 and confidence 86, and the security rationale says the fail-closed guarantees are preserved.
- It's large: about +2.7k lines, stacked on #2533 and merging together with it. The description lists where to start reviewing.
I'll re-check CI and the risk score once they update on the new head.
Yes, #2631 is ready for human review. Every CI test passes on the current head (524201d4c9), the only red check is the approval gate, and there are no open findings.
| Monitor | Before the last fixes | Now | Rationale |
|---|---|---|---|
| Risk (lower is better) | 83 | 68 | "Preserves fail-closed semantics". The remaining concern is async failure modes (held too long, delivered wrongly), which the worker tests and real-binary runs cover |
| Security (lower is better) | 24 | 35 | Everything uses the existing safety primitives. One nit: the spawn-marker helper doesn't validate its marker name, though every caller passes a fixed literal |
| Confidence (higher is better) | 86 | 73 | Credits the fuzz and unit coverage; takes points off for worker complexity and because the monitor itself can't run the CI gates |
| Drift (lower is better) | 14 | 12 | — |
The confidence drop cites nothing specific to fix, and these scores move ±20 between runs.
The security nit is a real, if small, gap: about 10 lines to have spawnmarker reject names that aren't plain words. I'd add it before review so the reviewer doesn't hit it.
Should I add it, or send #2631 to review as is? The Slack line:
OPF no longer makes
git pushwait: scanning happens in the background and checkpoints ship when it's done (both storage backends). https://github.com/entireio/cli/pull/2631
They don't have to merge together, but #2631 does depend on #2533.
- #2631 needs #2533. Its cache-only rewrite and delivery sit on #2533's per-ref loop and exact-generation queue. It can't target main on its own.
- #2533 can merge alone. After the rework it only fixes correctness, with main's caps and inline scan, so shipping it without #2631 makes nothing worse. My earlier "merge them together" dates from before the rework and no longer applies.
The options:
- Keep the stack (recommended). Review and merge #2533 first. GitHub retargets #2631 to main automatically once #2533's branch is merged and deleted. Two smaller reviews, and #2533 isn't blocked by #2631's larger review.
- One PR: point #2631 at main and close #2533. That's one review of about +4.3k lines, the size problem we split to avoid. It also means closing a PR that's already in review.
I'd keep the stack and drop the "merge together" line from both PR descriptions, since it now overstates the dependency. Should I update the descriptions, and add the small spawn-marker check to #2631?
Yes, both are ready for human review.
| PR | Head | CI | Findings | Merge conflicts |
|---|---|---|---|---|
| #2533 | faca8396c9 | 16 pass, 1 skipped | none | none |
| #2631 | 524201d4c9 | 12 pass | none | none |
The one red check on each is Entire Gates, which waits on a reviewer's approval.
Both PR descriptions still say "merge these together". That's no longer true: #2533 is safe to merge alone, and only #2631 needs #2533. Worth fixing so reviewers aren't misled; I can correct both descriptions if you want.