Fix Entire Search Native Repository Scope

Investigate the reported behavior: entire search errors on Entire-native repositories with "not a GitHub repository", even when the caller supplies an explicit --repo.
This is an investigation only. Do not implement a broad fix. Do not push, open or update a PR, merge, deploy, write to production, alter credentials, or change another agent/worktree. You may make local investigation notes in this isolated worktree and run routine read-only repository searches, exact CLI reproductions, and focused local tests. Preserve all unrelated files and sessions.
Verify first. Do not assume the report is still current. Before designing anything, confirm the problem is real and current: (1) Datadog logs/metrics for the reported window and the last 24h, if they materially distinguish server behavior from a local client rejection; keep any access read-only; (2) live cluster state via AWS SSO — kubectl / natsctl / entire api — and the deployed build, only if needed to distinguish supported backend capability from premature client validation, and read-only only; (3) the code path on current main as it is now, cited from this worktree; (4) Entire search (search_entire, search_entire_code, and the entire search skill) for prior checkpoints, sessions, trails, or a later fix on this symptom. If a source is not applicable because the failure occurs before any request leaves the client, say so explicitly and show the evidence. Write evidence and one explicit verdict — confirmed / already fixed / different than described — at the top of FINDINGS.md before any code changes. If already fixed or materially different, report that and stop; do not build the described fix.
Read AGENTS.md before work and follow its linked CLI conventions, API routing, testing, and git/subprocess guidance where relevant. This worktree is fresh from origin/main; do not switch branches or update main.
Investigation scope:
- Reproduce the exact native-repository error with implicit repository discovery and with explicit
--repoforms. Record the exact commands, environment assumptions, stdout/stderr, and exit status. Use a safe public/read-only native repo if one is already configured; do not create or mutate repositories. - Identify whether installed/released behavior differs from this worktree's current code. Record the exact CLI version/binary used for each reproduction.
- Trace where local remote/host validation happens and determine whether it incorrectly overrides or rejects explicit repository selection.
- Trace native repository identifier parsing and routing support end-to-end for semantic search, including accepted
et/...,/et/..., and full Entire clone URL forms where relevant. - Identify the relevant tests and coverage gaps. Run focused safe tests that prove the current behavior; do not run paid real-agent E2E tests.
- Separate two questions: whether semantic search for native repos is actually supported by the backend, and whether the CLI fails prematurely with GitHub-only validation before it could try the supported path.
- Compare with relevant completed work such as ENT-2195, ENT-2196, ENT-2365 and PRs #2245/#2247, but do not assume those changes cover search.
- Produce minimal fix options with tradeoffs and a recommended smallest change, but do not implement it unless explicitly asked in a later turn.
When root-caused, present the human-readable result using the architecture-findings skill: explain the players, normal request flow, where it breaks, whether backend capability exists, and where the minimal correction belongs, without file paths, line numbers, or symbol names in that narrative. Keep code-level references, commands, test names/results, and exact evidence in FINDINGS.md for follow-up.
When finished, stop and report the verdict, root cause, evidence, test results, and recommendation. Leave the Herdr session and worktree intact.
Using superpowers:systematic-debugging to structure the investigation, then search-test for Entire history.
Base directory for this skill: /Users/alisha/.claude/plugins/cache/claude-plugins-official/superpowers/6.4.1/skills/systematic-debugging
Systematic Debugging
Overview
Core principle: ALWAYS find root cause before attempting fixes. Symptom fixes are failure.
Violating the letter of this process is violating the spirit of debugging.
The Iron Law
If you haven't completed Phase 1, you cannot propose fixes.
When to Use
Use for ANY technical issue:
- Test failures
- Bugs in production
- Unexpected behavior
- Performance problems
- Build failures
- Integration issues
Use this ESPECIALLY when:
- Under time pressure (emergencies make guessing tempting)
- "Just one quick fix" seems obvious
- You've already tried multiple fixes
- Previous fix didn't work
- You don't fully understand the issue
Don't skip when:
- Issue seems simple (simple bugs have root causes too)
- You're in a hurry (rushing guarantees rework)
- Manager wants it fixed NOW (systematic is faster than thrashing)
The Four Phases
You MUST complete each phase before proceeding to the next.
Phase 1: Root Cause Investigation
BEFORE attempting ANY fix:
-
Read Error Messages Carefully
- Don't skip past errors or warnings
- They often contain the exact solution
- Read stack traces completely
- Note line numbers, file paths, error codes
-
Reproduce Consistently
- Can you trigger it reliably?
- What are the exact steps?
- Does it happen every time?
- If not reproducible → gather more data, don't guess
-
Check Recent Changes
- What changed that could cause this?
- Git diff, recent commits
- New dependencies, config changes
- Environmental differences
-
Gather Evidence in Multi-Component Systems
WHEN system has multiple components (CI → build → signing, API → service → database):
BEFORE proposing fixes, add diagnostic instrumentation:
Example (multi-layer system):
This reveals: Which layer fails (secrets → workflow ✓, workflow → build ✗)
-
Trace Data Flow
WHEN error is deep in call stack:
See
root-cause-tracing.mdin this directory for the complete backward tracing technique.Quick version:
- Where does bad value originate?
- What called this with bad value?
- Keep tracing up until you find the source
- Fix at source, not at symptom
Phase 2: Pattern Analysis
Find the pattern before fixing:
-
Find Working Examples
- Locate similar working code in same codebase
- What works that's similar to what's broken?
-
Compare Against References
- If implementing pattern, read reference implementation COMPLETELY
- Don't skim - read every line
- Understand the pattern fully before applying
-
Identify Differences
- What's different between working and broken?
- List every difference, however small
- Don't assume "that can't matter"
-
Understand Dependencies
- What other components does this need?
- What settings, config, environment?
- What assumptions does it make?
Phase 3: Hypothesis and Testing
Scientific method:
-
Form Single Hypothesis
- State clearly: "I think X is the root cause because Y"
- Write it down
- Be specific, not vague
-
Test Minimally
- Make the SMALLEST possible change to test hypothesis
- One variable at a time
- Don't fix multiple things at once
-
Verify Before Continuing
- Did it work? Yes → Phase 4
- Didn't work? Form NEW hypothesis
- DON'T add more fixes on top
-
When You Don't Know
- Say "I don't understand X"
- Don't pretend to know
- Ask for help
- Research more
Phase 4: Implementation
Fix the root cause, not the symptom:
-
Create Failing Test Case
- Simplest possible reproduction
- Automated test if possible
- One-off test script if no framework
- MUST have before fixing
- Use the
superpowers:test-driven-developmentskill for writing proper failing tests
-
Implement Single Fix
- Address the root cause identified
- ONE change at a time
- No "while I'm here" improvements
- No bundled refactoring
-
Verify Fix
- Test passes now?
- No other tests broken?
- Issue actually resolved?
- Use the
superpowers:verification-before-completionskill before claiming success
-
If Fix Doesn't Work
- STOP
- Count: How many fixes have you tried?
- If < 3: Return to Phase 1, re-analyze with new information
- If ≥ 3: STOP and question the architecture (step 5 below)
- DON'T attempt Fix #4 without architectural discussion
-
If 3+ Fixes Failed: Question Architecture
Pattern indicating architectural problem:
- Each fix reveals new shared state/coupling/problem in different place
- Fixes require "massive refactoring" to implement
- Each fix creates new symptoms elsewhere
STOP and question fundamentals:
- Is this pattern fundamentally sound?
- Are we "sticking with it through sheer inertia"?
- Should we refactor architecture vs. continue fixing symptoms?
Discuss with your human partner before attempting more fixes
This is NOT a failed hypothesis - this is a wrong architecture.
Red Flags - STOP and Follow Process
If you catch yourself thinking:
- "Quick fix for now, investigate later"
- "Just try changing X and see if it works"
- "Add multiple changes, run tests"
- "Skip the test, I'll manually verify"
- "It's probably X, let me fix that"
- "I don't fully understand but this might work"
- "Pattern says X but I'll adapt it differently"
- "Here are the main problems: [lists fixes without investigation]"
- Proposing solutions before tracing data flow
- "One more fix attempt" (when already tried 2+)
- Each fix reveals new problem in different place
ALL of these mean: STOP. Return to Phase 1.
If 3+ fixes failed: Question the architecture (see Phase 4.5)
your human partner's Signals You're Doing It Wrong
Watch for these redirections:
- "Is that not happening?" - You assumed without verifying
- "Will it show us...?" - You should have added evidence gathering
- "Stop guessing" - You're proposing fixes without understanding
- "Ultra-think this" - Question fundamentals, not just symptoms
- "We're stuck?" (frustrated) - Your approach isn't working
When you see these: STOP. Return to Phase 1.
Common Rationalizations
| Excuse | Reality |
|---|---|
| "Issue is simple, don't need process" | Simple issues have root causes too. Process is fast for simple bugs. |
| "Emergency, no time for process" | Systematic debugging is FASTER than guess-and-check thrashing. |
| "Just try this first, then investigate" | First fix sets the pattern. Do it right from the start. |
| "I'll write test after confirming fix works" | Untested fixes don't stick. Test first proves it. |
| "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. |
| "Reference too long, I'll adapt the pattern" | Partial understanding guarantees bugs. Read it completely. |
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause. |
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem. Question pattern, don't fix again. |
Quick Reference
| Phase | Key Activities | Success Criteria |
|---|---|---|
| 1. Root Cause | Read errors, reproduce, check changes, gather evidence | Understand WHAT and WHY |
| 2. Pattern | Find working examples, compare | Identify differences |
| 3. Hypothesis | Form theory, test minimally | Confirmed or new hypothesis |
| 4. Implementation | Create test, fix, verify | Bug resolved, tests pass |
When Process Reveals "No Root Cause"
If systematic investigation reveals issue is truly environmental, timing-dependent, or external:
- You've completed the process
- Document what you investigated
- Implement appropriate handling (retry, timeout, error message)
- Add monitoring/logging for future investigation
But: 95% of "no root cause" cases are incomplete investigation.
Supporting Techniques
These techniques are part of systematic debugging and available in this directory:
root-cause-tracing.md- Trace bugs backward through call stack to find original triggerdefense-in-depth.md- Add validation at multiple layers after finding root causecondition-based-waiting.md- Replace arbitrary timeouts with condition polling
Tracing where the GitHub-only check sits relative to --repo, and how the semantic searcher uses owner/repo. Pulling the remaining code, history, and prior Entire work in parallel.
Continue the investigation. The previous command was denied only because it wrote under ~/.claude. Use /tmp/native-search-repo-routing/entire for the worktree binary instead, then complete the requested exact reproductions, focused tests, FINDINGS.md, architecture-findings narrative, and final verdict. Keep all access read-only except the /tmp binary and FINDINGS.md in this isolated worktree; do not implement a fix.
Owner/repo from the GitHub check feed only the default scope, nothing on the wire. Checking the secondary --code index miss, then writing FINDINGS.md and loading the narrative skill.
Base directory for this skill: /Users/alisha/.claude/skills/architecture-findings
Architecture findings
What this is for
After an investigation, there are two very different documents you could hand someone:
- A lookup table — every claim pinned to a file and line so an engineer can jump straight to the code. Precise, but dense and exhausting to read.
- A map of the territory — who the moving parts are, how they normally fit together, where the break is, and whose job it is to fix. This is the one that actually transfers understanding.
This skill produces the second kind. The reader should come away able to explain the system to someone else, not just able to find the line you're talking about.
The core rule
No code coordinates in the body. No file paths, no line numbers, no function or variable names, no code snippets. Those are the texture of the lookup table, and they're exactly what makes findings hard to read.
What you do name freely: the components, services, and systems by their
real names (e.g. "the EU indexer", "entire-core", "the NATS bus"), the
concepts ("ref events", "the repo-access token"), and the environment
values that matter to the story only when named as a fact, not a citation
("the EU audience host is aws-eu-central-1.entire.io"). The test: would this
detail help someone picture the system, or is it just a pointer for someone
who wants to open a file? Keep the former, drop the latter.
You still did the rigorous, code-level investigation — this is just how you report it. If the user later asks "where exactly", give them the coordinates then. Don't pre-empt that with a wall of citations.
Structure
Adapt to the investigation, but this arc almost always works:
- Bottom line — 2–4 sentences. What's happening, why, and the one thing that matters most (e.g. "the fix isn't where you'd expect — it's upstream"). Someone who reads only this should get the headline right.
- The players — introduce each component in one line: what it is and what it's responsible for. This is the cast list; everything after refers back to it.
- The normal flow — how it works when it works. This is the reference the reader measures the break against. A simple arrow-chain diagram earns its place here (see below).
- Why it breaks — the failure, told as a story against that normal flow. Show the point where reality diverges from the happy path.
- The structural gap (when there is one) — the deeper "why can't we just fix it here" — the missing capability or wrong assumption underneath the symptom. This is often the most valuable section: it's what stops the reader proposing the naive fix.
- Where the fix belongs — which system owns it and why, what the options are, and explicitly what the wrong fix is if there's a tempting one.
- How other cases differ (when relevant) — the comparison that proves it's structural, not a one-off (other regions, other tenants, the working vs. broken side).
Drop sections that don't apply. A small investigation might be just bottom line + why it breaks + where the fix belongs.
Flow diagrams
Simple ASCII arrow chains are encouraged where a flow clarifies the story — they carry sequence better than a paragraph. Keep them component-level, never code-level. For example:
Then a second chain showing where the broken path diverges is often the clearest way to show a failure. Label them ("normal flow" / "what EU does"). Don't force a diagram where prose is already clear.
Voice
Write for a smart colleague who doesn't live in this code. Explain the why behind each mechanism, not just the what — "the STS refuses because that repo isn't an EU repo; that refusal is the residency guard doing its job" teaches more than "the STS returns an error". Prefer plain language over jargon, and when a term is load-bearing, define it once in passing. Keep it tight — this is a briefing, not an essay.
What to avoid
- File paths, line numbers, symbol names, code blocks (other than the component-level flow diagrams above).
- Turning "the players" into an inventory of every file touched.
- Hedging everything. You investigated; state what you found plainly, and flag genuine open questions as open questions.
- Burying the headline. The bottom line goes first.
Verdict: confirmed. Reproduced on both the worktree build of main and the installed nightly. Not fixed anywhere. Details, commands, versions, and tests are in FINDINGS.md in this worktree.
Bottom line
entire search dies inside the CLI before it talks to anything. It reads the clone's origin, insists the origin be GitHub-backed, and only then would look at --repo or --all-repos. An Entire-native clone has an origin that names an Entire cluster and the et forge, so the check fails every time, whatever flags you pass. The backend is not the problem: the identical binary, launched from a GitHub clone with --repo et/entirehq/entire-api, returns native-repo results from production. The fix belongs in the search command's scope resolution, not in the backend and not in the shared GitHub parser.
The players
- The search command in the CLI: turns the query and flags into a scope and hands it to the semantic searcher.
- The remote parser: understands plain GitHub URLs and Entire clone URLs, and reports which forge backs a clone,
ghoret. - The GitHub-only check: a thin policy on top of the parser that the search command, and the dispatch commands, use to insist on a GitHub origin.
- entire-core, the control plane: holds the repo index, where native repos are stored under forge-prefixed names, and resolves a name to a repo ID and home cell.
- query-serve in each cell: the semantic index that answers scoped searches by repo ID.
The normal flow
When no --repo is given, the scope is the current repo, derived from origin. Owner and repo from origin are used only to build that default name. They never reach the wire.
Why it breaks
The command runs the GitHub-only check on origin unconditionally, before it decides whether it even needs a default scope. For a native clone the parser correctly says "forge is et, no GitHub upstream", the check reports the cluster host as "not GitHub", and the command exits. Explicit --repo et/..., --repo gh/..., and --all-repos are all discarded because the error fires first. Two other input shapes, a leading-slash /et/... and a full clone URL, fail even earlier in filter validation, which only accepts the bare prefixed form.
The structural gap
The check was loosened in June to accept GitHub mirrors reached through an Entire clone URL, but kept its GitHub-only policy on purpose. When native repos were taught to the trail commands and to checkpoint explain in September, the search command's help text was updated to advertise native repos, but its origin handling was not. The code-search branch of the same command already does the right thing: it builds a forge-prefixed default name and never calls the GitHub check. The semantic branch simply predates that pattern.
One trap: skipping the check and keeping the bare owner/repo default would be wrong too. entire-core treats a bare pair as "either forge", so a native repo and a same-named GitHub mirror would both match. The default must carry the forge prefix, as the control plane and the index already do.
Where the fix belongs
The search command, at the point where it derives the default scope. Parse origin with the forge-aware parser, seed the default scope with the forge-prefixed name, and fail on an unparseable origin only when no explicit scope was supplied, mirroring the code-search branch. That one change fixes implicit, --repo, and --all-repos in native clones. Do not widen the shared GitHub-only check or teach the parser that et maps to a host: dispatch relies on that check staying strict.
Optional follow-up: accept /et/... and full clone URLs as filters, as checkpoint explain already does.
Remaining limitation
entire search --code against a native repo reaches the backend but reports no matching repository, from both native and GitHub clones. Code search matches filters against the unfiltered repo listing, while semantic search uses the server's exact-name filter. Not investigated further per your instruction; noted in FINDINGS.md as a separate ticket candidate.
Worktree, Herdr session, and all other files are untouched. The only artifacts are FINDINGS.md here and the binary at /tmp/native-search-repo-routing/entire.
Base directory for this skill: /Users/alisha/.claude/skills/rubber-duck
Rubber Duck
The user is looking at something dense — usually a long agent response or a findings dump — and wants to understand it, not act on it. Your whole job is to be the clear, patient explainer and then a sounding board for their questions. Think out loud with them, in plain language.
The one job
Turn the wall of text into something a smart person with no context can follow in under a minute, then answer whatever they ask next — still briefly.
Hard rules (do not break these)
- Do NOT write a plan, and do NOT enter plan mode. No ExitPlanMode, no writing-plans/brainstorming skills, no "here's how we'd fix it" roadmaps.
- Do NOT change code, run builds, or edit files. This is understanding only.
- Do NOT make a todo list. It's a conversation, not a task.
- Stay read-only. Reading a file or scrolling the text to explain it is fine. Anything that does something is not.
- Be concise — this is the point of the skill. You are the antidote to the wall of text, so don't produce another one. Lead with the answer. Short sentences. No preamble, no "great question", no restating the question back.
- Plain language over jargon. No file paths, line numbers, symbol names, or
code snippets unless the user asks for them. Names of systems/people are fine;
citations are not. (This is the same spirit as the
architecture-findingsskill — the story, not the lookup table.)
First explanation — the format
Keep it to roughly this, and stop:
- One-line TL;DR. The single most important thing, first.
- 3–5 short bullets covering: what's actually going on, why it matters, and (if relevant) what's broken or being decided. One idea per bullet.
- One line on "so what" — the upshot, or the decision in front of them.
If the source genuinely has more than fits, say "there's more detail under X if you want it" and let them pull it — don't push it.
Use an analogy when it makes a hard idea click. Skip it when it doesn't.
The follow-up loop
After the first explanation, they'll ask things. For each:
- Answer the specific question and nothing more. One concept at a time.
- Stay just as concise — don't let answers creep back into walls of text.
- It's fine to say "not sure / the text doesn't say" rather than guess.
- If a question needs you to look something up to answer honestly, look (read only), then answer plainly.
Keep going as long as they're asking. You're the duck — they're thinking out loud and you're helping them see it.
When they're ready to act
The moment the user shifts from "help me understand" to "okay let's fix/build/ plan this," rubber-duck mode is over. Don't quietly start doing it inside this mode. Say something like: "That's action territory — want me to pick up that work?" and let them start the real task (a plan, a dispatch, an edit) as its own thing. Your job here ends at understanding.
TL;DR: The search command checks "is your origin GitHub?" before it looks at anything else, so native-repo clones always fail locally. The backend is fine.
- What happens: In a native clone, origin points at an Entire cluster with an
etpath. The command reads that, says "not GitHub", and exits. Your--repoor--all-reposis never read because the error fires first. - Proof the backend works: Same binary, run from a GitHub clone with
--repo et/entirehq/entire-api, returns native-repo results from production. - Why it's there: That GitHub check was built when search only knew GitHub. In September native repos were added to trails and
checkpoint explain, and the search help text was updated, but this one check was missed. - The trap: You can't just delete the check. Without the
et/prefix, the default scope would be a bare "owner/repo", which the control plane treats as "either forge," so a native repo and a same-named GitHub mirror would both match. - Already-fixed sibling: The
--codebranch of the same command does it right: it builds the prefixed name and never runs the GitHub check.
So what: It's a small fix in one place, the search command's default-scope logic, mirroring what --code does. Nothing to change in the backend or the shared GitHub parser (dispatch still needs that strict).
There's a side note in FINDINGS.md about --code failing differently on native repos if you want it.
Fixed, mirroring the --code branch. Uncommitted in this worktree.
Change
cmd/entire/cli/search_cmd.go: the GitHub-only origin check is gone. Semantic search now derives the default scope from origin's forge-qualified coordinates via a newresolveDefaultSearchRepohelper. With an explicit--repo,repo:, or--all-repos, an unreadable origin is ignored. Without one, it errors with the same hint--codeuses: "could not determine current repository for search (use --repo or --all-repos)".cmd/entire/cli/search/search.go:Config.Forgeadded; the default slug is nowet/proj/repoorgh/owner/repo, so a native repo and a same-named mirror can't be conflated.ParseGitHubRemoteis untouched; dispatch still relies on it.
Tests (written first, failed, then passed)
- New
TestResolveDefaultSearchRepo: native, GitHub, mirror origins, explicit-scope, and no-origin with and without explicit scope. - Three new
ScopeSlugscases for forge-prefixed defaults. - Full
cmd/entire/cli,search, andgitremotepackages pass.mise run fmtclean,mise run lint0 issues.
End-to-end with the rebuilt binary from ~/Projects/entire-api: implicit, --repo et/..., --all-repos, and --repo gh/entireio/cli all exit 0 with correct hits. Implicit from a GitHub clone still works. Details appended to FINDINGS.md.
Not done: mise run check (integration + canary) and the commit. --code on native repos still reports no matching repo, which is the separate index issue noted in FINDINGS.md.
Yes. I rebuilt the binary from the fixed tree and ran it against production from the native clone at ~/Projects/entire-api (origin entire://aws-us-east-2.entire.io/et/entirehq/entire-api):
| Command (from the native clone) | Before | After |
|---|---|---|
search 'routing' --json --compact --limit 1 | exit 1, "not a GitHub repository" | exit 0, commit hit in entirehq/entire-api |
--repo et/entirehq/entire-api | exit 1 | exit 0, same native hit |
--all-repos | exit 1 | exit 0, hit in entirehq/company-knowledge (native) |
--repo gh/entireio/cli | exit 1 | exit 0, hit in entireio/cli |
Also checked from /tmp (no git repo): implicit gives the --code-style hint, --repo et/entirehq/entire-api returns results. And from this GitHub worktree, implicit still resolves to entireio/cli.
All read-only searches. Exact output is in the "Fix applied" section of FINDINGS.md.
Base directory for this skill: /Users/alisha/.claude/skills/pr-ready
Getting a PR Ready for Review
The goal is a branch that a reviewer can look at with confidence: covered by a real test, caught up with the base branch, no conflicts, no regressions, a PR open (ready for review), and CI green. This skill is the step before requesting review — it does not merge the PR and does not post it anywhere.
Work through the phases below in order, but scope to what's needed. If the branch is already merged-clean and only CI is red, skip straight to the CI phase. The one phase you never skip is the test-coverage check (Phase 2) — a branch with no test for what it changed isn't review-ready no matter how green CI is. If there's no remote branch yet, you'll still do everything. Tell the user which phases you're running and why.
Guiding principles
These matter more than the mechanics — internalize them, because the mechanics serve them:
- Never mask a problem to make a check pass. Deleting a failing test,
loosening an assertion, adding
// eslint-disableover a real lint error,--no-verify, orcontinue-on-errorturns a red check green while leaving the bug in place. That defeats the entire point of getting review-ready. Fix the underlying cause. If a check is failing for a reason genuinely unrelated to this branch (flaky test, pre-existing breakage on main), say so explicitly and ask the user rather than papering over it. - Conflict resolution preserves both intents. A merge conflict means two
changes touched the same place. Resolving it is not "pick a side" — it's
understanding what each side was trying to do and producing code that honors
both. Blindly taking
--oursor--theirsis how regressions sneak in. - A change without a test that would fail without it isn't done. The test is what stops this bug or feature from silently regressing in six months when someone refactors past it. A test that passes equally well against the pre-change code covers nothing, however thorough it looks — prove the red before you trust the green.
- Verify before you push, then verify again in CI. Local verification is fast and catches most regressions before they cost a CI cycle. CI is the source of truth. Both matter.
- Review your own diff before you ship it. Build/test/lint prove the code runs; they don't prove the change is clean. Read the full branch diff with fresh eyes for two things a green build hides: regressions the change slipped in (behavior quietly altered, a caller left on an old contract, a dropped edge case) and dead or unused code it created (functions/params/imports/types nothing references, unreachable branches, leftover debug or commented-out code, scaffolding wired to nothing). Ship the change, not the scaffolding around it.
- Stop and ask when the right answer is ambiguous, especially on conflict resolution where the two sides represent real product decisions. A wrong guess here is worse than a question.
Phase 1 — Orient
Establish the ground truth before changing anything. Run these read-only:
git rev-parse --abbrev-ref HEAD— the current branch. If it'smain/master, stop: there's nothing to ready-for-review. Ask the user which branch they mean.git status— confirm a clean working tree. If there are uncommitted changes, surface them and ask whether to commit, stash, or abort. Don't merge on top of a dirty tree — it tangles your changes with the merge.- Determine the base branch. It's usually
main, sometimesmasteror a release branch. Checkgit remote show origin | grep "HEAD branch"or look at what the existing PR targets (Phase 5). Don't assumemain. gh pr view --json number,state,isDraft,baseRefName,url 2>/dev/null— does a PR already exist for this branch, and what does it target? Remember the answer for Phase 5.
Report a one-line plan: which base, whether a PR exists, and which phases you'll run.
Phase 2 — Prove a real test covers the change
Do this before touching the branch. A change that ships without a test that exercises it will regress — someone will refactor past it, and nothing will notice. This is the first substantive step, ahead of the merge, so a missing test is written against the branch's own code rather than bolted on after a conflict resolution has muddied it.
Read the branch diff and identify what it actually changed:
For each behavioral change the branch makes (a bug fixed, a feature added, a
rule or edge case altered), find the test that covers it. Look in the diff first
(--diff-filter=AM over the repo's test paths), then in existing test files near
the changed code — an existing test extended with a new case counts.
A test only counts as real coverage if all of these hold:
- It would fail without the change. This is the whole bar. For a bugfix, the test reproduces the original symptom; for a feature, it asserts the new behavior. If it passes equally well against the pre-change code, it covers nothing.
- It asserts behavior, not shape. Checking that a function was called, that a struct has a field, or that a string matches the new copy is not coverage. Assert the outcome a user or caller would observe.
- It goes through the real path. A test that mocks out the very thing the change fixed proves only that the mock works. Exercise the actual code under test; mock at the system boundary, not around the change.
- It covers the edge case the bug lived in, not just the happy path. Bugs live in the empty list, the nil, the concurrent second call, the expired token.
Verify the "would fail without it" claim — don't assert it from reading. Cheapest reliable check: stash or revert the source change (not the test) and run the test, confirm it fails, then restore and confirm it passes.
If stashing is awkward (a large or tangled diff), invert the key assertion or temporarily neuter the fixed line instead — anything that gives you a real red before the green. Record what you saw; "the test passes" alone is not evidence.
If coverage is missing or weak, write the test now — before the merge, before the push. Follow the repo's existing test conventions (framework, fixtures, naming, table-driven style) rather than inventing a new pattern. Put it at the level that makes the guarantee durable: prefer a unit test close to the logic when that genuinely exercises the bug, and reach for an integration/e2e test when the bug only exists in the seam between components.
Then run the new test and confirm it fails against the unfixed code, per above.
Legitimate exceptions — state them explicitly rather than skipping silently:
- Pure refactors, renames, formatting, dependency bumps, or comment/doc-only changes that alter no behavior — the existing suite passing is the coverage.
- Changes to code the repo has no practical way to test (generated files, infra/config, a thin shell over an external API with no harness). Say so, and say what manual verification you did instead.
- The user explicitly waives it.
If you believe an exception applies but aren't sure, ask — don't decide for the user that a change is untestable. If a test is genuinely needed but writing it requires a product decision (what should the behavior be at the edge?), stop and ask.
Report, in one or two lines: which test covers each behavioral change, whether it already existed or you wrote it, and the red-then-green evidence you observed.
Phase 3 — Sync the base branch
Get the latest base so the merge reflects reality:
Fetching is enough — you merge origin/<base> directly in the next phase, so you
don't need to check out and fast-forward your local base branch. (If the user
specifically wants their local base updated too, git checkout <base> && git pull --ff-only && git checkout - does it without surprises.)
Phase 4 — Merge the base in and resolve conflicts
The user asked for a merge, not a rebase — keep it a merge so history and any in-flight review threads stay intact:
If it merges cleanly, go to Phase 5 after a quick local verification.
If there are conflicts, resolve them deliberately:
git statusandgit diffto see every conflicted file. Work through them one at a time — don't bulk-resolve.- For each conflict, understand both sides before writing the resolution:
git log --merge -p <file>shows the commits from each side touching the conflict region — this is the fastest way to see intent.- If the repo is Entire-enabled and the intent is still unclear, the
what-happened/explainskills can recover the reasoning behind a block. - Read enough surrounding code to know what each change was accomplishing.
- Write a resolution that keeps both intents. If side A renamed a function and side B added a new caller of the old name, the resolution updates the caller — it doesn't drop either change. If the two intents genuinely conflict (both changed the same value to different things), that's a product decision: stop and ask the user.
git addeach resolved file as you finish it.- When all conflicts are resolved, verify before committing the merge (next
phase). Don't
git committhe merge until the code at least builds — a merge commit that doesn't compile is a regression you're about to push.
Phase 5 — Verify locally & self-review the diff
A merge can produce code that compiles on each side but breaks when combined ("semantic conflicts" git can't detect). Catch these now, cheaply.
Run the project's real verification — discover it rather than guessing. Look for
the commands the repo actually uses (a project skill for running the app/tests,
Makefile/justfile targets, package.json scripts, go test ./..., CI config
under .github/workflows, CONTRIBUTING.md). Run, in roughly this order, what
exists: build → unit tests → lint/typecheck.
cli repo — do NOT run the e2e suite locally (temporary, 2026-07-20). The e2e suite (
mise run test:e2e*, andtest:ci/checkwhich call it) deletes the developer's real~/.claude/projectstranscripts via an unmerged HOME- isolation bug. The tasks are locally gated to SKIP (exit 0) — that skip is intentional; do not "fix" it or setENTIRE_ALLOW_LOCAL_E2E=1. Verify withmise run test(unit) +mise run lintonly; CI runs e2e on isolated runners.
- If verification passes, commit the merge (
git commitwith no-maccepts the default merge message, or write a short one) and continue. - If it fails, fix the cause — a semantic conflict you missed, an import that moved, a renamed symbol. Re-run until clean, then commit the merge.
Skipping this phase just moves the failure to CI, where it's slower to diagnose. It's worth doing even when the merge was conflict-free.
Self-review the branch diff — regressions & dead code
Verification proves the code runs; it doesn't prove the change is clean. Before pushing, read the branch's full diff against the base with fresh eyes — this is a review of the code this branch introduces, separate from the build/test pass above:
Look for two classes of problem a green build won't reliably flag:
- Regressions the change introduced — behavior that changed when it shouldn't have; a renamed or re-signatured function with a caller left on the old contract; an error path, guard, or edge case dropped in an edit or a conflict resolution; a default or constant that silently shifted. Trace each non-trivial hunk to the callers it affects and confirm they still hold.
- Dead / unused code the change created — new functions, methods, variables, parameters, struct fields, types, imports, or constants that nothing references; branches that can no longer be reached; a helper added but never wired in; a feature-flag path with no live caller; leftover debug logging, commented-out code, or TODO stubs. Review naturally sees the addition and rarely notices the thing is never used.
Lean on tooling to find the unused set, then confirm by hand — don't trust it alone:
- Compiler/linter unused checks (discover what the repo uses): Go flags unused
locals/imports at build and surfaces unused functions via
go vet/staticcheck/golangci-lint(andgolang.org/x/tools/cmd/deadcodefor whole-program dead functions); TypeScript viatsc --noUnusedLocals --noUnusedParametersand eslintno-unused-vars. - For each new exported symbol the diff adds, grep the repo for references — zero hits outside its own definition and test means it's dead.
- If a
dead-code-finderskill is available, point it at the branch's changes.
Resolve what you find: remove dead code this branch introduced (if it's deliberate scaffolding for an imminent follow-up, say so explicitly and keep it only with the user's nod); fix any regression at its root, never by loosening a test. If deleting suspected-dead code is risky — it may be reached via reflection, build tags, code generation, or an external consumer — confirm before removing it rather than guessing. Keep the scope to this branch's diff; don't go hunting pre-existing dead code across the repo — that's not what getting this PR ready is about.
Phase 6 — Push
Before you push, strip scratch working docs from the branch. Agents routinely
leave planning/findings scratch files in the worktree — FINDINGS.md, PLAN.md,
plan.md, *-seed.md, DAY_DRIVER_TASK.md, NOTES.md, TODO.md, design/
investigation write-ups — that are working artifacts, not part of the change.
They must never ship in the PR. Catch them here, before the branch goes remote.
- List
.mdfiles this branch adds vs the base, plus any untracked ones in the worktree: - Decide each candidate by intent, not just extension:
- Remove agent-generated scratch — findings / plan / seed / task-handoff / notes docs (typically at the repo root or a scratch dir) created by this branch and unrelated to the change under review.
- Keep genuine documentation — anything under
docs/,README*,CHANGELOG*, ADRs,CONTRIBUTING*, or a.mdthe change legitimately adds (e.g. docs for the feature you just built). When you're unsure whether a doc is intended, ask rather than delete.
- Remove the scratch docs so they can't reach the PR — but don't destroy work
you can't recover: if a findings/investigation write-up looks worth keeping,
copy it to a durable scratch location (
~/.claude/tmp/…) first, then delete.- Untracked (never committed): delete from disk (
rm <file>). - Already committed on this branch:
git rm <file>and commit the removal (git commit -m "chore: drop scratch planning/findings docs").
- Untracked (never committed): delete from disk (
Then push:
Never force-push — not --force, not --force-with-lease, not under any
circumstance. If a plain push is rejected because the remote branch has diverged,
that divergence is real work (someone — a person or another agent — pushed to
this branch), and overwriting it loses commits. Instead: fetch, inspect what's
there (git log HEAD..origin/HEAD --oneline), and integrate it by merging
origin/HEAD into your branch, resolving any conflicts exactly as in Phase 4.
Then push normally. If you can't reconcile it cleanly, stop and hand it to the
user — pushing over their work is never the answer.
Phase 7 — Open a PR (ready for review) if none exists
If Phase 1 found no PR, create one ready for review (not a draft):
--fill seeds the title/body from commits; improve the body if it's thin — a
reviewer should be able to tell what changed and why without reading the diff.
Capture the PR URL to report at the end.
Write the body short and plain. A PR/trail body is read by a busy human, not a compiler — optimize for fast understanding, not completeness. Aim for a screenful or less. Follow this shape:
- What was broken — the user-visible symptom, in one or two sentences, with a concrete example if there is one.
- Why — the root cause in plain language. Name mechanisms in words, not symbols; skip file paths, line numbers, function names, and constant names (the diff has those).
- The fix — what you changed, at a conceptual level. A short numbered list is fine.
- Verified — one line: what you ran (tests/lint) and any key behavioral proof.
Rules: no walls of text, no deep-dive rationale, no citing identifiersLikeThis
or path/to/file.go:123 in the body, no restating the same point three ways.
Keep the interesting-but-non-essential caveats out — if a reviewer needs them,
they'll ask or read the diff. Write like you're explaining it to a teammate in
Slack. If the drafted body is long or hard to follow, cut it down before opening
the PR.
Keep the trail out of the PR body. Do not put an entire-trail-link comment
block (or any trail link) in the GitHub PR body — the trail gets its own body in
the next step, not a link from the PR. If --fill or an existing body carried a
trail-link block over, strip it before settling the body.
If a PR already exists and is still a draft, mark it ready:
Write the Entire trail body
Do not copy the GitHub PR body onto the trail. The trail is read by reviewers and agents on Entire, and it should describe the branch as it is now, not as it was when the PR was opened. Write a fresh body from the current state of the branch: the diff against the base, the commits, what you verified, and what review has already said. Rewrite it whenever the branch changes materially (a fix pushed for CI, bugbot, or a trail finding), so it never goes stale.
Shape, in this order, a screenful or less:
- One-line summary — what this branch does, as a sentence a reviewer can repeat back.
- What changed — three to six bullets, conceptual, present tense. No file paths, line numbers, or identifiers; the diff has those.
- Why — one or two sentences on the problem or motivation. Link the ticket.
- Verified — what was run and passed (tests, lint, CI), and any behavioral proof worth naming.
- Review state — only when there is something to say: bugbot findings fixed or dismissed, trail findings resolved, known follow-ups deliberately left out. Omit the section when it would be empty.
Rules: plain language, one idea per bullet, no rationale essays, no duplicated points, no pasted PR text. If you find yourself copying a paragraph from the PR, rewrite it shorter.
Push it with the entire CLI (defaults to the current branch). Write the body to
a file first so apostrophes and backticks cannot garble the shell quoting, then
pass it through command substitution:
For a PR on a different branch, target it explicitly with --branch <headRefName>. If the update is refused because the description changed since
it was read, and you are deliberately replacing the whole body, add
--overwrite.
If entire trail update reports no trail for the branch, note it and move on —
don't create one just to set a body unless the user asks. If the entire CLI
isn't installed or the repo isn't Entire-enabled, skip this step silently.
Phase 8 — Make CI green
CI is the gate. Watch it, and fix what fails.
- Wait for checks to run and report status:
(Plain
gh pr checksfor a one-shot snapshot.) - For each failed check, get the actual failure — don't guess from the
check's name. Write the failing logs to a local file so you can read and
re-read them while iterating (this is the standing rule: debug from saved logs,
don't ask the user to paste them):
Then read that file and find the real failure (often buried above the final error). For richer context,
gh run view <run-id>lists the jobs and steps. - Diagnose and fix the root cause in the code, honoring the "never mask" principle above. A failing test means the code is wrong or the test needs a legitimate update because behavior intentionally changed — distinguish the two honestly. Reproduce locally first when you can; a fix you've confirmed locally is far more likely to turn the check green than a blind push.
- Commit the fix with a clear message, push, and return to step 1. Loop until all checks pass.
- Know when to stop. If a check fails for a reason outside this branch's control (infra flake, a secret the fork can't access, a failure already present on the base branch), don't thrash. State what's failing, why you believe it's unrelated, and hand it to the user.
Phase 9 — Bugbot review
With CI green, get an automated review pass from bugbot before handing off. Run this after CI is green so bugbot reviews the final state of the branch.
- Trigger bugbot by commenting on the PR:
Tell the user bugbot has been triggered and you're waiting for its review.
- Wait for the review. Bugbot typically takes ~5 minutes. Record the current
timestamp (ISO 8601) first, then get the repo slug:
Poll every 180 seconds (
sleep 180between checks) for new posts from bugbot. Its login iscursor[bot](older setups may show plaincursor) — match BOTH, and check BOTH post types (line comments AND reviews; a "no issues" verdict arrives only as a review body, so polling comments alone misses it):For a review with line comments, fetch them withgh api repos/{owner}/{repo}/pulls/{pr-number}/reviews/{review-id}/comments. If nothing appears after ~10 minutes, tell the user bugbot hasn't responded and ask how to proceed rather than waiting indefinitely. - Triage and fix each comment the same way the rest of this skill works — fix
the root cause, never mask it:
- Real bugs, security issues, correctness problems → fix them.
- Reasonable clarity/style improvements → fix them.
- False positives or suggestions that would make the code worse → skip them, and note why. Keep fixes minimal and focused on what bugbot flagged; don't refactor surrounding code or touch files bugbot didn't comment on. If a bugbot comment is ambiguous or represents a product decision, stop and ask the user.
- If you made fixes, verify locally (Phase 5), commit with a clear message, push (Phase 6), and return to Phase 8 — the fixes must pass CI too. Then you may re-trigger bugbot if the changes were substantial. If every comment was a false positive, say so and move on without committing.
If the repo doesn't use bugbot (commenting does nothing and no review ever arrives), don't block on it — note that bugbot isn't configured and finish.
Phase 10 — Check the trail slop meter and reduce slop
Entire runs a Slop monitor on the trail — a strict reviewer that rates the
branch's final diff small (Low), medium, or large for avoidable, low-trust
work: wrong assumptions about how the repo works, duplicate systems where an
established one existed, weak tests that match the code instead of proving
behavior, needless complexity, and over-commenting that narrates obvious control
flow. It judges the diff only — not the commit history, not how many rounds it
took.
Run this after bugbot's fixes have landed and CI is green, so the meter rates the final state. The runner triggers on push, so give it a moment after your last push before reading.
-
Read the slop rating. The monitor lives on the trail (key
slop). Fetch the trail's monitors — via the Entire MCP, the repo and trail number from the PR's branch:For the runner's full rationale (not just the value), fetch its latest run:
The trail is also viewable in the web UI if the API path is unavailable. If the repo has no
trail-sloprunner configured (check.entire/runners/trail-slop.json), or no trail exists, skip this phase and say so. If the run hasn't finished yet, wait and re-check (it takes a few minutes); don't block indefinitely — after ~10 minutes, report that the meter hasn't reported and move on. -
If the rating is
small(Low), note it and move on. Nothing to do. -
If the rating is
mediumorlarge, reduce the slop. This is not optional and it is not a style pass — the rationale names the two strongest reasons, and those are real maintenance costs a reviewer would otherwise pay. Read the rationale, then go fix what it names, at the root:- Wrong assumption / broken flow → trace the workflow end to end and correct the implementation, not the symptom.
- Duplicate system → use the repo's established way and delete the parallel one this branch added.
- Weak tests → strengthen them so they'd fail without the change (this is the Phase 2 bar; apply it again here).
- Needless complexity / oversized change → cut it down to what the result actually requires; delete scaffolding nothing uses.
- Over-commenting → remove comments this branch added that narrate obvious
control flow, translate code into English, or repeat nearby names, types, or
assertions. Keep comments that capture non-obvious intent, an invariant, an
external contract, a concurrency or ordering hazard, a security or
data-residency rule, a migration constraint, required exported-API docs, a
tool directive, or why a regression test exists. Don't touch pre-existing or
generated comments. If a
clean-commentsskill is available, it does exactly this pass.
-
Push the reduction and re-check. Verify locally (Phase 5), commit, push (Phase 6), return to Phase 8 for CI, then re-read the meter — the runner re-runs on push. Aim to land at
small. If it's stillmediumafter a genuine pass and you believe the remaining rationale is wrong (the complexity is necessary, the "duplicate" is deliberate), don't churn: say plainly what the meter flagged, what you fixed, what you disagree with and why, and hand the judgment to the user.
Treat the meter as a signal, not a boss: never game it by deleting tests, stripping legitimate comments, or shrinking the diff in ways that hurt the change. A lower rating bought that way is exactly the masking this skill forbids.
Phase 11 — Address reviews on the Entire trail
Reviewers (people and agents) leave review feedback on the branch's Entire trail as findings. Clear the open ones before handing off, the same root-cause way as the bugbot phase.
- List open findings for the branch's trail (defaults to the current branch,
--status open, current code version):If theentireCLI isn't installed, the repo isn't Entire-enabled, or there's no trail, skip this phase silently. If there are no open findings, note that and move on. Useentire trail finding show <finding-id>for the full detail of any finding. - Triage and fix each finding — fix the root cause, never mask it:
- Real bugs, security, correctness → fix them.
- Reasonable clarity/style improvements → fix them.
- False positives or suggestions that would make the code worse → dismiss with
a reason:
entire trail finding dismiss <finding-id> -m "why". When a finding carries a unified-diff suggestion that's correct, apply it directly (add--resolveto close it once the patch lands):
(--checkfirst if you want to confirm the patch applies cleanly.) When you fix a finding by hand instead, mark it resolved with a note:If a finding is ambiguous or represents a product decision, stop and ask the user rather than guessing. - If you made code fixes, verify locally (Phase 5), commit with a clear message, push (Phase 6), and return to Phase 8 — the fixes must pass CI too. Because the fixes changed the branch, rewrite the trail body (Phase 7) so it reflects the current state, including the findings you just resolved. If every finding was a false positive, dismiss them with reasons and move on without committing.
Done
Report concisely:
- Test coverage: which test covers each behavioral change, whether it already existed or you wrote it, and the red-then-green evidence — or the explicit reason no test was needed.
- Base merged in (and whether there were conflicts, and how you resolved any non-obvious ones).
- Local verification result.
- Diff self-review: any regressions caught and fixed, and any dead/unused code the branch introduced that you removed (name it) — or that the diff was clean.
- Any scratch planning/findings docs stripped before push (name each, and where you stashed a copy if one was worth keeping) — or that there were none.
- PR URL (and whether you created it or it already existed) — note it's ready for review (not a draft).
- Whether the Entire trail body was written fresh from the current branch state (and rewritten after later fixes), or why it was skipped.
- Final CI status: all green, or the specific checks still failing and why.
- Bugbot result: how many comments it raised, how many you fixed (and what), how many you skipped (and why) — or that bugbot isn't configured for the repo.
- Slop meter: the rating (
small/medium/large), and if it was medium or large, what you changed to reduce it and where it landed on the re-check — or that the repo has no slop runner / no trail. - Trail findings: how many open findings were on the trail, how many you resolved (and what), how many you dismissed (and why) — or that there was no trail / the repo isn't Entire-enabled.
If everything's green and the user wants to request review, point them at the
post-pr-review skill — that assigns the PR and posts it to Slack. This skill
deliberately stops short of that.
Blocked on one permission: opening the PR. Everything up to the push is done.
Done
- Test coverage: new
TestResolveDefaultSearchRepo(native, GitHub, mirror origins, explicit scope, no origin with and without explicit scope) and three forge cases inTestConfig_ScopeSlugs. Red with source reverted (build failure, the symbols are new), green with it restored. Behavioral red: the pre-fix binary exits 1 on the native clone where the fixed one exits 0. - Base: branch is 0 behind
origin/main, so no merge was needed. - Local verification:
mise run lint0 issues,mise run testpasses exceptTestOpenCodeSeedRepoPlantsDepsine2e/agents, which fails becauseopencode --versionprints nothing on this machine. The branch touches nothing undere2e. - Self-review: diff is clean. No dead code; the GitHub-only parser is still used by the two dispatch callers.
- Scratch docs:
FINDINGS.mdis excluded via.git/info/excludeand is not in the commit. It stays in the worktree. - Commit:
476493981pushed asorigin/native-search-repo-routing.
Blocked
gh pr createwas denied by the auto-mode classifier as an external write. The PR body is ready at/tmp/native-search-repo-routing/pr-body.md. To open it yourself:
- No trail exists yet for the branch, so the trail body (drafted at
/tmp/native-search-repo-routing/trail-body.md) could not be set. CI has not started on the branch either.
Once the PR exists, say so and I'll continue with the trail body, CI watch, bugbot, slop meter, and trail findings.
CI is running (watched in background), trail 1464 has 0 open findings so far, and bugbot runs as a PR check. I'll pick up the CI result, bugbot output, slop meter, and findings when the watch completes.
Finding resolved, trail body rewritten. CI on the new head is being watched in the background; I'll report when it finishes and then re-check the trail for any new findings on that head.