Fix Entire Search Native Repository Scope

Claude Code·Fable 5.1·alishakawaguchi·5d ago·54min·2 Checkpoints·5 file changes·+108/-24·54.9K tokens

Investigate the reported behavior: entire search errors on Entire-native repositories with "not a GitHub repository", even when the caller supplies an explicit --repo.

This is an investigation only. Do not implement a broad fix. Do not push, open or update a PR, merge, deploy, write to production, alter credentials, or change another agent/worktree. You may make local investigation notes in this isolated worktree and run routine read-only repository searches, exact CLI reproductions, and focused local tests. Preserve all unrelated files and sessions.

Verify first. Do not assume the report is still current. Before designing anything, confirm the problem is real and current: (1) Datadog logs/metrics for the reported window and the last 24h, if they materially distinguish server behavior from a local client rejection; keep any access read-only; (2) live cluster state via AWS SSO — kubectl / natsctl / entire api — and the deployed build, only if needed to distinguish supported backend capability from premature client validation, and read-only only; (3) the code path on current main as it is now, cited from this worktree; (4) Entire search (search_entire, search_entire_code, and the entire search skill) for prior checkpoints, sessions, trails, or a later fix on this symptom. If a source is not applicable because the failure occurs before any request leaves the client, say so explicitly and show the evidence. Write evidence and one explicit verdict — confirmed / already fixed / different than described — at the top of FINDINGS.md before any code changes. If already fixed or materially different, report that and stop; do not build the described fix.

Read AGENTS.md before work and follow its linked CLI conventions, API routing, testing, and git/subprocess guidance where relevant. This worktree is fresh from origin/main; do not switch branches or update main.

Investigation scope:

  • Reproduce the exact native-repository error with implicit repository discovery and with explicit --repo forms. Record the exact commands, environment assumptions, stdout/stderr, and exit status. Use a safe public/read-only native repo if one is already configured; do not create or mutate repositories.
  • Identify whether installed/released behavior differs from this worktree's current code. Record the exact CLI version/binary used for each reproduction.
  • Trace where local remote/host validation happens and determine whether it incorrectly overrides or rejects explicit repository selection.
  • Trace native repository identifier parsing and routing support end-to-end for semantic search, including accepted et/..., /et/..., and full Entire clone URL forms where relevant.
  • Identify the relevant tests and coverage gaps. Run focused safe tests that prove the current behavior; do not run paid real-agent E2E tests.
  • Separate two questions: whether semantic search for native repos is actually supported by the backend, and whether the CLI fails prematurely with GitHub-only validation before it could try the supported path.
  • Compare with relevant completed work such as ENT-2195, ENT-2196, ENT-2365 and PRs #2245/#2247, but do not assume those changes cover search.
  • Produce minimal fix options with tradeoffs and a recommended smallest change, but do not implement it unless explicitly asked in a later turn.

When root-caused, present the human-readable result using the architecture-findings skill: explain the players, normal request flow, where it breaks, whether backend capability exists, and where the minimal correction belongs, without file paths, line numbers, or symbol names in that narrative. Keep code-level references, commands, test names/results, and exact evidence in FINDINGS.md for follow-up.

When finished, stop and report the verdict, root cause, evidence, test results, and recommendation. Leave the Herdr session and worktree intact.

5d ago·4.5s

Using superpowers:systematic-debugging to structure the investigation, then search-test for Entire history.

Base directory for this skill: /Users/alisha/.claude/plugins/cache/claude-plugins-official/superpowers/6.4.1/skills/systematic-debugging

Systematic Debugging

Overview

Core principle: ALWAYS find root cause before attempting fixes. Symptom fixes are failure.

Violating the letter of this process is violating the spirit of debugging.

The Iron Law

If you haven't completed Phase 1, you cannot propose fixes.

When to Use

Use for ANY technical issue:

  • Test failures
  • Bugs in production
  • Unexpected behavior
  • Performance problems
  • Build failures
  • Integration issues

Use this ESPECIALLY when:

  • Under time pressure (emergencies make guessing tempting)
  • "Just one quick fix" seems obvious
  • You've already tried multiple fixes
  • Previous fix didn't work
  • You don't fully understand the issue

Don't skip when:

  • Issue seems simple (simple bugs have root causes too)
  • You're in a hurry (rushing guarantees rework)
  • Manager wants it fixed NOW (systematic is faster than thrashing)

The Four Phases

You MUST complete each phase before proceeding to the next.

Phase 1: Root Cause Investigation

BEFORE attempting ANY fix:

  1. Read Error Messages Carefully

    • Don't skip past errors or warnings
    • They often contain the exact solution
    • Read stack traces completely
    • Note line numbers, file paths, error codes
  2. Reproduce Consistently

    • Can you trigger it reliably?
    • What are the exact steps?
    • Does it happen every time?
    • If not reproducible → gather more data, don't guess
  3. Check Recent Changes

    • What changed that could cause this?
    • Git diff, recent commits
    • New dependencies, config changes
    • Environmental differences
  4. Gather Evidence in Multi-Component Systems

    WHEN system has multiple components (CI → build → signing, API → service → database):

    BEFORE proposing fixes, add diagnostic instrumentation:

    Example (multi-layer system):

    This reveals: Which layer fails (secrets → workflow ✓, workflow → build ✗)

  5. Trace Data Flow

    WHEN error is deep in call stack:

    See root-cause-tracing.md in this directory for the complete backward tracing technique.

    Quick version:

    • Where does bad value originate?
    • What called this with bad value?
    • Keep tracing up until you find the source
    • Fix at source, not at symptom

Phase 2: Pattern Analysis

Find the pattern before fixing:

  1. Find Working Examples

    • Locate similar working code in same codebase
    • What works that's similar to what's broken?
  2. Compare Against References

    • If implementing pattern, read reference implementation COMPLETELY
    • Don't skim - read every line
    • Understand the pattern fully before applying
  3. Identify Differences

    • What's different between working and broken?
    • List every difference, however small
    • Don't assume "that can't matter"
  4. Understand Dependencies

    • What other components does this need?
    • What settings, config, environment?
    • What assumptions does it make?

Phase 3: Hypothesis and Testing

Scientific method:

  1. Form Single Hypothesis

    • State clearly: "I think X is the root cause because Y"
    • Write it down
    • Be specific, not vague
  2. Test Minimally

    • Make the SMALLEST possible change to test hypothesis
    • One variable at a time
    • Don't fix multiple things at once
  3. Verify Before Continuing

    • Did it work? Yes → Phase 4
    • Didn't work? Form NEW hypothesis
    • DON'T add more fixes on top
  4. When You Don't Know

    • Say "I don't understand X"
    • Don't pretend to know
    • Ask for help
    • Research more

Phase 4: Implementation

Fix the root cause, not the symptom:

  1. Create Failing Test Case

    • Simplest possible reproduction
    • Automated test if possible
    • One-off test script if no framework
    • MUST have before fixing
    • Use the superpowers:test-driven-development skill for writing proper failing tests
  2. Implement Single Fix

    • Address the root cause identified
    • ONE change at a time
    • No "while I'm here" improvements
    • No bundled refactoring
  3. Verify Fix

    • Test passes now?
    • No other tests broken?
    • Issue actually resolved?
    • Use the superpowers:verification-before-completion skill before claiming success
  4. If Fix Doesn't Work

    • STOP
    • Count: How many fixes have you tried?
    • If < 3: Return to Phase 1, re-analyze with new information
    • If ≥ 3: STOP and question the architecture (step 5 below)
    • DON'T attempt Fix #4 without architectural discussion
  5. If 3+ Fixes Failed: Question Architecture

    Pattern indicating architectural problem:

    • Each fix reveals new shared state/coupling/problem in different place
    • Fixes require "massive refactoring" to implement
    • Each fix creates new symptoms elsewhere

    STOP and question fundamentals:

    • Is this pattern fundamentally sound?
    • Are we "sticking with it through sheer inertia"?
    • Should we refactor architecture vs. continue fixing symptoms?

    Discuss with your human partner before attempting more fixes

    This is NOT a failed hypothesis - this is a wrong architecture.

Red Flags - STOP and Follow Process

If you catch yourself thinking:

  • "Quick fix for now, investigate later"
  • "Just try changing X and see if it works"
  • "Add multiple changes, run tests"
  • "Skip the test, I'll manually verify"
  • "It's probably X, let me fix that"
  • "I don't fully understand but this might work"
  • "Pattern says X but I'll adapt it differently"
  • "Here are the main problems: [lists fixes without investigation]"
  • Proposing solutions before tracing data flow
  • "One more fix attempt" (when already tried 2+)
  • Each fix reveals new problem in different place

ALL of these mean: STOP. Return to Phase 1.

If 3+ fixes failed: Question the architecture (see Phase 4.5)

your human partner's Signals You're Doing It Wrong

Watch for these redirections:

  • "Is that not happening?" - You assumed without verifying
  • "Will it show us...?" - You should have added evidence gathering
  • "Stop guessing" - You're proposing fixes without understanding
  • "Ultra-think this" - Question fundamentals, not just symptoms
  • "We're stuck?" (frustrated) - Your approach isn't working

When you see these: STOP. Return to Phase 1.

Common Rationalizations

ExcuseReality
"Issue is simple, don't need process"Simple issues have root causes too. Process is fast for simple bugs.
"Emergency, no time for process"Systematic debugging is FASTER than guess-and-check thrashing.
"Just try this first, then investigate"First fix sets the pattern. Do it right from the start.
"I'll write test after confirming fix works"Untested fixes don't stick. Test first proves it.
"Multiple fixes at once saves time"Can't isolate what worked. Causes new bugs.
"Reference too long, I'll adapt the pattern"Partial understanding guarantees bugs. Read it completely.
"I see the problem, let me fix it"Seeing symptoms ≠ understanding root cause.
"One more fix attempt" (after 2+ failures)3+ failures = architectural problem. Question pattern, don't fix again.

Quick Reference

PhaseKey ActivitiesSuccess Criteria
1. Root CauseRead errors, reproduce, check changes, gather evidenceUnderstand WHAT and WHY
2. PatternFind working examples, compareIdentify differences
3. HypothesisForm theory, test minimallyConfirmed or new hypothesis
4. ImplementationCreate test, fix, verifyBug resolved, tests pass

When Process Reveals "No Root Cause"

If systematic investigation reveals issue is truly environmental, timing-dependent, or external:

  1. You've completed the process
  2. Document what you investigated
  3. Implement appropriate handling (retry, timeout, error message)
  4. Add monitoring/logging for future investigation

But: 95% of "no root cause" cases are incomplete investigation.

Supporting Techniques

These techniques are part of systematic debugging and available in this directory:

  • root-cause-tracing.md - Trace bugs backward through call stack to find original trigger
  • defense-in-depth.md - Add validation at multiple layers after finding root cause
  • condition-based-waiting.md - Replace arbitrary timeouts with condition polling
5d ago·5m

Tracing where the GitHub-only check sits relative to --repo, and how the semantic searcher uses owner/repo. Pulling the remaining code, history, and prior Entire work in parallel.

Continue the investigation. The previous command was denied only because it wrote under ~/.claude. Use /tmp/native-search-repo-routing/entire for the worktree binary instead, then complete the requested exact reproductions, focused tests, FINDINGS.md, architecture-findings narrative, and final verdict. Keep all access read-only except the /tmp binary and FINDINGS.md in this isolated worktree; do not implement a fix.

5d ago·9m

Owner/repo from the GitHub check feed only the default scope, nothing on the wire. Checking the secondary --code index miss, then writing FINDINGS.md and loading the narrative skill.

Base directory for this skill: /Users/alisha/.claude/skills/architecture-findings

Architecture findings

What this is for

After an investigation, there are two very different documents you could hand someone:

  1. A lookup table — every claim pinned to a file and line so an engineer can jump straight to the code. Precise, but dense and exhausting to read.
  2. A map of the territory — who the moving parts are, how they normally fit together, where the break is, and whose job it is to fix. This is the one that actually transfers understanding.

This skill produces the second kind. The reader should come away able to explain the system to someone else, not just able to find the line you're talking about.

The core rule

No code coordinates in the body. No file paths, no line numbers, no function or variable names, no code snippets. Those are the texture of the lookup table, and they're exactly what makes findings hard to read.

What you do name freely: the components, services, and systems by their real names (e.g. "the EU indexer", "entire-core", "the NATS bus"), the concepts ("ref events", "the repo-access token"), and the environment values that matter to the story only when named as a fact, not a citation ("the EU audience host is aws-eu-central-1.entire.io"). The test: would this detail help someone picture the system, or is it just a pointer for someone who wants to open a file? Keep the former, drop the latter.

You still did the rigorous, code-level investigation — this is just how you report it. If the user later asks "where exactly", give them the coordinates then. Don't pre-empt that with a wall of citations.

Structure

Adapt to the investigation, but this arc almost always works:

  • Bottom line — 2–4 sentences. What's happening, why, and the one thing that matters most (e.g. "the fix isn't where you'd expect — it's upstream"). Someone who reads only this should get the headline right.
  • The players — introduce each component in one line: what it is and what it's responsible for. This is the cast list; everything after refers back to it.
  • The normal flow — how it works when it works. This is the reference the reader measures the break against. A simple arrow-chain diagram earns its place here (see below).
  • Why it breaks — the failure, told as a story against that normal flow. Show the point where reality diverges from the happy path.
  • The structural gap (when there is one) — the deeper "why can't we just fix it here" — the missing capability or wrong assumption underneath the symptom. This is often the most valuable section: it's what stops the reader proposing the naive fix.
  • Where the fix belongs — which system owns it and why, what the options are, and explicitly what the wrong fix is if there's a tempting one.
  • How other cases differ (when relevant) — the comparison that proves it's structural, not a one-off (other regions, other tenants, the working vs. broken side).

Drop sections that don't apply. A small investigation might be just bottom line + why it breaks + where the fix belongs.

Flow diagrams

Simple ASCII arrow chains are encouraged where a flow clarifies the story — they carry sequence better than a paragraph. Keep them component-level, never code-level. For example:

Then a second chain showing where the broken path diverges is often the clearest way to show a failure. Label them ("normal flow" / "what EU does"). Don't force a diagram where prose is already clear.

Voice

Write for a smart colleague who doesn't live in this code. Explain the why behind each mechanism, not just the what — "the STS refuses because that repo isn't an EU repo; that refusal is the residency guard doing its job" teaches more than "the STS returns an error". Prefer plain language over jargon, and when a term is load-bearing, define it once in passing. Keep it tight — this is a briefing, not an essay.

What to avoid

  • File paths, line numbers, symbol names, code blocks (other than the component-level flow diagrams above).
  • Turning "the players" into an inventory of every file touched.
  • Hedging everything. You investigated; state what you found plainly, and flag genuine open questions as open questions.
  • Burying the headline. The bottom line goes first.
4d ago·2m

Verdict: confirmed. Reproduced on both the worktree build of main and the installed nightly. Not fixed anywhere. Details, commands, versions, and tests are in FINDINGS.md in this worktree.

Bottom line

entire search dies inside the CLI before it talks to anything. It reads the clone's origin, insists the origin be GitHub-backed, and only then would look at --repo or --all-repos. An Entire-native clone has an origin that names an Entire cluster and the et forge, so the check fails every time, whatever flags you pass. The backend is not the problem: the identical binary, launched from a GitHub clone with --repo et/entirehq/entire-api, returns native-repo results from production. The fix belongs in the search command's scope resolution, not in the backend and not in the shared GitHub parser.

The players

  • The search command in the CLI: turns the query and flags into a scope and hands it to the semantic searcher.
  • The remote parser: understands plain GitHub URLs and Entire clone URLs, and reports which forge backs a clone, gh or et.
  • The GitHub-only check: a thin policy on top of the parser that the search command, and the dispatch commands, use to insist on a GitHub origin.
  • entire-core, the control plane: holds the repo index, where native repos are stored under forge-prefixed names, and resolves a name to a repo ID and home cell.
  • query-serve in each cell: the semantic index that answers scoped searches by repo ID.

The normal flow

When no --repo is given, the scope is the current repo, derived from origin. Owner and repo from origin are used only to build that default name. They never reach the wire.

Why it breaks

The command runs the GitHub-only check on origin unconditionally, before it decides whether it even needs a default scope. For a native clone the parser correctly says "forge is et, no GitHub upstream", the check reports the cluster host as "not GitHub", and the command exits. Explicit --repo et/..., --repo gh/..., and --all-repos are all discarded because the error fires first. Two other input shapes, a leading-slash /et/... and a full clone URL, fail even earlier in filter validation, which only accepts the bare prefixed form.

The structural gap

The check was loosened in June to accept GitHub mirrors reached through an Entire clone URL, but kept its GitHub-only policy on purpose. When native repos were taught to the trail commands and to checkpoint explain in September, the search command's help text was updated to advertise native repos, but its origin handling was not. The code-search branch of the same command already does the right thing: it builds a forge-prefixed default name and never calls the GitHub check. The semantic branch simply predates that pattern.

One trap: skipping the check and keeping the bare owner/repo default would be wrong too. entire-core treats a bare pair as "either forge", so a native repo and a same-named GitHub mirror would both match. The default must carry the forge prefix, as the control plane and the index already do.

Where the fix belongs

The search command, at the point where it derives the default scope. Parse origin with the forge-aware parser, seed the default scope with the forge-prefixed name, and fail on an unparseable origin only when no explicit scope was supplied, mirroring the code-search branch. That one change fixes implicit, --repo, and --all-repos in native clones. Do not widen the shared GitHub-only check or teach the parser that et maps to a host: dispatch relies on that check staying strict.

Optional follow-up: accept /et/... and full clone URLs as filters, as checkpoint explain already does.

Remaining limitation

entire search --code against a native repo reaches the backend but reports no matching repository, from both native and GitHub clones. Code search matches filters against the unfiltered repo listing, while semantic search uses the server's exact-name filter. Not investigated further per your instruction; noted in FINDINGS.md as a separate ticket candidate.

Worktree, Herdr session, and all other files are untouched. The only artifacts are FINDINGS.md here and the binary at /tmp/native-search-repo-routing/entire.

Base directory for this skill: /Users/alisha/.claude/skills/rubber-duck

Rubber Duck

The user is looking at something dense — usually a long agent response or a findings dump — and wants to understand it, not act on it. Your whole job is to be the clear, patient explainer and then a sounding board for their questions. Think out loud with them, in plain language.

The one job

Turn the wall of text into something a smart person with no context can follow in under a minute, then answer whatever they ask next — still briefly.

Hard rules (do not break these)

  • Do NOT write a plan, and do NOT enter plan mode. No ExitPlanMode, no writing-plans/brainstorming skills, no "here's how we'd fix it" roadmaps.
  • Do NOT change code, run builds, or edit files. This is understanding only.
  • Do NOT make a todo list. It's a conversation, not a task.
  • Stay read-only. Reading a file or scrolling the text to explain it is fine. Anything that does something is not.
  • Be concise — this is the point of the skill. You are the antidote to the wall of text, so don't produce another one. Lead with the answer. Short sentences. No preamble, no "great question", no restating the question back.
  • Plain language over jargon. No file paths, line numbers, symbol names, or code snippets unless the user asks for them. Names of systems/people are fine; citations are not. (This is the same spirit as the architecture-findings skill — the story, not the lookup table.)

First explanation — the format

Keep it to roughly this, and stop:

  1. One-line TL;DR. The single most important thing, first.
  2. 3–5 short bullets covering: what's actually going on, why it matters, and (if relevant) what's broken or being decided. One idea per bullet.
  3. One line on "so what" — the upshot, or the decision in front of them.

If the source genuinely has more than fits, say "there's more detail under X if you want it" and let them pull it — don't push it.

Use an analogy when it makes a hard idea click. Skip it when it doesn't.

The follow-up loop

After the first explanation, they'll ask things. For each:

  • Answer the specific question and nothing more. One concept at a time.
  • Stay just as concise — don't let answers creep back into walls of text.
  • It's fine to say "not sure / the text doesn't say" rather than guess.
  • If a question needs you to look something up to answer honestly, look (read only), then answer plainly.

Keep going as long as they're asking. You're the duck — they're thinking out loud and you're helping them see it.

When they're ready to act

The moment the user shifts from "help me understand" to "okay let's fix/build/ plan this," rubber-duck mode is over. Don't quietly start doing it inside this mode. Say something like: "That's action territory — want me to pick up that work?" and let them start the real task (a plan, a dispatch, an edit) as its own thing. Your job here ends at understanding.

4d ago·8.7s

TL;DR: The search command checks "is your origin GitHub?" before it looks at anything else, so native-repo clones always fail locally. The backend is fine.

  • What happens: In a native clone, origin points at an Entire cluster with an et path. The command reads that, says "not GitHub", and exits. Your --repo or --all-repos is never read because the error fires first.
  • Proof the backend works: Same binary, run from a GitHub clone with --repo et/entirehq/entire-api, returns native-repo results from production.
  • Why it's there: That GitHub check was built when search only knew GitHub. In September native repos were added to trails and checkpoint explain, and the search help text was updated, but this one check was missed.
  • The trap: You can't just delete the check. Without the et/ prefix, the default scope would be a bare "owner/repo", which the control plane treats as "either forge," so a native repo and a same-named GitHub mirror would both match.
  • Already-fixed sibling: The --code branch of the same command does it right: it builds the prefixed name and never runs the GitHub check.

So what: It's a small fix in one place, the search command's default-scope logic, mirroring what --code does. Nothing to change in the backend or the shared GitHub parser (dispatch still needs that strict).

There's a side note in FINDINGS.md about --code failing differently on native repos if you want it.

fix it same way the code one did

4d ago·6m

Fixed, mirroring the --code branch. Uncommitted in this worktree.

Change

  • cmd/entire/cli/search_cmd.go: the GitHub-only origin check is gone. Semantic search now derives the default scope from origin's forge-qualified coordinates via a new resolveDefaultSearchRepo helper. With an explicit --repo, repo:, or --all-repos, an unreadable origin is ignored. Without one, it errors with the same hint --code uses: "could not determine current repository for search (use --repo or --all-repos)".
  • cmd/entire/cli/search/search.go: Config.Forge added; the default slug is now et/proj/repo or gh/owner/repo, so a native repo and a same-named mirror can't be conflated.
  • ParseGitHubRemote is untouched; dispatch still relies on it.

Tests (written first, failed, then passed)

  • New TestResolveDefaultSearchRepo: native, GitHub, mirror origins, explicit-scope, and no-origin with and without explicit scope.
  • Three new ScopeSlugs cases for forge-prefixed defaults.
  • Full cmd/entire/cli, search, and gitremote packages pass. mise run fmt clean, mise run lint 0 issues.

End-to-end with the rebuilt binary from ~/Projects/entire-api: implicit, --repo et/..., --all-repos, and --repo gh/entireio/cli all exit 0 with correct hits. Implicit from a GitHub clone still works. Details appended to FINDINGS.md.

Not done: mise run check (integration + canary) and the commit. --code on native repos still reports no matching repo, which is the separate index issue noted in FINDINGS.md.

did you test that the binary worked and you cna reach native repos

4d ago·6.8s

Yes. I rebuilt the binary from the fixed tree and ran it against production from the native clone at ~/Projects/entire-api (origin entire://aws-us-east-2.entire.io/et/entirehq/entire-api):

Command (from the native clone)BeforeAfter
search 'routing' --json --compact --limit 1exit 1, "not a GitHub repository"exit 0, commit hit in entirehq/entire-api
--repo et/entirehq/entire-apiexit 1exit 0, same native hit
--all-reposexit 1exit 0, hit in entirehq/company-knowledge (native)
--repo gh/entireio/cliexit 1exit 0, hit in entireio/cli

Also checked from /tmp (no git repo): implicit gives the --code-style hint, --repo et/entirehq/entire-api returns results. And from this GitHub worktree, implicit still resolves to entireio/cli.

All read-only searches. Exact output is in the "Fix applied" section of FINDINGS.md.

Base directory for this skill: /Users/alisha/.claude/skills/pr-ready

Getting a PR Ready for Review

The goal is a branch that a reviewer can look at with confidence: covered by a real test, caught up with the base branch, no conflicts, no regressions, a PR open (ready for review), and CI green. This skill is the step before requesting review — it does not merge the PR and does not post it anywhere.

Work through the phases below in order, but scope to what's needed. If the branch is already merged-clean and only CI is red, skip straight to the CI phase. The one phase you never skip is the test-coverage check (Phase 2) — a branch with no test for what it changed isn't review-ready no matter how green CI is. If there's no remote branch yet, you'll still do everything. Tell the user which phases you're running and why.

Guiding principles

These matter more than the mechanics — internalize them, because the mechanics serve them:

  • Never mask a problem to make a check pass. Deleting a failing test, loosening an assertion, adding // eslint-disable over a real lint error, --no-verify, or continue-on-error turns a red check green while leaving the bug in place. That defeats the entire point of getting review-ready. Fix the underlying cause. If a check is failing for a reason genuinely unrelated to this branch (flaky test, pre-existing breakage on main), say so explicitly and ask the user rather than papering over it.
  • Conflict resolution preserves both intents. A merge conflict means two changes touched the same place. Resolving it is not "pick a side" — it's understanding what each side was trying to do and producing code that honors both. Blindly taking --ours or --theirs is how regressions sneak in.
  • A change without a test that would fail without it isn't done. The test is what stops this bug or feature from silently regressing in six months when someone refactors past it. A test that passes equally well against the pre-change code covers nothing, however thorough it looks — prove the red before you trust the green.
  • Verify before you push, then verify again in CI. Local verification is fast and catches most regressions before they cost a CI cycle. CI is the source of truth. Both matter.
  • Review your own diff before you ship it. Build/test/lint prove the code runs; they don't prove the change is clean. Read the full branch diff with fresh eyes for two things a green build hides: regressions the change slipped in (behavior quietly altered, a caller left on an old contract, a dropped edge case) and dead or unused code it created (functions/params/imports/types nothing references, unreachable branches, leftover debug or commented-out code, scaffolding wired to nothing). Ship the change, not the scaffolding around it.
  • Stop and ask when the right answer is ambiguous, especially on conflict resolution where the two sides represent real product decisions. A wrong guess here is worse than a question.

Phase 1 — Orient

Establish the ground truth before changing anything. Run these read-only:

  • git rev-parse --abbrev-ref HEAD — the current branch. If it's main/master, stop: there's nothing to ready-for-review. Ask the user which branch they mean.
  • git status — confirm a clean working tree. If there are uncommitted changes, surface them and ask whether to commit, stash, or abort. Don't merge on top of a dirty tree — it tangles your changes with the merge.
  • Determine the base branch. It's usually main, sometimes master or a release branch. Check git remote show origin | grep "HEAD branch" or look at what the existing PR targets (Phase 5). Don't assume main.
  • gh pr view --json number,state,isDraft,baseRefName,url 2>/dev/null — does a PR already exist for this branch, and what does it target? Remember the answer for Phase 5.

Report a one-line plan: which base, whether a PR exists, and which phases you'll run.

Phase 2 — Prove a real test covers the change

Do this before touching the branch. A change that ships without a test that exercises it will regress — someone will refactor past it, and nothing will notice. This is the first substantive step, ahead of the merge, so a missing test is written against the branch's own code rather than bolted on after a conflict resolution has muddied it.

Read the branch diff and identify what it actually changed:

For each behavioral change the branch makes (a bug fixed, a feature added, a rule or edge case altered), find the test that covers it. Look in the diff first (--diff-filter=AM over the repo's test paths), then in existing test files near the changed code — an existing test extended with a new case counts.

A test only counts as real coverage if all of these hold:

  • It would fail without the change. This is the whole bar. For a bugfix, the test reproduces the original symptom; for a feature, it asserts the new behavior. If it passes equally well against the pre-change code, it covers nothing.
  • It asserts behavior, not shape. Checking that a function was called, that a struct has a field, or that a string matches the new copy is not coverage. Assert the outcome a user or caller would observe.
  • It goes through the real path. A test that mocks out the very thing the change fixed proves only that the mock works. Exercise the actual code under test; mock at the system boundary, not around the change.
  • It covers the edge case the bug lived in, not just the happy path. Bugs live in the empty list, the nil, the concurrent second call, the expired token.

Verify the "would fail without it" claim — don't assert it from reading. Cheapest reliable check: stash or revert the source change (not the test) and run the test, confirm it fails, then restore and confirm it passes.

If stashing is awkward (a large or tangled diff), invert the key assertion or temporarily neuter the fixed line instead — anything that gives you a real red before the green. Record what you saw; "the test passes" alone is not evidence.

If coverage is missing or weak, write the test now — before the merge, before the push. Follow the repo's existing test conventions (framework, fixtures, naming, table-driven style) rather than inventing a new pattern. Put it at the level that makes the guarantee durable: prefer a unit test close to the logic when that genuinely exercises the bug, and reach for an integration/e2e test when the bug only exists in the seam between components.

Then run the new test and confirm it fails against the unfixed code, per above.

Legitimate exceptions — state them explicitly rather than skipping silently:

  • Pure refactors, renames, formatting, dependency bumps, or comment/doc-only changes that alter no behavior — the existing suite passing is the coverage.
  • Changes to code the repo has no practical way to test (generated files, infra/config, a thin shell over an external API with no harness). Say so, and say what manual verification you did instead.
  • The user explicitly waives it.

If you believe an exception applies but aren't sure, ask — don't decide for the user that a change is untestable. If a test is genuinely needed but writing it requires a product decision (what should the behavior be at the edge?), stop and ask.

Report, in one or two lines: which test covers each behavioral change, whether it already existed or you wrote it, and the red-then-green evidence you observed.

Phase 3 — Sync the base branch

Get the latest base so the merge reflects reality:

Fetching is enough — you merge origin/<base> directly in the next phase, so you don't need to check out and fast-forward your local base branch. (If the user specifically wants their local base updated too, git checkout <base> && git pull --ff-only && git checkout - does it without surprises.)

Phase 4 — Merge the base in and resolve conflicts

The user asked for a merge, not a rebase — keep it a merge so history and any in-flight review threads stay intact:

If it merges cleanly, go to Phase 5 after a quick local verification.

If there are conflicts, resolve them deliberately:

  1. git status and git diff to see every conflicted file. Work through them one at a time — don't bulk-resolve.
  2. For each conflict, understand both sides before writing the resolution:
    • git log --merge -p <file> shows the commits from each side touching the conflict region — this is the fastest way to see intent.
    • If the repo is Entire-enabled and the intent is still unclear, the what-happened / explain skills can recover the reasoning behind a block.
    • Read enough surrounding code to know what each change was accomplishing.
  3. Write a resolution that keeps both intents. If side A renamed a function and side B added a new caller of the old name, the resolution updates the caller — it doesn't drop either change. If the two intents genuinely conflict (both changed the same value to different things), that's a product decision: stop and ask the user.
  4. git add each resolved file as you finish it.
  5. When all conflicts are resolved, verify before committing the merge (next phase). Don't git commit the merge until the code at least builds — a merge commit that doesn't compile is a regression you're about to push.

Phase 5 — Verify locally & self-review the diff

A merge can produce code that compiles on each side but breaks when combined ("semantic conflicts" git can't detect). Catch these now, cheaply.

Run the project's real verification — discover it rather than guessing. Look for the commands the repo actually uses (a project skill for running the app/tests, Makefile/justfile targets, package.json scripts, go test ./..., CI config under .github/workflows, CONTRIBUTING.md). Run, in roughly this order, what exists: build → unit tests → lint/typecheck.

cli repo — do NOT run the e2e suite locally (temporary, 2026-07-20). The e2e suite (mise run test:e2e*, and test:ci/check which call it) deletes the developer's real ~/.claude/projects transcripts via an unmerged HOME- isolation bug. The tasks are locally gated to SKIP (exit 0) — that skip is intentional; do not "fix" it or set ENTIRE_ALLOW_LOCAL_E2E=1. Verify with mise run test (unit) + mise run lint only; CI runs e2e on isolated runners.

  • If verification passes, commit the merge (git commit with no -m accepts the default merge message, or write a short one) and continue.
  • If it fails, fix the cause — a semantic conflict you missed, an import that moved, a renamed symbol. Re-run until clean, then commit the merge.

Skipping this phase just moves the failure to CI, where it's slower to diagnose. It's worth doing even when the merge was conflict-free.

Self-review the branch diff — regressions & dead code

Verification proves the code runs; it doesn't prove the change is clean. Before pushing, read the branch's full diff against the base with fresh eyes — this is a review of the code this branch introduces, separate from the build/test pass above:

Look for two classes of problem a green build won't reliably flag:

  • Regressions the change introduced — behavior that changed when it shouldn't have; a renamed or re-signatured function with a caller left on the old contract; an error path, guard, or edge case dropped in an edit or a conflict resolution; a default or constant that silently shifted. Trace each non-trivial hunk to the callers it affects and confirm they still hold.
  • Dead / unused code the change created — new functions, methods, variables, parameters, struct fields, types, imports, or constants that nothing references; branches that can no longer be reached; a helper added but never wired in; a feature-flag path with no live caller; leftover debug logging, commented-out code, or TODO stubs. Review naturally sees the addition and rarely notices the thing is never used.

Lean on tooling to find the unused set, then confirm by hand — don't trust it alone:

  • Compiler/linter unused checks (discover what the repo uses): Go flags unused locals/imports at build and surfaces unused functions via go vet / staticcheck / golangci-lint (and golang.org/x/tools/cmd/deadcode for whole-program dead functions); TypeScript via tsc --noUnusedLocals --noUnusedParameters and eslint no-unused-vars.
  • For each new exported symbol the diff adds, grep the repo for references — zero hits outside its own definition and test means it's dead.
  • If a dead-code-finder skill is available, point it at the branch's changes.

Resolve what you find: remove dead code this branch introduced (if it's deliberate scaffolding for an imminent follow-up, say so explicitly and keep it only with the user's nod); fix any regression at its root, never by loosening a test. If deleting suspected-dead code is risky — it may be reached via reflection, build tags, code generation, or an external consumer — confirm before removing it rather than guessing. Keep the scope to this branch's diff; don't go hunting pre-existing dead code across the repo — that's not what getting this PR ready is about.

Phase 6 — Push

Before you push, strip scratch working docs from the branch. Agents routinely leave planning/findings scratch files in the worktree — FINDINGS.md, PLAN.md, plan.md, *-seed.md, DAY_DRIVER_TASK.md, NOTES.md, TODO.md, design/ investigation write-ups — that are working artifacts, not part of the change. They must never ship in the PR. Catch them here, before the branch goes remote.

  1. List .md files this branch adds vs the base, plus any untracked ones in the worktree:
  2. Decide each candidate by intent, not just extension:
    • Remove agent-generated scratch — findings / plan / seed / task-handoff / notes docs (typically at the repo root or a scratch dir) created by this branch and unrelated to the change under review.
    • Keep genuine documentation — anything under docs/, README*, CHANGELOG*, ADRs, CONTRIBUTING*, or a .md the change legitimately adds (e.g. docs for the feature you just built). When you're unsure whether a doc is intended, ask rather than delete.
  3. Remove the scratch docs so they can't reach the PR — but don't destroy work you can't recover: if a findings/investigation write-up looks worth keeping, copy it to a durable scratch location (~/.claude/tmp/…) first, then delete.
    • Untracked (never committed): delete from disk (rm <file>).
    • Already committed on this branch: git rm <file> and commit the removal (git commit -m "chore: drop scratch planning/findings docs").

Then push:

Never force-push — not --force, not --force-with-lease, not under any circumstance. If a plain push is rejected because the remote branch has diverged, that divergence is real work (someone — a person or another agent — pushed to this branch), and overwriting it loses commits. Instead: fetch, inspect what's there (git log HEAD..origin/HEAD --oneline), and integrate it by merging origin/HEAD into your branch, resolving any conflicts exactly as in Phase 4. Then push normally. If you can't reconcile it cleanly, stop and hand it to the user — pushing over their work is never the answer.

Phase 7 — Open a PR (ready for review) if none exists

If Phase 1 found no PR, create one ready for review (not a draft):

--fill seeds the title/body from commits; improve the body if it's thin — a reviewer should be able to tell what changed and why without reading the diff. Capture the PR URL to report at the end.

Write the body short and plain. A PR/trail body is read by a busy human, not a compiler — optimize for fast understanding, not completeness. Aim for a screenful or less. Follow this shape:

  • What was broken — the user-visible symptom, in one or two sentences, with a concrete example if there is one.
  • Why — the root cause in plain language. Name mechanisms in words, not symbols; skip file paths, line numbers, function names, and constant names (the diff has those).
  • The fix — what you changed, at a conceptual level. A short numbered list is fine.
  • Verified — one line: what you ran (tests/lint) and any key behavioral proof.

Rules: no walls of text, no deep-dive rationale, no citing identifiersLikeThis or path/to/file.go:123 in the body, no restating the same point three ways. Keep the interesting-but-non-essential caveats out — if a reviewer needs them, they'll ask or read the diff. Write like you're explaining it to a teammate in Slack. If the drafted body is long or hard to follow, cut it down before opening the PR.

Keep the trail out of the PR body. Do not put an entire-trail-link comment block (or any trail link) in the GitHub PR body — the trail gets its own body in the next step, not a link from the PR. If --fill or an existing body carried a trail-link block over, strip it before settling the body.

If a PR already exists and is still a draft, mark it ready:

Write the Entire trail body

Do not copy the GitHub PR body onto the trail. The trail is read by reviewers and agents on Entire, and it should describe the branch as it is now, not as it was when the PR was opened. Write a fresh body from the current state of the branch: the diff against the base, the commits, what you verified, and what review has already said. Rewrite it whenever the branch changes materially (a fix pushed for CI, bugbot, or a trail finding), so it never goes stale.

Shape, in this order, a screenful or less:

  • One-line summary — what this branch does, as a sentence a reviewer can repeat back.
  • What changed — three to six bullets, conceptual, present tense. No file paths, line numbers, or identifiers; the diff has those.
  • Why — one or two sentences on the problem or motivation. Link the ticket.
  • Verified — what was run and passed (tests, lint, CI), and any behavioral proof worth naming.
  • Review state — only when there is something to say: bugbot findings fixed or dismissed, trail findings resolved, known follow-ups deliberately left out. Omit the section when it would be empty.

Rules: plain language, one idea per bullet, no rationale essays, no duplicated points, no pasted PR text. If you find yourself copying a paragraph from the PR, rewrite it shorter.

Push it with the entire CLI (defaults to the current branch). Write the body to a file first so apostrophes and backticks cannot garble the shell quoting, then pass it through command substitution:

For a PR on a different branch, target it explicitly with --branch <headRefName>. If the update is refused because the description changed since it was read, and you are deliberately replacing the whole body, add --overwrite.

If entire trail update reports no trail for the branch, note it and move on — don't create one just to set a body unless the user asks. If the entire CLI isn't installed or the repo isn't Entire-enabled, skip this step silently.

Phase 8 — Make CI green

CI is the gate. Watch it, and fix what fails.

  1. Wait for checks to run and report status:
    (Plain gh pr checks for a one-shot snapshot.)
  2. For each failed check, get the actual failure — don't guess from the check's name. Write the failing logs to a local file so you can read and re-read them while iterating (this is the standing rule: debug from saved logs, don't ask the user to paste them):
    Then read that file and find the real failure (often buried above the final error). For richer context, gh run view <run-id> lists the jobs and steps.
  3. Diagnose and fix the root cause in the code, honoring the "never mask" principle above. A failing test means the code is wrong or the test needs a legitimate update because behavior intentionally changed — distinguish the two honestly. Reproduce locally first when you can; a fix you've confirmed locally is far more likely to turn the check green than a blind push.
  4. Commit the fix with a clear message, push, and return to step 1. Loop until all checks pass.
  5. Know when to stop. If a check fails for a reason outside this branch's control (infra flake, a secret the fork can't access, a failure already present on the base branch), don't thrash. State what's failing, why you believe it's unrelated, and hand it to the user.

Phase 9 — Bugbot review

With CI green, get an automated review pass from bugbot before handing off. Run this after CI is green so bugbot reviews the final state of the branch.

  1. Trigger bugbot by commenting on the PR:
    Tell the user bugbot has been triggered and you're waiting for its review.
  2. Wait for the review. Bugbot typically takes ~5 minutes. Record the current timestamp (ISO 8601) first, then get the repo slug:
    Poll every 180 seconds (sleep 180 between checks) for new posts from bugbot. Its login is cursor[bot] (older setups may show plain cursor) — match BOTH, and check BOTH post types (line comments AND reviews; a "no issues" verdict arrives only as a review body, so polling comments alone misses it):
    For a review with line comments, fetch them with gh api repos/{owner}/{repo}/pulls/{pr-number}/reviews/{review-id}/comments. If nothing appears after ~10 minutes, tell the user bugbot hasn't responded and ask how to proceed rather than waiting indefinitely.
  3. Triage and fix each comment the same way the rest of this skill works — fix the root cause, never mask it:
    • Real bugs, security issues, correctness problems → fix them.
    • Reasonable clarity/style improvements → fix them.
    • False positives or suggestions that would make the code worse → skip them, and note why. Keep fixes minimal and focused on what bugbot flagged; don't refactor surrounding code or touch files bugbot didn't comment on. If a bugbot comment is ambiguous or represents a product decision, stop and ask the user.
  4. If you made fixes, verify locally (Phase 5), commit with a clear message, push (Phase 6), and return to Phase 8 — the fixes must pass CI too. Then you may re-trigger bugbot if the changes were substantial. If every comment was a false positive, say so and move on without committing.

If the repo doesn't use bugbot (commenting does nothing and no review ever arrives), don't block on it — note that bugbot isn't configured and finish.

Phase 10 — Check the trail slop meter and reduce slop

Entire runs a Slop monitor on the trail — a strict reviewer that rates the branch's final diff small (Low), medium, or large for avoidable, low-trust work: wrong assumptions about how the repo works, duplicate systems where an established one existed, weak tests that match the code instead of proving behavior, needless complexity, and over-commenting that narrates obvious control flow. It judges the diff only — not the commit history, not how many rounds it took.

Run this after bugbot's fixes have landed and CI is green, so the meter rates the final state. The runner triggers on push, so give it a moment after your last push before reading.

  1. Read the slop rating. The monitor lives on the trail (key slop). Fetch the trail's monitors — via the Entire MCP, the repo and trail number from the PR's branch:

    For the runner's full rationale (not just the value), fetch its latest run:

    The trail is also viewable in the web UI if the API path is unavailable. If the repo has no trail-slop runner configured (check .entire/runners/trail-slop.json), or no trail exists, skip this phase and say so. If the run hasn't finished yet, wait and re-check (it takes a few minutes); don't block indefinitely — after ~10 minutes, report that the meter hasn't reported and move on.

  2. If the rating is small (Low), note it and move on. Nothing to do.

  3. If the rating is medium or large, reduce the slop. This is not optional and it is not a style pass — the rationale names the two strongest reasons, and those are real maintenance costs a reviewer would otherwise pay. Read the rationale, then go fix what it names, at the root:

    • Wrong assumption / broken flow → trace the workflow end to end and correct the implementation, not the symptom.
    • Duplicate system → use the repo's established way and delete the parallel one this branch added.
    • Weak tests → strengthen them so they'd fail without the change (this is the Phase 2 bar; apply it again here).
    • Needless complexity / oversized change → cut it down to what the result actually requires; delete scaffolding nothing uses.
    • Over-commenting → remove comments this branch added that narrate obvious control flow, translate code into English, or repeat nearby names, types, or assertions. Keep comments that capture non-obvious intent, an invariant, an external contract, a concurrency or ordering hazard, a security or data-residency rule, a migration constraint, required exported-API docs, a tool directive, or why a regression test exists. Don't touch pre-existing or generated comments. If a clean-comments skill is available, it does exactly this pass.
  4. Push the reduction and re-check. Verify locally (Phase 5), commit, push (Phase 6), return to Phase 8 for CI, then re-read the meter — the runner re-runs on push. Aim to land at small. If it's still medium after a genuine pass and you believe the remaining rationale is wrong (the complexity is necessary, the "duplicate" is deliberate), don't churn: say plainly what the meter flagged, what you fixed, what you disagree with and why, and hand the judgment to the user.

Treat the meter as a signal, not a boss: never game it by deleting tests, stripping legitimate comments, or shrinking the diff in ways that hurt the change. A lower rating bought that way is exactly the masking this skill forbids.

Phase 11 — Address reviews on the Entire trail

Reviewers (people and agents) leave review feedback on the branch's Entire trail as findings. Clear the open ones before handing off, the same root-cause way as the bugbot phase.

  1. List open findings for the branch's trail (defaults to the current branch, --status open, current code version):
    If the entire CLI isn't installed, the repo isn't Entire-enabled, or there's no trail, skip this phase silently. If there are no open findings, note that and move on. Use entire trail finding show <finding-id> for the full detail of any finding.
  2. Triage and fix each finding — fix the root cause, never mask it:
    • Real bugs, security, correctness → fix them.
    • Reasonable clarity/style improvements → fix them.
    • False positives or suggestions that would make the code worse → dismiss with a reason: entire trail finding dismiss <finding-id> -m "why". When a finding carries a unified-diff suggestion that's correct, apply it directly (add --resolve to close it once the patch lands):
    (--check first if you want to confirm the patch applies cleanly.) When you fix a finding by hand instead, mark it resolved with a note:
    If a finding is ambiguous or represents a product decision, stop and ask the user rather than guessing.
  3. If you made code fixes, verify locally (Phase 5), commit with a clear message, push (Phase 6), and return to Phase 8 — the fixes must pass CI too. Because the fixes changed the branch, rewrite the trail body (Phase 7) so it reflects the current state, including the findings you just resolved. If every finding was a false positive, dismiss them with reasons and move on without committing.

Done

Report concisely:

  • Test coverage: which test covers each behavioral change, whether it already existed or you wrote it, and the red-then-green evidence — or the explicit reason no test was needed.
  • Base merged in (and whether there were conflicts, and how you resolved any non-obvious ones).
  • Local verification result.
  • Diff self-review: any regressions caught and fixed, and any dead/unused code the branch introduced that you removed (name it) — or that the diff was clean.
  • Any scratch planning/findings docs stripped before push (name each, and where you stashed a copy if one was worth keeping) — or that there were none.
  • PR URL (and whether you created it or it already existed) — note it's ready for review (not a draft).
  • Whether the Entire trail body was written fresh from the current branch state (and rewritten after later fixes), or why it was skipped.
  • Final CI status: all green, or the specific checks still failing and why.
  • Bugbot result: how many comments it raised, how many you fixed (and what), how many you skipped (and why) — or that bugbot isn't configured for the repo.
  • Slop meter: the rating (small/medium/large), and if it was medium or large, what you changed to reduce it and where it landed on the re-check — or that the repo has no slop runner / no trail.
  • Trail findings: how many open findings were on the trail, how many you resolved (and what), how many you dismissed (and why) — or that there was no trail / the repo isn't Entire-enabled.

If everything's green and the user wants to request review, point them at the post-pr-review skill — that assigns the PR and posts it to Slack. This skill deliberately stops short of that.

4d ago·4m

Blocked on one permission: opening the PR. Everything up to the push is done.

Done

  • Test coverage: new TestResolveDefaultSearchRepo (native, GitHub, mirror origins, explicit scope, no origin with and without explicit scope) and three forge cases in TestConfig_ScopeSlugs. Red with source reverted (build failure, the symbols are new), green with it restored. Behavioral red: the pre-fix binary exits 1 on the native clone where the fixed one exits 0.
  • Base: branch is 0 behind origin/main, so no merge was needed.
  • Local verification: mise run lint 0 issues, mise run test passes except TestOpenCodeSeedRepoPlantsDeps in e2e/agents, which fails because opencode --version prints nothing on this machine. The branch touches nothing under e2e.
  • Self-review: diff is clean. No dead code; the GitHub-only parser is still used by the two dispatch callers.
  • Scratch docs: FINDINGS.md is excluded via .git/info/exclude and is not in the commit. It stays in the worktree.
  • Commit: 476493981 pushed as origin/native-search-repo-routing.

Blocked

  • gh pr create was denied by the auto-mode classifier as an external write. The PR body is ready at /tmp/native-search-repo-routing/pr-body.md. To open it yourself:
  • No trail exists yet for the branch, so the trail body (drafted at /tmp/native-search-repo-routing/trail-body.md) could not be set. CI has not started on the branch either.

Once the PR exists, say so and I'll continue with the trail body, CI watch, bugbot, slop meter, and trail findings.

you do it

4d ago·2m

CI is running (watched in background), trail 1464 has 0 open findings so far, and bugbot runs as a PR check. I'll pick up the CI result, bugbot output, slop meter, and findings when the watch completes.

4d ago·2m

Finding resolved, trail body rewritten. CI on the new head is being watched in the background; I'll report when it finishes and then re-check the trail for any new findings on that head.