Debug and Fix Dispatch on Native Repos

Using superpowers:systematic-debugging to structure the investigation, then reproduce, check Datadog, trace the code, and search Entire history.
Base directory for this skill: /Users/alisha/.claude/plugins/cache/claude-plugins-official/superpowers/6.4.1/skills/systematic-debugging
Systematic Debugging
Overview
Core principle: ALWAYS find root cause before attempting fixes. Symptom fixes are failure.
Violating the letter of this process is violating the spirit of debugging.
The Iron Law
If you haven't completed Phase 1, you cannot propose fixes.
When to Use
Use for ANY technical issue:
- Test failures
- Bugs in production
- Unexpected behavior
- Performance problems
- Build failures
- Integration issues
Use this ESPECIALLY when:
- Under time pressure (emergencies make guessing tempting)
- "Just one quick fix" seems obvious
- You've already tried multiple fixes
- Previous fix didn't work
- You don't fully understand the issue
Don't skip when:
- Issue seems simple (simple bugs have root causes too)
- You're in a hurry (rushing guarantees rework)
- Manager wants it fixed NOW (systematic is faster than thrashing)
The Four Phases
You MUST complete each phase before proceeding to the next.
Phase 1: Root Cause Investigation
BEFORE attempting ANY fix:
-
Read Error Messages Carefully
- Don't skip past errors or warnings
- They often contain the exact solution
- Read stack traces completely
- Note line numbers, file paths, error codes
-
Reproduce Consistently
- Can you trigger it reliably?
- What are the exact steps?
- Does it happen every time?
- If not reproducible → gather more data, don't guess
-
Check Recent Changes
- What changed that could cause this?
- Git diff, recent commits
- New dependencies, config changes
- Environmental differences
-
Gather Evidence in Multi-Component Systems
WHEN system has multiple components (CI → build → signing, API → service → database):
BEFORE proposing fixes, add diagnostic instrumentation:
Example (multi-layer system):
This reveals: Which layer fails (secrets → workflow ✓, workflow → build ✗)
-
Trace Data Flow
WHEN error is deep in call stack:
See
root-cause-tracing.mdin this directory for the complete backward tracing technique.Quick version:
- Where does bad value originate?
- What called this with bad value?
- Keep tracing up until you find the source
- Fix at source, not at symptom
Phase 2: Pattern Analysis
Find the pattern before fixing:
-
Find Working Examples
- Locate similar working code in same codebase
- What works that's similar to what's broken?
-
Compare Against References
- If implementing pattern, read reference implementation COMPLETELY
- Don't skim - read every line
- Understand the pattern fully before applying
-
Identify Differences
- What's different between working and broken?
- List every difference, however small
- Don't assume "that can't matter"
-
Understand Dependencies
- What other components does this need?
- What settings, config, environment?
- What assumptions does it make?
Phase 3: Hypothesis and Testing
Scientific method:
-
Form Single Hypothesis
- State clearly: "I think X is the root cause because Y"
- Write it down
- Be specific, not vague
-
Test Minimally
- Make the SMALLEST possible change to test hypothesis
- One variable at a time
- Don't fix multiple things at once
-
Verify Before Continuing
- Did it work? Yes → Phase 4
- Didn't work? Form NEW hypothesis
- DON'T add more fixes on top
-
When You Don't Know
- Say "I don't understand X"
- Don't pretend to know
- Ask for help
- Research more
Phase 4: Implementation
Fix the root cause, not the symptom:
-
Create Failing Test Case
- Simplest possible reproduction
- Automated test if possible
- One-off test script if no framework
- MUST have before fixing
- Use the
superpowers:test-driven-developmentskill for writing proper failing tests
-
Implement Single Fix
- Address the root cause identified
- ONE change at a time
- No "while I'm here" improvements
- No bundled refactoring
-
Verify Fix
- Test passes now?
- No other tests broken?
- Issue actually resolved?
- Use the
superpowers:verification-before-completionskill before claiming success
-
If Fix Doesn't Work
- STOP
- Count: How many fixes have you tried?
- If < 3: Return to Phase 1, re-analyze with new information
- If ≥ 3: STOP and question the architecture (step 5 below)
- DON'T attempt Fix #4 without architectural discussion
-
If 3+ Fixes Failed: Question Architecture
Pattern indicating architectural problem:
- Each fix reveals new shared state/coupling/problem in different place
- Fixes require "massive refactoring" to implement
- Each fix creates new symptoms elsewhere
STOP and question fundamentals:
- Is this pattern fundamentally sound?
- Are we "sticking with it through sheer inertia"?
- Should we refactor architecture vs. continue fixing symptoms?
Discuss with your human partner before attempting more fixes
This is NOT a failed hypothesis - this is a wrong architecture.
Red Flags - STOP and Follow Process
If you catch yourself thinking:
- "Quick fix for now, investigate later"
- "Just try changing X and see if it works"
- "Add multiple changes, run tests"
- "Skip the test, I'll manually verify"
- "It's probably X, let me fix that"
- "I don't fully understand but this might work"
- "Pattern says X but I'll adapt it differently"
- "Here are the main problems: [lists fixes without investigation]"
- Proposing solutions before tracing data flow
- "One more fix attempt" (when already tried 2+)
- Each fix reveals new problem in different place
ALL of these mean: STOP. Return to Phase 1.
If 3+ fixes failed: Question the architecture (see Phase 4.5)
your human partner's Signals You're Doing It Wrong
Watch for these redirections:
- "Is that not happening?" - You assumed without verifying
- "Will it show us...?" - You should have added evidence gathering
- "Stop guessing" - You're proposing fixes without understanding
- "Ultra-think this" - Question fundamentals, not just symptoms
- "We're stuck?" (frustrated) - Your approach isn't working
When you see these: STOP. Return to Phase 1.
Common Rationalizations
| Excuse | Reality |
|---|---|
| "Issue is simple, don't need process" | Simple issues have root causes too. Process is fast for simple bugs. |
| "Emergency, no time for process" | Systematic debugging is FASTER than guess-and-check thrashing. |
| "Just try this first, then investigate" | First fix sets the pattern. Do it right from the start. |
| "I'll write test after confirming fix works" | Untested fixes don't stick. Test first proves it. |
| "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. |
| "Reference too long, I'll adapt the pattern" | Partial understanding guarantees bugs. Read it completely. |
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause. |
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem. Question pattern, don't fix again. |
Quick Reference
| Phase | Key Activities | Success Criteria |
|---|---|---|
| 1. Root Cause | Read errors, reproduce, check changes, gather evidence | Understand WHAT and WHY |
| 2. Pattern | Find working examples, compare | Identify differences |
| 3. Hypothesis | Form theory, test minimally | Confirmed or new hypothesis |
| 4. Implementation | Create test, fix, verify | Bug resolved, tests pass |
When Process Reveals "No Root Cause"
If systematic investigation reveals issue is truly environmental, timing-dependent, or external:
- You've completed the process
- Document what you investigated
- Implement appropriate handling (retry, timeout, error message)
- Add monitoring/logging for future investigation
But: 95% of "no root cause" cases are incomplete investigation.
Supporting Techniques
These techniques are part of systematic debugging and available in this directory:
root-cause-tracing.md- Trace bugs backward through call stack to find original triggerdefense-in-depth.md- Add validation at multiple layers after finding root causecondition-based-waiting.md- Replace arbitrary timeouts with condition polling
All new and existing dispatch tests pass. Now verifying end-to-end with the rebuilt binary, running lint, and loading the write-up skill.
Base directory for this skill: /Users/alisha/.claude/skills/architecture-findings
Architecture findings
What this is for
After an investigation, there are two very different documents you could hand someone:
- A lookup table — every claim pinned to a file and line so an engineer can jump straight to the code. Precise, but dense and exhausting to read.
- A map of the territory — who the moving parts are, how they normally fit together, where the break is, and whose job it is to fix. This is the one that actually transfers understanding.
This skill produces the second kind. The reader should come away able to explain the system to someone else, not just able to find the line you're talking about.
The core rule
No code coordinates in the body. No file paths, no line numbers, no function or variable names, no code snippets. Those are the texture of the lookup table, and they're exactly what makes findings hard to read.
What you do name freely: the components, services, and systems by their
real names (e.g. "the EU indexer", "entire-core", "the NATS bus"), the
concepts ("ref events", "the repo-access token"), and the environment
values that matter to the story only when named as a fact, not a citation
("the EU audience host is aws-eu-central-1.entire.io"). The test: would this
detail help someone picture the system, or is it just a pointer for someone
who wants to open a file? Keep the former, drop the latter.
You still did the rigorous, code-level investigation — this is just how you report it. If the user later asks "where exactly", give them the coordinates then. Don't pre-empt that with a wall of citations.
Structure
Adapt to the investigation, but this arc almost always works:
- Bottom line — 2–4 sentences. What's happening, why, and the one thing that matters most (e.g. "the fix isn't where you'd expect — it's upstream"). Someone who reads only this should get the headline right.
- The players — introduce each component in one line: what it is and what it's responsible for. This is the cast list; everything after refers back to it.
- The normal flow — how it works when it works. This is the reference the reader measures the break against. A simple arrow-chain diagram earns its place here (see below).
- Why it breaks — the failure, told as a story against that normal flow. Show the point where reality diverges from the happy path.
- The structural gap (when there is one) — the deeper "why can't we just fix it here" — the missing capability or wrong assumption underneath the symptom. This is often the most valuable section: it's what stops the reader proposing the naive fix.
- Where the fix belongs — which system owns it and why, what the options are, and explicitly what the wrong fix is if there's a tempting one.
- How other cases differ (when relevant) — the comparison that proves it's structural, not a one-off (other regions, other tenants, the working vs. broken side).
Drop sections that don't apply. A small investigation might be just bottom line + why it breaks + where the fix belongs.
Flow diagrams
Simple ASCII arrow chains are encouraged where a flow clarifies the story — they carry sequence better than a paragraph. Keep them component-level, never code-level. For example:
Then a second chain showing where the broken path diverges is often the clearest way to show a failure. Label them ("normal flow" / "what EU does"). Don't force a diagram where prose is already clear.
Voice
Write for a smart colleague who doesn't live in this code. Explain the why behind each mechanism, not just the what — "the STS refuses because that repo isn't an EU repo; that refusal is the residency guard doing its job" teaches more than "the STS returns an error". Prefer plain language over jargon, and when a term is load-bearing, define it once in passing. Keep it tight — this is a briefing, not an essay.
What to avoid
- File paths, line numbers, symbol names, code blocks (other than the component-level flow diagrams above).
- Turning "the players" into an inventory of every file touched.
- Hedging everything. You investigated; state what you found plainly, and flag genuine open questions as open questions.
- Burying the headline. The bottom line goes first.
Full mise run check is running in the background (lint passed, unit tests in progress). Once it finishes I'll remove the now-orphaned GitHub-only parser, re-run its package tests, and commit.
Re-running the full pre-commit check after the dead-code removal; I'll commit once it's green and then deliver the write-up.
Verdict: confirmed, and the fix is in the CLI. Committed on dispatch-native-repos as efb0f1748, not pushed. Code-level detail is in FINDINGS.md (gitignored, in the worktree).
Bottom line
entire dispatch in an Entire-native checkout dies before it ever talks to the server. The command works out which repo to summarize by reading the origin remote, and that step only knew how to read GitHub. A native checkout's origin points at an Entire cell, so the command refused it as "not a GitHub repository". The server side is fine: naming the same repo explicitly as et/entirehq/entire-api produced a full dispatch during the investigation. The gap was entirely in how the CLI names the repo it is standing in.
The players
- The dispatch command in the CLI. Collects a window and a set of repos, then either asks Entire to generate the summary (cloud mode) or builds one from local checkpoints (local mode).
- The repo slug. Since late September every repo a dispatch covers is addressed by a forge-qualified name:
gh/owner/repofor GitHub-backed repos,et/project/repofor Entire-native ones. A bareowner/repois refused because the same name can exist on both forges. - The origin remote. For a GitHub checkout it points at github.com or at an Entire mirror of a GitHub repo. For a native checkout it points at an Entire cell, with the path spelling out which forge the repo belongs to.
- The remote URL parser shared across the CLI. It already understands every one of those URL shapes and reports which forge a remote belongs to.
entire searchwas switched onto it earlier today for exactly this class of bug. - The dispatch service in entire-api. Accepts
gh/andet/slugs, resolves each on its own forge, and generates the summary.
The normal flow
With --repos, the first two steps are skipped and the user's slugs go straight to the service.
Why it broke
The slug derivation step used a GitHub-only parser left over from when dispatch was GitHub-only. It accepted github.com remotes and Entire mirrors of GitHub repos, and rejected any other host. It then prefixed the result with gh/ unconditionally, on the stated assumption that "the origin remote is GitHub by construction".
That assumption stopped being true the moment native repos existed, but nothing forced the derivation to change when --repos learned about et/. The same derivation sits behind local mode and the wizard's on-disk repo discovery, so --local failed with the same message and the interactive picker silently left native checkouts out.
Datadog could not show this failure because it never reaches a server. The dispatch endpoint saw twelve requests in the last day, all GitHub-shaped or explicit, with successes and some rate-limiting. The only "dispatch failed" noise in logs is entiredb's git protocol dispatcher, which is unrelated.
Where the fix belongs, and what was done
The fix belongs in the CLI, at the one step that turns a remote into a slug. It now uses the shared remote parser and emits whatever forge the URL actually names. GitHub checkouts behave exactly as before. Native checkouts dispatch as et/project/repo in cloud mode, are named that way in local mode with no GitHub link, and are offered by the wizard's disk discovery. A remote on a host Entire does not serve is still refused, and the error now points the user at --repos. The GitHub-only parser had no remaining callers and was removed.
The wrong fix would have been to teach the server to guess a forge from a bare name. That is the collision the forge prefix exists to prevent.
Verification. Five new tests fail on main with the reproduced error and pass with the fix. The full mise run check passed. The rebuilt binary generated a dispatch from the native entire-api checkout with no flags, and --local runs there too. One unrelated placement-latency test flaked once on port reuse and passed three times in isolation.
Two gaps left open
- The wizard's authenticated repo list labels every row
gh/, because the listing it calls returns names without a forge. Correct labelling needs a forge field on that response or a switch to the control-plane repo index, which does carry a provider. Until then the picker offers native repos only when that listing fails and disk discovery runs. - The jurisdiction hints (placement join and the not-found hint) are keyed by GitHub name, so a native repo falls back to the home jurisdiction. Harmless, but the same control-plane provider field could key
et/rows too.
Want this as a shareable page?
Base directory for this skill: /Users/alisha/.claude/skills/rubber-duck
Rubber Duck
The user is looking at something dense — usually a long agent response or a findings dump — and wants to understand it, not act on it. Your whole job is to be the clear, patient explainer and then a sounding board for their questions. Think out loud with them, in plain language.
The one job
Turn the wall of text into something a smart person with no context can follow in under a minute, then answer whatever they ask next — still briefly.
Hard rules (do not break these)
- Do NOT write a plan, and do NOT enter plan mode. No ExitPlanMode, no writing-plans/brainstorming skills, no "here's how we'd fix it" roadmaps.
- Do NOT change code, run builds, or edit files. This is understanding only.
- Do NOT make a todo list. It's a conversation, not a task.
- Stay read-only. Reading a file or scrolling the text to explain it is fine. Anything that does something is not.
- Be concise — this is the point of the skill. You are the antidote to the wall of text, so don't produce another one. Lead with the answer. Short sentences. No preamble, no "great question", no restating the question back.
- Plain language over jargon. No file paths, line numbers, symbol names, or
code snippets unless the user asks for them. Names of systems/people are fine;
citations are not. (This is the same spirit as the
architecture-findingsskill — the story, not the lookup table.)
First explanation — the format
Keep it to roughly this, and stop:
- One-line TL;DR. The single most important thing, first.
- 3–5 short bullets covering: what's actually going on, why it matters, and (if relevant) what's broken or being decided. One idea per bullet.
- One line on "so what" — the upshot, or the decision in front of them.
If the source genuinely has more than fits, say "there's more detail under X if you want it" and let them pull it — don't push it.
Use an analogy when it makes a hard idea click. Skip it when it doesn't.
The follow-up loop
After the first explanation, they'll ask things. For each:
- Answer the specific question and nothing more. One concept at a time.
- Stay just as concise — don't let answers creep back into walls of text.
- It's fine to say "not sure / the text doesn't say" rather than guess.
- If a question needs you to look something up to answer honestly, look (read only), then answer plainly.
Keep going as long as they're asking. You're the duck — they're thinking out loud and you're helping them see it.
When they're ready to act
The moment the user shifts from "help me understand" to "okay let's fix/build/ plan this," rubber-duck mode is over. Don't quietly start doing it inside this mode. Say something like: "That's action territory — want me to pick up that work?" and let them start the real task (a plan, a dispatch, an edit) as its own thing. Your job here ends at understanding.
TL;DR: entire dispatch broke in native repos because the CLI assumed your origin remote was always GitHub. Fixed in the CLI, committed, not pushed.
- What was happening: With no
--reposflag, the CLI reads your gitoriginto figure out which repo to summarize. That reader only understood GitHub URLs, so a native repo'sentire://…/et/…origin got rejected before any request was sent. - Why it looked like a server problem but wasn't: Passing
--repos et/entirehq/entire-apiexplicitly worked fine the whole time. The server already handles native repos. Only the "guess the repo from where I'm standing" step was broken. - The fix: Swap the GitHub-only reader for the shared URL parser that already knows both forges (
entire searchgot this same fix earlier today). Now a native checkout dispatches aset/…, GitHub staysgh/…, and anything else gets a clear error pointing at--repos. - Proof: Five new tests reproduce the old error on main and pass with the fix. The rebuilt binary generated a real dispatch from the native entire-api checkout with no flags.
- Still open: The interactive wizard's repo list can't tell forges apart because the API it calls doesn't say which forge a repo is on. That needs an API change, not a CLI one.
So what: Nothing left to decide for the core bug. The branch is ready for you to review and push whenever. The wizard gap is a separate follow-up if you want it.
does this pr fix it? https://github.com/entireio/cli/pull/2639
No. That PR is the entire search fix for the same class of bug, merged this morning. It only touched the search command's files.
Your worktree already includes it (main at 55a695e43 is past that merge), and entire dispatch still failed there with the exact "not a GitHub repository" error. Dispatch has its own copy of the origin-reading step and #2639 never reached it.
The analogy: two rooms with the same broken lock. #2639 fixed the search room. This branch fixes the dispatch room with the same key.
okay lets fix it and is there a way to add a test to catch when cli changes breaks dispatch or search commands?
Done. Two commits on dispatch-native-repos, not pushed:
- efb0f1748 — the dispatch fix (origin's forge drives the slug; GitHub-only parser removed).
- 40e36a8de — the regression guard you asked for.
How the guard works. A ledger test git-greps every non-test file that reads the origin remote. Each one must have an entry saying what it does with the forge, and name tests that run it against an entire://…/et/… origin. Entries may instead say DROPS or NOT PINNED with a written reason. The test fails on an unregistered reader, a stale entry, a missing pinning test, or an unjustified unpinned entry. I verified it by dropping in a throwaway file that calls the origin resolver: the build failed naming that file.
What it found along the way. Three more commands handle native checkouts imperfectly and are now recorded as known gaps in the ledger rather than fixed:
entire expertsand the recap data-API scope use a bare owner/repo, so a native repo resolves GitHub-first.- Repo-scoped recap's mirror lookup knows GitHub only, so a native checkout gets a personal-only recap.
Both mise run check runs passed. The only flake was an unrelated placement-latency port-reuse test that passed three times in isolation.
Next: say the word and I'll push and open the PR.