Fix FirstRun Planner Against Invented Steps

Claude Code·Sonnet 5.5·HighnessAtharva·yesterday·3min·1 Checkpoint·4 file changes·+306/-15·18.1K tokens

You are fixing the FirstRun planner with a fix that came from neatlogs. Repo: this folder (D:\Atharva\firstrun), branch main. Python venv: .venv\Scripts\python.exe. Never print or commit a key. Never copy text from plan/ into the repo.

  1. Use the neatlogs MCP tools first: call search_traces / get_trace_context for the planner spans flagged by the detection "planner invented a step" (list_detections, get_detection_trend), and call triage_list. If a triage item about the planner exists, read it with triage_get and triage_get_investigation and use it. The Investigate answer Atharva ran is saved in evidence/neatlogs-investigation.md; use it as the fix spec either way.

  2. Locate code with entire graph search --repo . --format agent --query "planner system prompt and runbook validation" before any grep.

  3. Apply the fix in src/firstrun/planner.py: a. Replace SYSTEM_PROMPT with the neatlogs prompt, adapted to our schema field names (id, intent, command, files[{path, content}], expected{exit_code, output_contains, file_exists}, source_line, skipped_out_of_scope, skip_reason) and keeping our existing rules that still hold (copy commands exactly, never fix a command, image choice, skip GUI/signup steps as skipped_out_of_scope, 10 step cap). Keep it tight. b. Add code guards that run on every planned runbook, because a prompt alone cannot stop a model from slipping: drop no-op commands (true, :, empty); drop or rewrite fake-secret exports (values like sk-..., xxx, your-key, <...>) to export NAME="${NAME}" only if the docs show that export, else drop; drop invented checks (command -v, which, test -f, test -e, ls) unless that exact text appears in the docs markdown; drop any other command whose normalized text does not appear in the docs markdown, except python <file> / node <file> / bash <file> that run a file the step itself creates. Log every dropped or rewritten step with neatlogs.log("planner guard dropped step: {command} ({why})", ...) inside the planner span and record the count as a span attribute neatlogs.planner.guard_dropped. c. Determinism: cache the runbook keyed by sha256 of the docs markdown plus the planner prompt version in config.RUNS_DIR/.plan_cache/<hash>.json, so the same page gives the same plan. Add a --replan flag only if trivial; otherwise an env var FIRSTRUN_REPLAN=1 bypasses the cache. d. Steps that need an API key stay in the runbook (do not skip them).

  4. Tests in tests/test_planner_guard.py: each of the v1 mistakes is dropped (true, command -v entire, test -f first_trace.py, export OPENAI_API_KEY='sk-...'), a real docs command is kept, python <file> with files is kept, the cache returns the same runbook twice. Update any existing test that asserted old behavior. Run the full suite: .venv\Scripts\python -m pytest -q -p no:cacheprovider and keep it green.

  5. Commit only these files (planner.py, the tests, evidence/neatlogs-investigation.md if unchanged is fine to include) with this exact message body style, and with NO Co-Authored-By line and no other trailer you add yourself: Planner fix from neatlogs Investigate: copy commands only, guard against invented steps

    The neatlogs detection "planner invented a step" flagged the planner spans of 3 v1 runs. neatlogs Investigate traced it to the planner prompt and proposed a stricter one (evidence/neatlogs-investigation.md). This applies it and adds code guards that drop no-op, invented and fake-secret steps, plus a plan cache so the same page gives the same plan.

  6. Run entire checkpoint list and print the last 3 lines.

  7. If a triage item existed, call triage_update_status to set it in_progress (not resolved yet; the lead resolves it after v2 proves the fix). Finish with 6 lines: MCP calls made and what they returned, files changed, test result, commit SHA, guard counts on the 3 v1 planner outputs if you replayed them (optional), triage item status.

yesterday·3m

Planner fix committed as 3ca6f11 and the full suite is green. Nothing was replayed against the 3 v1 outputs.

  1. MCP calls:
    • list_detections returned 9 detections, including "planner invented a step".
    • get_detection_trend showed 3 occurrences on 2026-10-10.
    • triage_list returned 0 cards, so I used evidence/neatlogs-investigation.md as the spec.
    • search_traces with fields failed with not_enabled (NL-1923 rollout). A retry without fields, filtered by the detection name, returned 0 results, so I read no trace spans.
  2. Files:
    • src/firstrun/planner.py has the new prompt (PROMPT_VERSION "2"), the guards and the cache.
    • tests/test_planner_guard.py is new.
    • tests/test_edge_planner.py has 3 tests updated because their docs text lacked the commands.
    • evidence/neatlogs-investigation.md is included unchanged.
  3. Tests: 175 passed, 24 skipped, 1 xfailed.
  4. Commit: 3ca6f11, no Co-Authored-By line from me. Entire's hook added its own Entire-Checkpoint trailer. entire checkpoint list ended with 01M4JBCP5ZFNH2JYXAFR6V910P "Only operate on the main branch for this project!!..." and 10-10 12:58 (81d6403) Contracts: pydantic models and module stubs for the seven tracks.
  5. Guard counts: not replayed. Each run records neatlogs.planner.guard_dropped on the planner span. I also added neatlogs.planner.guard_rewritten.
  6. Triage: no item existed, so triage_update_status was not called.

I made four calls the spec did not cover:

  • Cache scope: the cache is only used when no llm is injected, so tests and replays never touch it. FIRSTRUN_REPLAN=1 bypasses it.
  • Dropped step files: files from a dropped step move to the next kept step, so a later python <file> still finds them. Without this, the true step with files would lose its file.
  • Skipped steps: steps with skipped_out_of_scope are not guarded, because they never run.
  • Chained commands: a && chain counts as in the docs when every part appears there.

Next: run firstrun on the Chroma page once to see real guard counts in neatlogs.