Fix FirstRun Planner Against Invented Steps

You are fixing the FirstRun planner with a fix that came from neatlogs. Repo: this folder (D:\Atharva\firstrun), branch main. Python venv: .venv\Scripts\python.exe. Never print or commit a key. Never copy text from plan/ into the repo.
-
Use the neatlogs MCP tools first: call search_traces / get_trace_context for the planner spans flagged by the detection "planner invented a step" (list_detections, get_detection_trend), and call triage_list. If a triage item about the planner exists, read it with triage_get and triage_get_investigation and use it. The Investigate answer Atharva ran is saved in evidence/neatlogs-investigation.md; use it as the fix spec either way.
-
Locate code with
entire graph search --repo . --format agent --query "planner system prompt and runbook validation"before any grep. -
Apply the fix in src/firstrun/planner.py: a. Replace SYSTEM_PROMPT with the neatlogs prompt, adapted to our schema field names (id, intent, command, files[{path, content}], expected{exit_code, output_contains, file_exists}, source_line, skipped_out_of_scope, skip_reason) and keeping our existing rules that still hold (copy commands exactly, never fix a command, image choice, skip GUI/signup steps as skipped_out_of_scope, 10 step cap). Keep it tight. b. Add code guards that run on every planned runbook, because a prompt alone cannot stop a model from slipping: drop no-op commands (
true,:, empty); drop or rewrite fake-secret exports (values like sk-..., xxx, your-key, <...>) toexport NAME="${NAME}"only if the docs show that export, else drop; drop invented checks (command -v, which, test -f, test -e, ls) unless that exact text appears in the docs markdown; drop any other command whose normalized text does not appear in the docs markdown, exceptpython <file>/node <file>/bash <file>that run a file the step itself creates. Log every dropped or rewritten step with neatlogs.log("planner guard dropped step: {command} ({why})", ...) inside the planner span and record the count as a span attribute neatlogs.planner.guard_dropped. c. Determinism: cache the runbook keyed by sha256 of the docs markdown plus the planner prompt version in config.RUNS_DIR/.plan_cache/<hash>.json, so the same page gives the same plan. Add a --replan flag only if trivial; otherwise an env var FIRSTRUN_REPLAN=1 bypasses the cache. d. Steps that need an API key stay in the runbook (do not skip them). -
Tests in tests/test_planner_guard.py: each of the v1 mistakes is dropped (true, command -v entire, test -f first_trace.py, export OPENAI_API_KEY='sk-...'), a real docs command is kept, python <file> with files is kept, the cache returns the same runbook twice. Update any existing test that asserted old behavior. Run the full suite:
.venv\Scripts\python -m pytest -q -p no:cacheproviderand keep it green. -
Commit only these files (planner.py, the tests, evidence/neatlogs-investigation.md if unchanged is fine to include) with this exact message body style, and with NO Co-Authored-By line and no other trailer you add yourself: Planner fix from neatlogs Investigate: copy commands only, guard against invented steps
The neatlogs detection "planner invented a step" flagged the planner spans of 3 v1 runs. neatlogs Investigate traced it to the planner prompt and proposed a stricter one (evidence/neatlogs-investigation.md). This applies it and adds code guards that drop no-op, invented and fake-secret steps, plus a plan cache so the same page gives the same plan.
-
Run
entire checkpoint listand print the last 3 lines. -
If a triage item existed, call triage_update_status to set it in_progress (not resolved yet; the lead resolves it after v2 proves the fix). Finish with 6 lines: MCP calls made and what they returned, files changed, test result, commit SHA, guard counts on the 3 v1 planner outputs if you replayed them (optional), triage item status.
Planner fix committed as 3ca6f11 and the full suite is green. Nothing was replayed against the 3 v1 outputs.
- MCP calls:
list_detectionsreturned 9 detections, including "planner invented a step".get_detection_trendshowed 3 occurrences on 2026-10-10.triage_listreturned 0 cards, so I usedevidence/neatlogs-investigation.mdas the spec.search_traceswithfieldsfailed withnot_enabled(NL-1923 rollout). A retry withoutfields, filtered by the detection name, returned 0 results, so I read no trace spans.
- Files:
src/firstrun/planner.pyhas the new prompt (PROMPT_VERSION"2"), the guards and the cache.tests/test_planner_guard.pyis new.tests/test_edge_planner.pyhas 3 tests updated because their docs text lacked the commands.evidence/neatlogs-investigation.mdis included unchanged.
- Tests: 175 passed, 24 skipped, 1 xfailed.
- Commit: 3ca6f11, no Co-Authored-By line from me. Entire's hook added its own
Entire-Checkpointtrailer.entire checkpoint listended with01M4JBCP5ZFNH2JYXAFR6V910P "Only operate on the main branch for this project!!..."and10-10 12:58 (81d6403) Contracts: pydantic models and module stubs for the seven tracks. - Guard counts: not replayed. Each run records
neatlogs.planner.guard_droppedon the planner span. I also addedneatlogs.planner.guard_rewritten. - Triage: no item existed, so
triage_update_statuswas not called.
I made four calls the spec did not cover:
- Cache scope: the cache is only used when no
llmis injected, so tests and replays never touch it.FIRSTRUN_REPLAN=1bypasses it. - Dropped step files: files from a dropped step move to the next kept step, so a later
python <file>still finds them. Without this, thetruestep with files would lose its file. - Skipped steps: steps with
skipped_out_of_scopeare not guarded, because they never run. - Chained commands: a
&&chain counts as in the docs when every part appears there.
Next: run firstrun on the Chroma page once to see real guard counts in neatlogs.