Activity
Contributors
README
FirstRun
Docs that work. Literally.
FirstRun runs your quickstart in a clean container, finds the first step that breaks, and proves the fix with a second clean run.
Live site neatHack 2026 Tests Python
Demo video (2:59): neatlogs, Entire and cfo.ai, architecture to result · Try it in the browser · Watch a real run · See the 4 upstream PRs · For judges · Follow the build on X
Before and after: v1, the baseline, finished 2 of 4 quickstarts and recovered 0 failed steps. v2 finished 3 of 4, recovered 3 failed steps and proved 1 fix in a fresh container. See the comparison.
<br> <table> <tr> <td align="center" width="25%"><a href="https://www.youtube.com/watch?v=JJNWWhSV9pc"><img src="docs/assets/vid-doctest.webp" alt="DocTest demo thumbnail: step 10 of a quickstart fails, then the fix passes 9 of 9 steps" width="210"></a><br><sub><b>DocTest demo (2:59)</b></sub></td> <td align="center" width="25%"><a href="https://www.youtube.com/watch?v=2vZb6pZAF8Q"><img src="docs/assets/vid-docaudit-journey.webp" alt="DocAudit thumbnail: 15 stale pages found in dbt-core, each fix goes to the person who knows the code" width="210"></a><br><sub><b>DocAudit journey (1:56)</b></sub></td> <td align="center" width="25%"><a href="https://www.youtube.com/watch?v=jKL8Ocb0Ki8"><img src="docs/assets/vid-teaser.webp" alt="Teaser thumbnail: step 10 of the dbt DuckDB guide fails" width="210"></a><br><sub><b>Teaser (0:36)</b></sub></td> <td align="center" width="25%"><a href="https://www.youtube.com/watch?v=zo94agwRAXc"><img src="docs/assets/vid-docaudit.webp" alt="DocAudit thumbnail: a dbt-core page is 453 days stale" width="210"></a><br><sub><b>DocAudit tour (0:57)</b></sub></td> </tr> </table> <br> <video src="docs/assets/firstrun-launch.mp4" poster="docs/assets/launch-poster.jpg" width="860" muted loop playsinline controls> <img src="docs/assets/launch-poster.jpg" alt="FirstRun launch film: a quickstart fails at step 3 of 6" width="860"> </video><sub>The 44-second launch film. The video does not play for you? Open it on the site or download the MP4.</sub>
</div> <br>The Demo Video Runs 2 Minutes 59 Seconds and Gives Each Tool a Chapter
Watch the demo video on YouTube (2:59). It runs FirstRun on the dbt guide at 0:12 and shows Bisect at 0:36. The neatlogs chapter starts at 1:07, Entire at 1:51 and cfo.ai at 2:19.
Any Coding Agent Can Start From agent.md
Give Claude Code, Cursor, Codex or your own agent this one line:
- Open Agent Ready in the site navigation to copy the prompt. It adds an example task and shows the 6 steps your agent runs.
- docs/agent.md holds the preflight, the install, a recipe for each task, the JSON keys, the exit and error codes, and the safety rules.
- llms.txt lists the agent files, and llms-full.txt puts them in one file.
firstrun run,audit,receiptandprereqstake--json, which prints one JSON object on stdout.firstrun doctor --jsonchecks Python, git, Docker, the keys and the network, and prints no key value.- AGENTS.md tells a coding agent how to build and test this repo.
- skills/firstrun is an Agent Skill. Copy the folder into
.claude/skills/in a project, or into~/.claude/skills/for every project, and Claude Code loads it.
Break, Fix, Prove Takes 195 Seconds
<div align="center"> <img src="docs/assets/step-prove.webp" alt="Break. Fix. Prove. Step 3 of 6 fails, the docs fix is applied, and a second clean container passes 6 of 6 steps" width="760"> </div> <table> <tr> <td width="33%" valign="top"><img src="docs/assets/step-run.webp" alt="A firstrun command above six numbered quickstart steps"><br><b>1. Read</b><br>FirstRun reads the page and plans one step per command.</td> <td width="33%" valign="top"><img src="docs/assets/step-3.webp" alt="Six quickstart steps. Step 3 fails for a new developer"><br><b>2. Run</b><br>Every step runs in a clean container, like a new developer. The first failure stops the run.</td> <td width="33%" valign="top"><img src="docs/assets/step-prove.webp" alt="A second clean container passes 6 of 6 steps"><br><b>3. Prove</b><br>A fix is done only when a second clean container passes the whole quickstart.</td> </tr> </table>The input below is a known-broken test fixture (tests/fixtures/docs/broken_quickstart_FIXTURE.md), not a real project's docs. Step 3 pins an old chromadb that cannot install on Python 3.11.
FirstRun writes report.md, patch.diff and pr_body.md into the run folder. The patch is one line.
Seven Teams Get a Different Job Out of the Same Engine
| # | You are | The problem | Run this | You get |
|---|---|---|---|---|
| 1 | A DevRel or docs lead | Nobody tests the quickstart after launch day | firstrun <url> | A pass or a verified fix, with a PR draft |
| 2 | A release manager | A new release of your package broke the docs and you do not know which one | firstrun bisect <run> --package <pkg> | The last good and first bad release |
| 3 | A technical writer | The docs assume tools the reader does not have | firstrun prereqs <run> | A Prerequisites block, proven in a clean container |
| 4 | An open-source contributor | You want a first PR that a maintainer merges | firstrun pr <run> | The gh commands for the docs PR, source check included |
| 5 | An engineer picking a tool | You want to know if the vendor's quickstart works before you commit | firstrun <url> | A time to first success and the first step that breaks |
| 6 | A coding agent | You need a loop that checks its own docs work | firstrun run --resume <id> | State after every step and a model-free verify |
| 7 | A head of docs | Hundreds of pages, and the code moved under some of them | firstrun audit owner/repo | Stale pages with the commits behind them, features with no docs, proposal cards and drafted release notes |
One blobless clone gives the whole history. The audit then works offline and costs $0 in model calls. Jev sorts commits with no conventional prefix when TYPESAFE_API_KEY is set, and every audit is one neatlogs trace.
- Oldest first. Every page, ranked by the date of its last edit.
- Stale or quiet. A page is stale when product commits on its topic landed after its last edit, and the report names those commits. An old page that nothing changed under is quiet, so the team can leave it alone.
- Features with no docs. Each user-facing commit in 120 days is checked against the docs text. A flag or name that no page mentions becomes a "Document this" card.
- Proposal cards. Document this, refresh this page, feature this, and automate the release notes.
- What to ship next. One ranked list of up to 30 tasks. Each task has a score from 0 to 100, an effort of S, M or L, and one sentence that cites the commits and dates behind it. The score adds reader impact, urgency, evidence, whether Jev checked it, and ease, with fixed weights you can recompute by hand.
- Who works on what. A map of the people who made the last 12 months of commits and the parts of the code they work on. Each task suggests an owner and gives the reason, such as "Maya made 14 of the last 20 commits to docs/use-cases/".
- Drafts. Release notes for the newest stable tag and up to 3 product highlights, written from commit subjects and pull request numbers.
firstrun audit-issue turns the newest audit into issue.md, one checkbox for each of the top 10 tasks, and prints the gh issue create command. It opens nothing on its own. A suggested owner appears as a name and a login without the @, so nobody gets a GitHub notification.
Team map and privacy. The map reads public commits: author names, GitHub logins and avatars, commit counts and dates. It never shows an email, raw or hashed. Emails stay in memory to join one person's commits, and a sanitizer drops any email before a board is written.
Bots are left out. A commit counts by the files it changed, never by lines, and a commit 90 days old counts half. Every owner is a suggestion, not an assignment. Four opt-outs turn the map off or hide one person:
- Run
firstrun audit --no-team, or setFIRSTRUN_TEAM=0. - A repo owner adds
team: falseto.github/firstrun.ymlor.firstrun.yml. - An audit request ticks "Leave out the team map" on the issue form.
firstrun audit-submitleaves the map out unless you add--with-team. - A person adds
sha256:<hex>of their lowercase login or name tohiddenindocs/audits/team-optout.json. Their work moves into "Everyone else" and no task suggests them.
To audit every week, copy examples/github-action/firstrun-audit.yml into .github/workflows/. Each Monday it audits the repo and updates one issue labeled firstrun-audit, so the docs team gets a fresh work list without anyone running a command.
--docs-repo reads the docs from a second repo when the docs site lives apart from the code. Inside this repo, each audit is published to docs/audits/ and opens at audit.html.
A DevRel team ships a quickstart once and a dependency moves the week after. The first person to find the break is a new user, and most new users leave instead of filing a bug.
A clean pass ends as passed. A break ends as verified once a second container proves the fix. The nightly check costs $0.0178 per run on average over 88 runs, measured by neatlogs at list price.
FirstRun binary-searches the package's releases, one fresh container per probe, with the version pinned through PIP_CONSTRAINT. It replays stored commands, so a bisect makes no model calls.
Real result: the dbt DuckDB guide never worked in a clean environment. Its step jafgen --years 6 passes on jafgen 0.3.1 (2023-01-17) and fails on 0.4.6 (2024-04-07). PyPI has no release between them. The guide merged on 2025-03-14, 11 months after --years was gone. 5 probes found it.
FirstRun turns the container fixes of a run into a Prerequisites block. Then it proves the block. A fresh container with only those items runs the docs as written.
| Quickstart | The docs assume | Proof in a clean python:3.11-slim |
|---|---|---|
| AccuKnox knoxctl | curl, ca-certificates, gnupg, sudo | proven, 5/5 steps |
| MkDocs | curl | proven, 9/9 steps |
| Entire install | curl, git, ~/.local/bin on PATH | proven, 4/4 steps, git found in round 2 |
| dbt + DuckDB | Python 3.12 or later, git | steps 1 to 9 pass on python:3.12-slim, step 10 is the jafgen docs bug |
FirstRun prints the gh commands and a PR body. Every draft gets a source check first, because a docs note can look like a docs bug. FirstRun opened 4 real PRs this way. They are listed below.
Point FirstRun at any public quickstart. The report shows how long a new developer needs to reach a working result, and where they get stuck first.
</details> <details> <summary><b>6. Give your coding agent the loop.</b> Plan, run, recover, verify, with a trace.</summary> <br>The site hands you a prompt for Claude Code, Cursor or Codex. The agent installs FirstRun, checks Docker, runs your quickstart, proves the fix and opens the PR when you say yes. state.json is written after every step, so a run resumes with firstrun run --resume RUN_ID.
Radar Picks the Quickstarts Worth a Run, and the Run Log Picks the Buyers
FirstRun spends model money on every run, so it should only run where a break is likely and someone will read the fix.
- Find. Exa search returns candidate quickstart pages at no cost.
- Rank.
scripts/radar.pyasks Jev 4 typed questions per page in one call: can it run unattended in a container, is a step likely to break, how many developers read it, and is it the official docs. Code multiplies the answers into one priority. No model reads the pages one by one. - Run.
firstrun run --targets targets-radar.txt --tag radar --tier teamruns the top pages in 3 parallel lanes under the plan-set cost cap. Every run is logged on the runs page. - Sell.
scripts/buyers.pymaps each docs site FirstRun ran on or audited to its company, enriches it through treg at $0.0018 a lookup, and asks Jev whether the company fits the $49 to $299 plan. The result is business/buyers.md: companies FirstRun already has proof for, each linked to its replay. It holds company data only, with no people and no emails.
FirstRun Found 4 Real Docs Bugs and Opened 4 Upstream PRs
FirstRun ran five open-source quickstarts whose docs live in public repos. Every finding came from a verified run, where a second clean container passed with the fix.
<table> <tr> <td width="33%" valign="top"><a href="https://github.com/dbt-labs/docs.getdbt.com/pull/10149"><img src="docs/assets/pr-10149.webp" alt="Pull request 10149 on dbt-labs/docs.getdbt.com"></a><br><b>dbt-labs/docs.getdbt.com#10149</b><br>Fix the jafgen command in the DuckDB quickstart.</td> <td width="33%" valign="top"><a href="https://github.com/dbt-labs/docs.getdbt.com/pull/10150"><img src="docs/assets/pr-10150.webp" alt="Pull request 10150 on dbt-labs/docs.getdbt.com"></a><br><b>dbt-labs/docs.getdbt.com#10150</b><br>State the Python version the DuckDB quickstart needs.</td> <td width="33%" valign="top"><a href="https://github.com/entireio/cli/pull/2724"><img src="docs/assets/pr-2724.webp" alt="Pull request 2724 on entireio/cli"></a><br><b>entireio/cli#2724</b><br>Say where install.sh puts entire on Linux.</td> </tr> </table>| Quickstart | Result | What FirstRun found | Upstream |
|---|---|---|---|
| dbt Core with DuckDB | verified 9/10 | jafgen --years 6 fails because years is now positional | dbt-labs/docs.getdbt.com#10149 |
| same guide | verified | the cloned repo pins networkx==3.7, which needs Python 3.12, and the guide names no version | dbt-labs/docs.getdbt.com#10150 |
| Entire CLI install | verified 4/4 | install.sh puts entire in ~/.local/bin, which a clean shell lacks on PATH | entireio/cli#2724 |
| AccuKnox knoxctl | verified 5/5 | the page names 0.9.0 as latest while the installer fetches 0.9.58, and the APT steps need curl, gnupg and sudo | accuknox/help#671 |
| MkDocs getting started | verified 9/10 | the docs commands work, the container needed curl | none needed |
| Dagster quickstart | verified 10/10 | the docs already cover both gaps FirstRun hit: install create-dagster first, and pip 25.1 for --group | none needed |
On #10149, a bisect comment shows jafgen 0.4.6 dropped --years 11 months before the guide shipped. Replay any run in the browser at firstrun.atharvashah.com.
Built for #neatHack With Three Required Tools
<div align="center">A huge thanks to the neatlogs team for #neatHack. One weekend, one solo build, and three tools that each did one job. No FirstRun code existed before kickoff on 2026-10-10 at 12:30 PM IST.
</div> <table> <tr> <td width="33%" valign="top" align="center"> <a href="https://neatlogs.com"><img src="docs/assets/logo-neatlogs.webp" alt="neatlogs logo" width="56"></a> <h3>neatlogs</h3> <b>Traces every run.</b><br> The planner, executor and verifier show up as spans, so a bad run has a cause you can read. </td> <td width="33%" valign="top" align="center"> <a href="https://entire.io"><img src="docs/assets/logo-entire.webp" alt="Entire logo" width="56"></a> <h3>Entire</h3> <b>Records every session.</b><br> Each agent commit carries the Claude Code session that wrote it, so its lines trace to a prompt. </td> <td width="33%" valign="top" align="center"> <a href="https://cfo.ai"><img src="docs/assets/logo-cfoai.webp" alt="cfo.ai logo" width="56"></a> <h3>cfo.ai</h3> <b>Turns cost into a plan.</b><br> The measured cost per run from neatlogs goes into the pricing model. </td> </tr> </table>neatlogs Shows What the Agent Did
Each run is one firstrun.run trace. The span kinds are WORKFLOW, AGENT, TOOL and GUARDRAIL. Claude calls appear as typed LLM spans.
Two detections run on the traces: "docs step broke" and "planner invented a step". The second one flagged 3 planner spans in v1. Investigate traced them to the planner prompt and proposed a stricter one (evidence/neatlogs-investigation.md). A Claude Code session read the detection and its trend over the neatlogs MCP server and applied the fix in commit 3ca6f11. The fix adds code guards that drop invented, placeholder and fake-secret steps.
The integration check read 3 real runs back from persisted traces. See docs/neatlogs-integration.md.
Entire Records How the Build Happened
<table> <tr> <td width="33%" valign="top"><img src="docs/assets/t-entire-checkpoint.webp" alt="A commit message with an Entire-Checkpoint ID"><br><b>A checkpoint on each agent commit.</b> The commit maps back to the session that wrote it.</td> <td width="33%" valign="top"><img src="docs/assets/t-entire-gate.webp" alt="An Entire merge gate blocking a merge on a high finding"><br><b>A merge gate.</b> The merge stays blocked while a high finding is open.</td> <td width="33%" valign="top"><img src="docs/assets/t-entire-review.webp" alt="An Entire inline review finding on recovery.py"><br><b>Inline review.</b> The finding sits on the line in <code>src/firstrun/recovery.py</code>.</td> </tr> </table>- Checkpoints: 91 of 132 non-merge commits on main up to
f5d73ebcarry anEntire-Checkpoint:trailer, counted withgit log --no-mergeson 2026-10-11. The 41 without are the 4 setup commits before the first agent session and site commits from 2026-10-10 and 2026-10-11. On many of those the Entire commit hook hung and was stopped so the commit could land. Merge commits carry none, because Entire's hook skips merges. The neatlogs fix commit shows the full agent session behind it: checkpoint 01M4JGJ6NTE7YNB9Z0YW5FYVHN. - Trail reviews: Entire runners reviewed every layer on Trails 3 to 7. They raised 2 real bugs on Layer 1, a secret retry loop and a Windows-only Docker default, both fixed in e2aacb7.
- Trails: Trail 1 covers Layer 1. Trail pages need an entire.io login. The commit and session pages open without one.
- Graph: the build agents searched the repo through the graph. 4 graph calls saved about 11,662 tokens.
- Brain: it holds 7 durable facts, such as why recovery stops at 3 attempts. Its last refresh indexed 63 files and 1,056 symbols.
cfo.ai Prices the Nightly Check
<table> <tr> <td width="45%" valign="top"><img src="docs/assets/t-cfo-ari.webp" alt="Ari, the cfo.ai agent, greets Atharva on day one"><br><sub>Ari, the cfo.ai agent, on day one.</sub></td> <td valign="top">The business plan sets the agent's spending limit. firstrun econ reads 88 neatlogs traces and writes business/drivers.csv for cfo.ai. The mean run costs $0.0178, and the p95 run costs $0.0529. A 70% margin floor turns those costs into a cap per tier, and firstrun run --tier scale stops any run that spends past the cap.
| Tier | Price | Quickstarts | Cap per run | Runs over the cap |
|---|---|---|---|---|
| Free | $0 | 1, weekly | $0.0135 | 50.0% |
| Team | $49 a month | 10, nightly | $0.0470 | 9.1% |
| Scale | $299 a month | 100, nightly, with PRs, Bisect and Prerequisite Proof | $0.0279 | 26.1% |
The public cfo.ai model covers the four parts of the plan.
| Part | Figure | Source |
|---|---|---|
| Pricing | Free $0, Team $49 a month, Scale $299 a month | business/unit_economics.json |
| Costs | $0.01862 per run, which is the trace mean plus compute. Fixed costs are $25 a month, and Stripe takes 2.9% plus $0.30 per invoice | business/BUSINESS_PLAN.md |
| Revenue | $2,069 MRR in month 18 of the Base Scenario, and $5,427 in the Growth Scenario | docs/data/cfo-model.json |
| Runway | Cash starts at $5,000 on 2026-11-01 and ends month 18 at $16,685 in Base. Cash stays positive, so cfo.ai shows no runway limit | docs/data/cfo-model.json, business/BUSINESS_PLAN.md |
The old fixed cap of $0.50 let one bad month cost a Scale customer's fee five times over. The plan, the Scale fix and every assumption are in business/BUSINESS_PLAN.md. The cfo.ai model is public at https://cfo.ai/s/_u5lbrHIl9KX_JQE. Under the old cap, cash goes negative in month 14.
The plan page shows the caps, the measured cost spread and the cash per Scenario. Every doc-test run is tagged against the caps. A neatlogs alert uses the Scale cap as its threshold, and scripts/neatlogs_issues.py mirrors each detection and alert as a GitHub issue labeled neatlogs.
Every Run and Audit Ends With a Receipt
A receipt turns one unit of agent work into a priced, capped, verified and checkable record. A sample bill for the AccuKnox knoxctl run shows the parts:
- Priced. Model usage $0.0453 plus compute $0.0020 gives a total of $0.0473. The prices are list prices, not billed.
- Capped. Bars show the run against each tier's cap from the cfo.ai plan. The model cost is 96.4% of the Team cap and over the Scale and Free caps.
- Proven. A stamp says whether a second clean container reran the patched steps with no model calls. A run that ended any other way says "Not verified".
- Checkable. The receipt carries its own hash and the SHA-256 of
state.json,runbook.jsonandreport.md.
--check fails when the receipt or one of its evidence files changed after it was issued. firstrun receipt --all-published writes a receipt for every run and audit on the site, and publishing a run or an audit writes its receipt too. The Doc-test and Doc-audit pages show a "View receipt" button with the amount on every row.
A doc audit prices its Jev reads at TypeSafe's list price of $0.042 per million input tokens, and the clone and the analysis cost nothing. New audits record their Jev calls and tokens. The 7 audits published before cost tracking show an estimate, marked "est.", measured by re-running each one with scripts/audit_cost_estimates.py. The estimates run from $0.0001 (fastapi/typer) to $0.0055 (dbt-core).
The Build Happened in Public on X
Every step is a post on @cultist_dev. The five below tell the story in order.
<table> <tr> <td width="50%" valign="top">kickoff<br>
</td> <td width="50%" valign="top">#neatHack kickoff, 12:30 IST. FirstRun code before kickoff: 0 lines. neatlogs doctor probe: 7/7 PASS. Solo build. Submission opens Monday, 7 PM IST.
day one<br>
</td> </tr> <tr> <td width="50%" valign="top">#neatHack, day one. all three required tools are up. Entire: trail #1 created. neatlogs: first agent trace, 3 steps, 12s. cfo.ai: Ari read my public profile.
the pitch<br>
</td> <td width="50%" valign="top">your quickstart is the only part of your product tested by people who leave instead of filing a bug. so for #neatHack i'm building FirstRun. break. fix. prove.
the plan<br>
</td> </tr> <tr> <td colspan="2" valign="top">a quickstart breaks the way an agent breaks. quietly. so the plan for this weekend: every FirstRun run becomes a trace in @neatlogs.
the first catch<br>
</td> </tr> </table>my neatlogs hackathon project has already caught doc issues for a chroma (24K+ github stars). chromadb's deprecated-config error sends you to its own migration guide. that page returns 404.
Run Your First Check in 5 Minutes
You need:
- Python 3.11 or newer.
- Git.
pip install git+https://...clones the repo, so it fails without git. - Docker. On Windows, run Docker inside WSL with the
Ubuntu-22.04distro. - A Claude subscription token in
CLAUDE_CODE_OAUTH_TOKEN. Create one withclaude setup-token. - Optional:
NEATLOGS_API_KEYfor traces. - Optional:
TYPESAFE_API_KEYfor Jev. Without it, Claude classifies the failure.
Install:
FirstRun reads keys from .env in the folder you run it from. To use another file, set FIRSTRUN_ENV_FILE:
FirstRun uses docker from your PATH. On Windows without it, FirstRun calls Docker inside WSL with wsl -d Ubuntu-22.04 docker. Set FIRSTRUN_DOCKER to use another command.
Run it on a quickstart URL:
A URL tested in the last 24 hours returns the cached result. Add --fresh to run it again.
Read the result in runs/<run-id>/report.md. Build the HTML report with:
Run it inside a clone of this repo and every run is logged for the public site. FirstRun writes the replay to docs/runs/ and adds a row to docs/runs/index.json. Commit those files and the run appears on the runs page. Rebuild the log from the runs/ folder with:
The offline test suite needs no network, Docker or keys. The tests ship with the repo, not with the pip package, so clone it first:
FirstRun ran this section on itself in a clean python:3.11-slim container. The install failed until git was present, which is why git is listed above. The run step then stopped with the missing-token message, because a clean container has no Claude token or Docker.
One Trace Follows the Docs Page From Read to Verified Fix
A budget GUARDRAIL span runs before every recovery. The root span firstrun.run is a WORKFLOW span that holds the whole tree. docs/architecture.md lists the states, the failure classes, the budgets and every span.
Each of the 8 Focus Areas Has a Feature You Can Open
| Focus area | Feature | Where you see it |
|---|---|---|
| Planning | The planner turns the docs page into a runbook | runbook.json, the planner span |
| Tool use | Firecrawl reads, Docker runs, Jev classifies | One TOOL span per call |
| State | state.json is written after every step | firstrun run --resume RUN_ID |
| Context | The planner reads only the quickstart page | read_docs span input |
| Execution | One fresh container per run | step.N spans with exit codes |
| Recovery | Classify, search, patch, retry up to 3 times | recover.N span |
| Verification | A model-free re-run in a new container | verify span, status verified |
| Iteration | v1 traces, detections, a fix, v2 traces | The results table below |
<a id="version-2-finishes-3-of-4-quickstarts-and-recovers-3-failed-steps"></a>
Before and After, Finished Quickstarts Went From 2 of 4 to 3 of 4
v1 ran with recovery off. v2 ran after the neatlogs planner fix (3ca6f11), with recovery and verification on. The same four targets ran both times. Cost and duration come from the neatlogs traces in evidence/before_after_traces.json. Cost is the Sonnet 5.5 list price and is not billed, because the runs use a subscription token.
| Target | Before (v1, baseline, recovery off) | After (v2, neatlogs fix, recovery, verify) |
|---|---|---|
| neatlogs Python SDK | passed 3/3 | passed 3/3 |
| Entire CLI install | failed at step 1, curl: command not found | verified 4/4 after 2 fixes, one of them a real docs fix |
| Chroma getting started | passed 3/3 | passed 2/2 |
| OpenAI v0 README (fixture, Wayback 2023-08-02) | failed at step 6, openai: command not found | pinned openai==0.28.1, then stopped honestly: the example needs an OpenAI key |
| Metric | Before (v1, baseline) | After (v2) | Change |
|---|---|---|---|
| Quickstarts that end passed or verified | 2 of 4 | 3 of 4 | +1 |
| Failed steps recovered | 0 | 3 | +3 |
| Fixes proven in a fresh container | 0 | 1 | +1 |
| Average tool spans per run | 4.25 | 6.0 | +1.75 |
| Average cost per run, from neatlogs | $0.0262 | $0.0200 | -$0.0062 |
| Average run time, verification included | 128 s | 201 s | +73 s |
v2 takes longer because a verified fix runs the whole quickstart a second time in a new container.
The six mistakes the agent made in v1:
- The plan changes between runs. The same neatlogs page gave 4 steps, then 3 steps, 5 minutes later.
- The planner invents steps. It added
command -v entireandtest -f first_trace.py, which the docs never say. - The planner writes placeholder steps. Chroma step 2 is
true. - A fake secret passes as a step.
export OPENAI_API_KEY='sk-...'exits 0. - The planner skips the steps that matter, because they need a key.
- 9 of 16 steps check only the exit code.
A Quickstart Step Failed, and FirstRun Fixed It and Proved the Fix
The dbt DuckDB guide ran on 2026-10-10 as a 10-step plan. Step 1 failed because the container had no git, step 3 hit a networkx pin that needs Python 3.12, which the guide never states, and step 10 used a removed flag (jafgen --years 6). Steps 3 and 10 are the two docs bugs. FirstRun classed each failure, patched the command and retried it. Step 7, dbt docs serve, starts a web server, so FirstRun skipped it and said why. A fresh container then passed 9 of 9 runnable steps with no model calls. The run record is docs/examples/dbt.txt, a share code in the format of docs/share-format.md. Open the dbt run on the site.
- The step fails.
jafgen --years 6exits 2 withNo such option: --years. The loop in src/firstrun/executor.py lines 257 to 294 runs the step and sees the failure. Line 278 checks the budget before recovery starts. - Recovery classes the failure.
_resolvein src/firstrun/recovery.py lines 87 to 118 looks up the error signature inprecedents.jsonfirst. That signature was stored asdocs_bug, so the class came back with probability 1.0 and no model call. A new error goes to Jev, then Claude, then a regex table (src/firstrun/classify.py lines 128 to 140). Below 0.7 confidence, FirstRun asks the human once (recovery.pylines 98 to 118). - Recovery writes the patch. A docs bug gets one Claude call with WebSearch and WebFetch (recovery.py lines 385 to 396). It read the jafgen usage line and returned
jafgen 6, with the jaffle-shop-generator repo as its source. Recovery rejects a patch that keeps less than 60% of the original command (lines 254 to 260 and 284 to 287). - The step retries.
executor.pylines 293 and 294 store the patch and rerun the step with the new command.jafgen 6exits 0. Each step gets 3 attempts at most (src/firstrun/models.pyline 111,src/firstrun/budget.pyline 17). - A fresh container proves the fix. src/firstrun/verify.py applies the patches (lines 20 to 29) and names a new container (line 54). It runs every step inside
config.block_model_calls()(line 57), where a Claude or Jev call raisesModelCallBlocked. The run endsverifiedonly when every runnable step passes (lines 89 to 92). tests/test_verify.py line 89 fails ifverify.pyimports the planner, recovery or llm module.
The same run fixed 2 more steps. Step 1 lacked git, which is an environment failure, so recovery added apt-get install -y -qq git with no model call. Step 3 hit No matching distribution found for networkx==3.7, and Jev classed it as a docs bug at 0.87.
The neatlogs trace shows the same loop on the Entire install. evidence/trace-recovered-entire-v2.md exports trace 5520c5cff5e0ea61f130eced082b20c4. Step 1 failed with curl: command not found. Recovery reused the stored precedent (span 12), which classes that error as environment. The apt table in recovery.py lines 139 to 146 then added curl, and lines 370 to 376 built the patch. Step 2 failed until ~/.local/bin was on PATH (spans 14 to 23). The verify span (26) then ran all 4 steps in a fresh container (spans 27 to 30). evidence/neatlogs.md says why the trace status reads ERROR while the run ended verified.
What Broke While Building
- Git 2.33 hung Entire checkpoints. The post-commit hook waited 206 s on
git update-ref --stdin, and the checkpoint was never written. Git 2.55.0 fixed it, and the next hook took 3 s. - One run split into 3 traces. The first live run on the neatlogs docs produced 3 traces. The CLI now opens one
firstrun.runtrace around reading, planning, every step and verification. - Ask-once crashed background runs. The prompt raised
EOFErrorwith no terminal. It now returns nothing, and recovery uses the top class. - Tests exported junk traces. Each
pytestrun created about 50 emptyneatlogs-apptraces. The test fixture now sets an empty key, and a run after the fix created 0 traces. - A docs note looked like a docs bug. On Dagster, FirstRun labeled two gaps as docs bugs. A check of the source showed the page already covers both, in a prerequisite line and a note the planner skipped. FirstRun now opens no PR on Dagster. Every PR draft gets a source check before it ships.
- A cd broke steps three ways. A leading cd, a cd inside a patched chain, and a file check after cd each failed a step that worked. The container now treats a cd into the folder it is already in as a no-op, and file checks also look from the run root.
- The Entire Brain watcher does not run on Windows. Brain runs with
--no-daemonand is refreshed by hand.
Known limits are listed in docs/known-limits.md. Examples: a hung step is not retried, and version pinning covers pip only.
Parallel Agents Built It and Entire Reviewed Every Layer
- Build: a lead session planned the work and ran worker agents in parallel, one track each: fixtures, planner, executor, recovery, verification, report page, then edge cases. All agents ran on Sonnet 5.5 at medium effort.
- Review: each layer is an Entire Trail, so Entire runners reviewed every commit since kickoff. They found 3 real bugs, all fixed: a secret retry loop, a Windows-only Docker default, and a cd guard that was written but never called.
- CI: GitHub Actions runs the offline suite on every push and pull request. On the launch branch on 2026-10-11 the offline suite gave 384 passed, 68 skipped (they need
FIRSTRUN_LIVE=1or a human check) and 1 expected failure. - Jury self-test: the Entire judge plugin, run on this repo with the real start time, ranked it with no excluded work: 3.93 of 5, authenticity 4.5, integrity 5.0. The report is evidence/jury-selftest.txt.
Planning happened before the event in a private repo. The build log is hackathon.md.
<div align="center"> <br>Made solo by Atharva Shah for #neatHack by neatlogs.
</div>