Demo Experience UI Copy and Layout Refinements

Cursor·grok-4.7-high·kernelKain·yesterday·21min·5 Checkpoints·22 file changes·+616/-84·510.5K tokens

PROMPT 2 — Plan Execution and Delivery

ROLE

You are my senior engineering partner executing a locked implementation plan.

Prompt 0 selected and shaped the product. Prompt 1 completed and locked the specification, architecture, scope, stack, acceptance criteria, phase plan, and numbered execution queue. Prompt 2 executes that plan.

Optimize for:

  • A working Hook
  • Required sponsor integration
  • A publicly reachable submission
  • Reliable fixture fallback
  • Passing acceptance criteria
  • Completion before feature freeze and deadline
  • Portfolio-grade implementation quality
  • The smallest complete product, not the largest feature set

Do not reopen settled product or planning decisions without new evidence.

The numbered execution queue in HANDOFF.md Section 24 is the default work queue. The change policy in Section 25 controls permitted deviations.

RESPONSE CONTRACT

Each reply may contain only one of:

  • A request for the missing HANDOFF.md
  • A request for one genuinely blocking fact or approval
  • A resume sheet followed by execution of the next authorized unit
  • The result of one supervised step
  • The result of an autonomous phase
  • One packet part when output must be split
  • A blocker or failed-verification report
  • A phase-completion handoff

Do not provide speculative implementation essays before working.

End every reply with:

CHECK: questions_this_turn={0|1} step={id|none} autonomy={A1|A2|A3} packet={k/m|none} stage=C

REQUIRED INPUT

Execution requires the complete locked HANDOFF.md produced by Prompt 1.

If HANDOFF.md is missing, ask only:

“Please paste the complete HANDOFF.md produced by Prompt 1. A repository URL may accompany it, but it does not replace the locked plan.”

Then stop.

Do not ask whether the project is new. Do not ask for the idea. Do not create a new implementation plan. Do not infer the execution queue from a README.

If the HANDOFF was lost but a repository exists, direct me back to Prompt 1 continuation mode to reconstruct and lock a new handoff.

HARNESS DETECTION

Determine the harness without asking me:

HARNESS=agent Repository, file, shell, git, and related tools are available.

HARNESS=paste You cannot directly inspect or modify the repository.

State the harness in the first execution reply.

For HARNESS=agent:

  • Inspect and edit files directly
  • Run relevant local commands and tests
  • Use specialized file tools when available
  • Show concise summaries instead of printing complete files
  • Respect autonomy restrictions on git and external actions

For HARNESS=paste:

  • Request only the files or command output required for the active step
  • Return complete files, never fragments
  • Use the packet protocol when needed
  • Give exact commands for me to run
  • Wait for actual output before diagnosing failures

AUTHORITY AND TRUST ORDER

Use this authority order:

  1. My latest explicit instruction
  2. Official hackathon rules
  3. Locked product, architecture, scope, and acceptance criteria in HANDOFF.md
  4. Current repository and deployment evidence
  5. Execution state recorded in HANDOFF.md
  6. Previous assistant summaries
  7. Assumptions

Repository evidence is authoritative for what currently exists and works.

HANDOFF.md is authoritative for:

  • Intended product
  • Locked architecture
  • Stack
  • Scope
  • Acceptance criteria
  • Phase plan
  • Execution queue
  • Change policy

If git and HANDOFF.md disagree:

  • Do not choose one silently
  • State the exact discrepancy
  • Use git for current implementation facts
  • Use HANDOFF for intended plan
  • Ask one question only if the discrepancy changes the next action

Never overwrite or discard user changes to make the repository match the handoff.

EXECUTION BOUNDARY

Allowed:

  • Implementing an authorized execution-queue step
  • Reading repository and deployment state
  • Running local builds, tests, linters, and formatters
  • Installing locked dependencies when required by the active step
  • Making small implementation decisions inside locked architecture
  • Debugging failures caused by the active step
  • Updating uncommitted HANDOFF.md execution state
  • Performing actions allowed by the current autonomy level
  • Applying the Section 25 change policy

Forbidden unless explicitly approved:

  • Changing the product, user, problem, or Hook
  • Adding user-facing scope
  • Removing required acceptance criteria
  • Replacing required sponsor technology
  • Adding an external API
  • Adding another backend language
  • Changing architecture boundaries
  • Adding paid services
  • Increasing the spend cap
  • Changing privacy or data assumptions
  • Moving feature freeze later
  • Upgrading outside the locked toolchain policy
  • Destructive git or data operations
  • Exposing secrets
  • Implementing unrelated cleanup

RESUME PROCEDURE

When HANDOFF.md is present:

  1. Parse the locked plan and execution state.
  2. Inspect the repository when HARNESS=agent.
  3. Determine the current branch, status, commits, and existing files.
  4. Check whether the recorded last step is actually complete.
  5. Locate the next incomplete execution-queue step.
  6. Calculate remaining budget only from known timing information.
  7. Identify any user changes that must be preserved.
  8. Determine current autonomy.
  9. Print the resume sheet.
  10. Immediately execute the next authorized unit unless blocked.

Do not wait for “ok” before the first unit unless it requires:

  • A secret
  • Paid spend
  • Destructive action
  • Merge or deployment without A3 authority
  • Missing architecture-changing information
  • Resolution of conflicting user changes
  • A scope-changing decision

RESUME SHEET

Keep the resume sheet to one screen:

  • Project
  • Product thesis
  • Locked Hook
  • Harness
  • Autonomy
  • Repository and branch
  • Working-tree state
  • Last completed step
  • Next step
  • Current phase
  • Deadline
  • Feature-freeze status
  • Remaining known budget
  • Live URL
  • Fixture status
  • Largest current risk
  • Quality items already dropped
  • Any HANDOFF/repository discrepancy

Then begin the next authorized unit.

AUTONOMY LEVELS

A1 — SUPERVISED EXECUTION

For HARNESS=agent:

  • Execute one numbered step per reply
  • Read and edit files directly
  • Run local verification
  • Do not commit, push, open PRs, merge, or deploy
  • Stop after reporting the step result

For HARNESS=paste:

  • Provide complete files and exact commands for one step
  • Wait for me to apply and verify them

A2 — BRANCH AND PR

  • Execute the remainder of the authorized phase without routine confirmation
  • Use the branch strategy locked in HANDOFF.md
  • Edit, test, commit, push, and open or update a PR
  • Do not merge
  • Do not deploy to production
  • Preview deployment is allowed only if explicitly authorized in HANDOFF
  • Stop only for an abort condition or required approval

A3 — MERGE AND DEPLOY

  • Includes all A2 authority
  • May merge using the locked merge strategy
  • May deploy or redeploy
  • Must verify the live URL
  • May roll back its own deployment when verification fails
  • Does not authorize destructive data operations or spend above the cap

AUTONOMY COMMANDS

Interpret these commands:

“ok” Execute the next A1 step.

“A1” or “hands-on” Use supervised one-step execution.

“autonomous” Use A2 for the remainder of the current phase.

“autonomous A3” Use A3 for the remainder of the current phase.

“autonomous phase [id]” Use A2 for the named phase.

“phase [id] autonomy: A1|A2|A3” Set autonomy for that phase.

“stay autonomous” Preserve the current A2 or A3 level for subsequent phases.

“pause” Stop safely after the current atomic operation and update execution state.

“switching harness” Produce a complete updated HANDOFF.md and Resume Block, then stop.

“changed: [description]” Treat my described edits as new repository evidence and inspect them.

“add step: [description]” Request an explicit addition to the execution queue.

“skip step [id]” Evaluate its acceptance-criteria impact before skipping it.

“stop” Stop without beginning another action.

After an autonomous phase, revert to A1 unless I said “stay autonomous” or configured autonomy for the next phase.

PREFLIGHT SAFETY CHECK

Before changing files, inspect:

  • Current repository root
  • Current branch
  • Git status
  • Untracked files relevant to the step
  • Active step and dependencies
  • Locked toolchain
  • Relevant existing files
  • Acceptance criteria covered by the step
  • Available remaining time
  • Known user edits

If the working tree contains unrelated changes:

  • Preserve them
  • Do not overwrite, revert, stash, or commit them without approval
  • Limit edits to the active step
  • Flag overlap only if the active step touches the same lines or files

Never use destructive commands such as:

  • git reset --hard
  • git clean -fd
  • force push
  • destructive database reset
  • deleting user data

unless I explicitly approve the exact action.

EXECUTION UNIT

The execution unit depends on autonomy:

A1: One numbered step from HANDOFF Section 24.

A2/A3: The remainder of the active phase, including only its authorized steps and permitted change-policy substeps.

PACKET MODE: One complete packet part.

Never mix unrelated phases in one unit.

If the active step is already complete:

  • Verify its done-when conditions
  • Mark it complete with evidence
  • Continue only within the current authorized unit

STEP EXECUTION LOOP

For every execution-queue step:

  1. ORIENT
  • Read the step contract
  • Read its dependencies
  • Read relevant files
  • Identify acceptance criteria
  • Check the locked decision and toolchain entries
  1. PROVE FIRST

For risky integrations or assumptions:

  • Run the smallest safe proof first
  • Verify access, compatibility, or contract behavior
  • Do not build dependent layers before the proof succeeds
  • Use the planned fallback when the proof fails
  1. IMPLEMENT
  • Make the smallest complete change satisfying the step
  • Follow existing repository conventions
  • Keep architecture boundaries locked
  • Avoid unrelated refactors
  • Do not leave pseudocode, placeholders, or hidden TODOs
  • Do not fake live success
  1. VERIFY

Run verification from narrowest to broadest:

  • Targeted test
  • Relevant type or static check
  • Relevant integration or contract test
  • Build
  • Broader suite when proportionate
  • Browser or deployment verification when required

Use actual command output.

Do not claim a command passed unless it was run successfully or I supplied successful output.

  1. REVIEW

Inspect the resulting diff for:

  • Scope creep
  • Secrets
  • Debug output
  • Placeholder data
  • Unhandled failure states
  • Accidental toolchain changes
  • User changes accidentally included
  • Acceptance-criteria gaps
  1. RECORD

Update execution state with:

  • Step status
  • Evidence
  • Files changed
  • Verification run
  • Decisions made
  • New risks
  • Actual time if known
  • Next step
  1. REPORT

Report only information needed to understand, verify, or continue the work.

CHANGE POLICY

Follow HANDOFF.md Section 25.

WITHOUT APPROVAL, you may:

  • Perform read-only reconnaissance
  • Debug within the active step
  • Make implementation-detail choices preserving contracts
  • Resolve dependency patch versions within the locked version rule
  • Fix behavior required by an existing acceptance criterion
  • Add a necessary emergency substep under the active parent ID

Emergency substep format:

[parent-id]a — [specific necessary correction]

An emergency substep must:

  • Serve an existing acceptance criterion
  • Stay inside the locked architecture
  • Add no user-facing scope
  • Add no paid service or external API
  • Be recorded in HANDOFF.md
  • Include a concise reason and verification

APPROVAL IS REQUIRED before:

  • Product or Hook changes
  • New user-facing capabilities
  • Architecture-boundary changes
  • Required acceptance-criteria removal
  • External API additions
  • Backend-language additions
  • Paid-service additions
  • Spend-cap increases
  • Data/privacy changes
  • Sponsor-technology replacement
  • Feature-freeze extension
  • Destructive action

When approval is required, ask one focused question containing:

  • Evidence
  • Impact
  • Recommended choice
  • Safe fallback

BRANCH AND GIT RULES

Follow the branch strategy locked in HANDOFF.md.

Do not assume every phase requires a separate branch if Prompt 1 selected a lighter workflow for a short build.

When a phase branch is required:

  • Use its exact planned name
  • Reuse it for every step in that phase
  • Do not create a branch per step
  • Do not rename it without a real conflict
  • Branch from the planned base after confirming it is current

Before committing:

  • Inspect the complete diff
  • Exclude unrelated user changes
  • Exclude secrets and local environment files
  • Exclude HANDOFF.md when the handoff policy says it is uncommitted
  • Run required checks
  • Use the commit convention locked in HANDOFF

Never:

  • Force-push without explicit approval
  • Rewrite shared history
  • Amend a user commit
  • Commit unrelated working-tree changes
  • Claim a PR is green without checking its current status

HANDOFF.md FILE POLICY

docs/HANDOFF.md is an execution-state artifact.

Default policy:

  • Keep it local and uncommitted
  • Ensure it is ignored by git
  • Never include secrets
  • Update it after each completed step when practical
  • Print the complete updated contents at phase boundaries, interruptions, harness switches, or when I request it

If the locked HANDOFF explicitly chooses a committed planning document:

  • Preserve the planning sections in the committed document
  • Keep volatile execution state in a separate ignored handoff
  • Never commit secrets or local-only account details

Do not silently change the handoff policy.

DEPENDENCY AND TOOLCHAIN RULES

  • Use the locked language, framework, package manager, and version policy
  • Do not perform broad upgrades
  • Do not regenerate lockfiles unnecessarily
  • Do not introduce a second backend runtime
  • Do not add a library when existing capabilities are sufficient
  • Check licenses for newly required dependencies
  • Record every added direct dependency and its purpose
  • Use exact lockfile-resolved versions
  • If a locked package is unavailable, gather evidence before proposing change
  • Security patch changes outside the locked rule require approval unless the current version cannot safely or successfully build

SECRETS AND EXTERNAL SERVICES

Never:

  • Ask me to paste a secret into chat
  • Print secret values
  • Add secrets to source files
  • Commit real environment files
  • Put server secrets in frontend bundles
  • Expose secrets in logs, screenshots, commands, or HANDOFF.md
  • Spend money beyond the locked cap
  • Create external resources outside the current autonomy

When a secret is needed:

  • Ask me to configure the named environment variable locally or in the host
  • Give exact location or dashboard instructions
  • Continue after I confirm configuration
  • Verify only presence or successful authenticated behavior

Environment examples must contain empty or clearly fake values.

LIVE AND FIXTURE PATHS

The live integration and fixture fallback must preserve the same visible user flow whenever hackathon rules permit.

Requirements:

  • Fixture data must be original or licensed
  • Synthetic data must be visibly labeled
  • Fixture mode must not pretend to be a live response
  • Live and fixture outputs must follow the same internal contract
  • Fixture mode must work without secrets
  • The Hook characterization test must cover the shared behavior
  • A sponsor integration required by rules must still be genuinely implemented and demonstrable; fixtures cannot replace eligibility requirements

DEBUGGING AND FAILURE POLICY

Diagnose from actual evidence.

For a failing check:

  1. Capture the exact error.
  2. Form one evidence-based hypothesis.
  3. Apply the smallest correction.
  4. Rerun the narrow failing check.
  5. Run broader checks after it passes.

Do not perform two speculative fixes in succession without new evidence.

After two unsuccessful verify-fix cycles for the same failure:

  • Stop modifying that area
  • Preserve all user work
  • Report the exact commands and errors
  • Explain the two attempted fixes
  • State the most likely cause
  • Recommend the planned fallback or one next diagnostic
  • Ask one focused question only if user input is required

Do not automatically revert to an earlier commit if doing so might discard user or unrelated work.

A safe rollback may be used only when:

  • The changes being rolled back were created entirely in the current unit
  • No user changes are included
  • The rollback is non-destructive
  • The reason is recorded

ABORT CONDITIONS

Stop and ask before proceeding when:

  • A required secret is missing
  • Paid spend would exceed the cap
  • A destructive action is necessary
  • Force-push appears necessary
  • Merge or deployment is required without A3
  • User edits conflict with the active step
  • Official rules contradict the plan
  • The required sponsor capability does not exist
  • A must-build acceptance criterion appears infeasible
  • A data or privacy assumption is unsafe
  • The next action changes locked scope or architecture
  • Continuing would cross feature freeze with new feature work

Do not stop for routine reversible implementation choices.

CLOCK AND FEATURE FREEZE

Use actual known time. Do not invent elapsed hours.

After each phase, report:

  • Planned phase budget
  • Known actual active time
  • Remaining build hours
  • Deadline status
  • Feature-freeze status
  • Buffer status

At feature freeze:

Allowed without additional approval:

  • Bug fixes
  • Acceptance-criteria completion
  • Tests
  • Reliability improvements
  • Security fixes
  • Accessibility fixes
  • Deployment fixes
  • Documentation
  • Submission artifacts

Not allowed without approval:

  • New features
  • New screens
  • New integrations
  • Architecture rewrites
  • Cosmetic work that risks the Hook

If behind schedule:

  1. Remove Could Build work
  2. Remove incomplete Should Build work not supporting the Hook
  3. Apply the locked drop order
  4. Preserve every Never Drop item
  5. Use documented fallbacks
  6. Report what was cut and why

Do not quietly extend the schedule.

QUALITY FLOOR

Never drop:

  • Working Hook
  • Required sponsor integration
  • Fixture fallback where permitted
  • Public URL when required
  • Secrets hygiene
  • Must-be-true claim
  • Hook characterization test
  • Recoverable error behavior
  • Complete required submission artifacts
  • Updated execution handoff
  • 1280px layout correctness

Keep until feature freeze when relevant:

  • Shared design tokens
  • Loading state
  • Empty state
  • Partial-result state
  • Error state with next action
  • Success state
  • Visible keyboard focus
  • Accessible contrast
  • Provenance labels
  • Responsive behavior
  • Reduced-motion support
  • Deployment smoke test

Drop first when behind:

  1. Optional motion
  2. Extra screenshots beyond the hero
  3. Secondary filters or views
  4. Optional analytics
  5. Blog polish
  6. Non-required integrations
  7. Mobile visual polish while keeping mobile functional
  8. All Could Build features

Never ship:

  • Lorem ipsum
  • Unexplained fake metrics
  • Fake users or testimonials
  • Unlabeled synthetic data
  • Debug logs
  • Missing focus styles
  • Inaccessible color-only state indicators
  • Accidental default component-library styling
  • Praise-copy instead of product information

PERCEIVED PERFORMANCE

When UI is involved:

  • Give visible feedback within approximately 100ms
  • Reserve layout space to avoid large shifts
  • Use a layout-matching loading state
  • Stream or progressively reveal genuinely long operations when supported
  • Provide cancellation or retry when relevant
  • Avoid fake progress indicators
  • Keep the Hook path visually clear

A UI step is incomplete until its planned states and viewport checks pass.

TESTING POLICY

Tests must protect behavior and risk, not chase coverage percentages.

Prioritize:

  1. Hook characterization
  2. Sponsor/live contract
  3. Fixture parity
  4. Core transformations
  5. Failure behavior
  6. API boundary
  7. End-to-end demo path
  8. Deployment smoke path

When behavior changes:

  • Add or update the smallest test proving the behavior
  • Keep existing relevant tests green
  • Do not delete a failing test merely to pass CI
  • Do not weaken assertions without evidence that the contract changed
  • Record intentionally deferred test gaps in HANDOFF.md

BROWSER AND UI VERIFICATION

When browser tools are available and the step affects UI:

  • Verify the actual rendered application
  • Check relevant viewport sizes
  • Inspect loading, empty, error, partial, and success states
  • Check browser console errors
  • Exercise the keyboard path for the Hook
  • Capture required screenshots only when scheduled
  • Verify visible sponsor attribution when required

Do not claim visual correctness from source inspection alone.

When browser tools are unavailable:

  • Provide exact run and verification steps
  • Request only the resulting screenshot, console output, or behavior needed

DEPLOYMENT RULES

Deployment requires A3 or explicit deployment authorization.

Before deployment:

  • Run required tests and build
  • Confirm environment-variable names
  • Confirm no secrets are in client output
  • Confirm deployment target and spend
  • Confirm health and smoke-check commands
  • Confirm rollback path

After deployment:

  • Verify the health endpoint
  • Verify the public URL
  • Exercise the Hook
  • Exercise or confirm fixture fallback
  • Check the deployed browser console
  • Record the deployment URL and evidence
  • Update HANDOFF.md

A URL existing is not proof that the current Hook works.

PULL REQUEST RULES

When A2 or A3 requires a PR, include:

  • What changed
  • Why it changed
  • Execution step IDs
  • Acceptance criteria covered
  • Verification commands and actual results
  • Screenshots when required
  • Known limitations
  • Risk or fallback notes
  • AI-use disclosure required by the project

Do not claim CI is green until current checks complete successfully.

Do not blanket-accept automated review suggestions.

For each review comment:

  • Confirm it applies
  • Fix valid correctness, security, or requirement issues
  • Reject irrelevant or scope-expanding suggestions with a reason
  • Rerun affected verification

PACKET PROTOCOL

Use packet mode only for HARNESS=paste or when user-visible output would otherwise truncate.

Start each packet with:

PART [k]/[m] Steps: [step IDs] Files: [files included]

Then provide:

  • Complete files
  • Exact commands
  • Verification instructions

End each packet with:

END PART [k]/[m] Reply “next” for PART [k+1]/[m].

Rules:

  • Split only on file boundaries
  • Never split one file across packets
  • Make no new design decisions between packets
  • Do not repeat completed files
  • Do not mark the step complete until all packets are applied and verified
  • Agent harness edits files directly and does not print full files merely for visibility

A1 OUTPUT CONTRACT

For one supervised step, report:

STEP [ID] — [Name] Phase: [phase] Autonomy: A1 Status: COMPLETE | BLOCKED | PARTIAL

Goal: [One concise sentence]

Changes:

  • [file or system area]: [what changed and why]

Verification:

  • [command/check]: [actual result]
  • Acceptance criteria: [IDs and status]

Decisions:

  • [implementation detail decided, if material]

Risks or limitations:

  • [only real remaining concerns, or NONE]

HANDOFF update:

  • Completed: [step ID or NONE]
  • Next: [step ID]
  • Clock/freeze status: [known status]

For HARNESS=paste, include complete files and commands using the packet protocol when necessary.

Stop after one step and wait for “ok,” unless I authorized another mode.

A2/A3 OUTPUT CONTRACT

For an autonomous phase, begin with:

PHASE [ID] — [Name] Autonomy: [A2|A3] Authorized steps: [IDs]

Then execute without routine progress questions.

At completion report:

  1. Phase outcome
  2. Steps completed
  3. Files and systems changed
  4. Verification and actual results
  5. Acceptance criteria completed
  6. Branch and commits
  7. PR status
  8. Deployment and live verification for A3
  9. Risks, fallbacks, and deferred work
  10. Clock and feature-freeze status
  11. Phase exit checklist
  12. Complete updated HANDOFF.md
  13. Learning document
  14. Next step

Stop only for abort conditions or required approval.

PHASE EXIT CHECKLIST

At every phase boundary, mark each item:

DONE SKIPPED — [reason] NOT APPLICABLE — [reason] BLOCKED — [reason]

Checklist:

  1. Planned phase steps resolved
  2. Phase done-when conditions satisfied
  3. Relevant acceptance criteria verified
  4. Hook characterization test passes
  5. New behavior tests pass
  6. Build, lint, and type checks pass where configured
  7. Live/fixture parity checked where relevant
  8. UI states and viewport checks completed where relevant
  9. No secrets or debug artifacts in diff
  10. Dependency and toolchain policy preserved
  11. Required documentation reflects current reality
  12. AI-use disclosure updated when required
  13. Branch and commit policy followed
  14. PR created or updated when authorized
  15. CI status checked when applicable
  16. Automated review comments triaged when applicable
  17. Deployment verified when authorized
  18. HANDOFF.md execution state updated
  19. Clock, cuts, and freeze status recorded
  20. Next step is unambiguous

Do not mark the phase complete when a required done-when condition remains unmet. Mark it BLOCKED or use the documented fallback.

LEARNING DOCUMENT

At the end of each A2 or A3 phase, provide a concise learning document.

Learning Document — Phase [ID]

Material Commands

List material commands in execution order:

  • Command with secrets redacted
  • What it did
  • Why it was needed
  • Result

Do not include every trivial navigation or read command.

Files Changed

For each meaningful file:

  • Path
  • NEW, MODIFIED, MOVED, or DELETED
  • What changed
  • Why it changed

How the Phase Changed the Project

Explain:

  • Capability now available
  • Connection to earlier work
  • What the next phase depends on
  • Side effects or limitations

Failures and Recoveries

Include:

  • Important failure
  • Evidence
  • Root cause
  • Fix or fallback
  • Lesson for later phases

Write “None” if there were no meaningful failures.

Review Notes

List 1–3 honest trade-offs or concerns a reviewer should know.

The learning document:

  • Is for understanding and continuity
  • Must not contain secrets
  • Is not automatically committed or submitted
  • Should remain concise

JUDGE-DAY RUNBOOK

The phase assigned by HANDOFF must produce and verify:

  • 60-second script with timestamps
  • Cold-start procedure
  • Demo-data reset
  • Live sponsor path
  • Fixture fallback path
  • 1280px laptop path
  • Required login preparation
  • Network-failure fallback
  • Public-URL failure fallback
  • Backup recording
  • Hero screenshot
  • Required submission links
  • Final five-minute smoke checklist

The runbook must reflect actual behavior, not planned behavior.

INTERRUPTION AND HARNESS SWITCH

When stopping mid-phase because of a blocker, pause, or harness switch:

  1. Finish or safely stop the current atomic file operation.
  2. Do not begin another execution step.
  3. Record partial work accurately.
  4. Record commands run and verification state.
  5. Record uncommitted files.
  6. Record the exact blocker.
  7. Set the next action.
  8. Print the complete updated HANDOFF.md.
  9. Print the Resume Block.
  10. Stop.

Do not mark a partially completed step DONE.

COMPLETION CONDITION

The project is complete only when:

  • Every required acceptance criterion passes
  • The Hook works
  • Required sponsor technology is demonstrably used
  • Fixture fallback works where permitted
  • Public deployment works when required
  • Tests and smoke checks pass
  • Required documentation is accurate
  • Demo and submission artifacts are complete
  • Judge-day runbook is verified
  • No secret or critical blocker remains
  • Final HANDOFF execution state says COMPLETE

At completion, provide:

  • Final outcome
  • Public URL
  • Repository/PR/release links when available
  • Acceptance-criteria summary
  • Verification evidence
  • Known limitations
  • Submission checklist
  • Final HANDOFF.md

Do not add celebratory filler or propose new features.

START HERE

If HANDOFF.md is missing:

  • Ask only for the complete Prompt 1 HANDOFF.md
  • Stop

If HANDOFF.md is present:

  • Detect the harness
  • Reconstruct repository state
  • Print the resume sheet
  • Execute the next authorized unit immediately unless blocked

Do not ask for the idea. Do not redesign the plan. Do not begin an unauthorized step. Do not claim success without verification.

CHECK: questions_this_turn=1 step=none autonomy=A1 packet=none stage=C

I'll locate the locked handoff and the repository state, then continue from the next incomplete step.

The progress log and the locked plan disagree on how far the build has gone. I'll check git and the next step contract before changing anything.

The working tree is clean on demo-experience, ahead of the progress log. I'll check whether the deploy step is actually done before starting the next unit.

The sponsor slice is merged and the deploy step is still open. I'll verify that revision locally, then stop where deployment needs you.

Local checks passed on the merged sponsor revision. Deployment is still yours, so I’ll record that and stop before the next phase.

HARNESS=agent

Project: Happen
Product thesis: Pick one evidence-backed restaurant and arrival window for an Indiranagar evening.
Locked Hook: Three restaurant timelines resolve across the evening until one interval is the recommended visit, with separate fit and confidence and a distinct fallback.
Harness: agent
Autonomy: A1
Repository and branch: kernelKain/happen on demo-experience at 91b7b94 (origin/main). No upstream.
Working-tree state: Clean at the start of this turn. This turn only updated docs/HANDOFF2.md and docs/BUILD_LOG.md.
Last completed step: P3.5 — Evaluate and optionally tune Gemma. Untuned 270M stayed; the extraction gate was missed.
Next step: P3.6 — Deploy and smoke the sponsor vertical slice.
Current phase: P3 — live sponsor Hook. Code is merged. Public smoke is not done.
Deadline: October 5, 2026 at 06:59 UTC (12:29 PM IST). About 13 hours 40 minutes of calendar time remain. Measured build hours are not recorded.
Feature-freeze status: Not reached.
Remaining known budget: $0 spent. Render 1c-2g is $25/month from the existing $50 credits and does not sleep.
Live URL: Not deployed.
Fixture status: Captured Indiranagar fixture is on main and is what Find the moment scores when live mode is off.
Largest current risk: There is still no public URL, and a real submit still returns three timelines and no winner because the model kept no review spans.
Quality items already dropped: None.
HANDOFF / repository discrepancy: docs/HANDOFF.md section 27 still says the build has not started. That section is the locked planning snapshot. docs/HANDOFF2.md plus git are the progress record. The progress log still said live-sponsor and an open pull request 6. Git shows pull request 7 merged, and demo-experience is origin/main with no commits of its own. live-sponsor (d05d024) and main (91b7b94) contain the same files.

STEP P3.6 — Deploy and smoke the sponsor vertical slice
Phase: P3 — live sponsor Hook
Autonomy: A1
Status: BLOCKED

Goal:
Confirm the merged sponsor slice is safe to deploy, then stop because creating Render services is not allowed at A1.

Changes:

  • docs/HANDOFF2.md: recorded the merged revision, the local checks, and the Render steps still waiting on you.
  • docs/BUILD_LOG.md: recorded the same readiness check. No application code changed.

Verification:

  • uv run pytest in backend: 119 passed.
  • npm test in frontend: 7 passed.
  • python3 scripts/scan-secrets.py: passed.
  • GitHub Actions run 37219473378 on main: succeeded.
  • Acceptance criteria: AC-04 through AC-14 and AC-17 are not met in production. The local fixture and live routes are present; the public smoke has not run.

Decisions:

  • Leave HAPPEN_LIVE_ENABLED false in render.yaml until you turn it on for one public live run.
  • Do not start the demo-experience work while this deploy is still open. demo-experience has no commits yet.

Risks or limitations:

  • render.yaml still names branch live-sponsor. That matches main today. After the services exist, point both at main.
  • The local .env has SERPAPI_API_KEY set. It was not printed and it is not in the blueprint.

HANDOFF update:

  • Completed: NONE. P3.6 is blocked on Render.
  • Next: P3.6, your deploy.
  • Clock/freeze status: feature freeze not reached. Deadline is 12:29 PM IST on October 5.

Your side:

  1. In Render, create a Blueprint from this repo’s render.yaml. Do not create a second copy by hand. After the services exist, point both at main.
  2. When Render prompts, set SERPAPI_API_KEY from the local ignored .env. Leave HF_TOKEN unset.
  3. Leave HAPPEN_LIVE_ENABLED false for the first check. Open the frontend URL and https://<api-host>/healthz. Paste those two URLs back here. Do not paste the key.
  4. For one bounded live run, set HAPPEN_LIVE_ENABLED to true, redeploy the API, and submit once from the public page. Then paste whether the page showed SerpApi sources, the Gemma version, and a live or captured label.
  5. Suspend happen-api after that smoke if you are not ready to leave it running. The static site is free.

CHECK: questions_this_turn=0 step=P3.6 autonomy=A1 packet=none stage=C

Execute the selected diff-tab commit action.

I'll inspect the branch diff and recent commit style, then commit only the notes from this step.

Committed the deploy-readiness notes on demo-experience as 1a9803b with the message “Record that the merged sponsor slice is ready and still waiting on Render.”

The commit includes only docs/BUILD_LOG.md and docs/HANDOFF2.md. The working tree is clean, and nothing was pushed.

Possible Play: modiqo/list-top-committers — say search Plays to review matches.

ok

You confirmed the next supervised step. I'll read the result-state contract and the current page, then implement only that step.

The public deploy is still waiting on you. The next supervised step is the visible result states, and the local fixture path is enough to build those.

Playwright is looking in a sandbox cache that does not have Chromium. I'll point it at the browser that is already installed and rerun the tests.

The browser tests ran against the previous production build. I'll rebuild the frontend and run those checks again.

HARNESS=agent

Project: Happen
Locked Hook: Three restaurant timelines resolve to one evidence-backed visit, with separate fit and confidence and a distinct fallback.
Autonomy: A1
Branch: demo-experience, latest commit 1a9803b. Result-state edits are uncommitted.
Last completed step: P4.1 — Complete all visible result states
Next step: P4.2 — Finish evidence and methodology experience
Live URL: Not deployed. P3.6 is still waiting on Render.
Largest current risk: No public URL, and a real submit can still return no winner.

STEP P4.1 — Complete all visible result states
Phase: P4 — complete demo experience
Autonomy: A1
Status: COMPLETE

Goal:
Give every required result state specific copy and a next action.

Changes:

  • frontend/src/app/App.tsx: loading tells the user to wait. A model that is waking up or unavailable is named, and the unavailable state retries metadata. Quota and invalid-input errors hide Retry. Invalid fields are listed.
  • frontend/src/app/result/MomentResult.tsx: a partial result keeps three timelines and says Unknown intervals are not a recommendation. An insufficient result says to start over and restore the demo preset. A captured fixture says it is saved evidence, not a live search.
  • frontend/src/lib/api/recommendation.ts: the client now keeps retryable, field errors, and retry delay.
  • frontend/src/lib/resultStates.ts: names the public error states.
  • Tests cover partial, quota, invalid input, and an unavailable model.
  • docs/HANDOFF2.md and docs/BUILD_LOG.md: recorded the step. The public deploy is still open.

Verification:

  • npm test: 9 passed.
  • npm run check and npm run build: passed.
  • npx playwright test: 16 passed, including the Hook, partial, quota, invalid input, model unavailable, and the 1280 and 390 shell checks.
  • The built page at http://127.0.0.1:4173/?layout=sample showed Courtyard Lantern as the recommended moment, North Gallery Supper as the fallback, and three timelines. The local API was not running, so that browser stayed on the service check. Playwright covered the new clicks.
  • Acceptance criteria: AC-08 and AC-12 are covered locally for these states. AC-15 was checked on the existing initial and error states, not on every new panel.

Decisions:

  • Retry stays off when the server marks the error as not retryable. The captured-fixture button stays a separate action.

Risks or limitations:

  • Render is still not created, so these states are verified locally only.

HANDOFF update:

  • Completed: P4.1
  • Next: P4.2
  • Clock/freeze status: feature freeze is not reached. Deadline remains 12:29 PM IST on October 5.

CHECK: questions_this_turn=0 step=P4.1 autonomy=A1 packet=none stage=C

Briefly inform the user about the task result and perform any follow-up actions (if needed). If there's no follow-ups needed, don't explicitly say that.

The local preview on port 4173 stopped because it was shut down after the result-state check. The result-state work is still in the working tree on demo-experience, and the next step is the evidence view.

Execute the selected diff-tab commit action.

I'll review the branch diff and commit only the result-state work.

Committed the result-state work on demo-experience as 88e080f with the message “Give each result state specific copy and a next action.”

The commit includes the planner, result view, error parsing, tests, and the two progress notes. The working tree is clean, and nothing was pushed.

Possible Play: modiqo/list-top-committers — say search Plays to review matches.

OK

The next supervised step is the evidence view. I'll read the current panel against that contract and finish only what it still lacks.

The evidence panel already shows quotes and source links. I'll add the methodology, conflict treatment, rejected-span explanation, and a keyboard close.

HARNESS=agent

Project: Happen
Locked Hook: Three restaurant timelines resolve to one evidence-backed visit, with separate fit and confidence and a distinct fallback.
Autonomy: A1
Branch: demo-experience. Latest commit is 88e080f. Evidence-panel edits are uncommitted.
Last completed step: P4.2 — Finish evidence and methodology experience
Next step: P4.3 — Finish reveal, responsive, and reduced motion
Live URL: Not deployed. P3.6 is still waiting on Render.

STEP P4.2 — Finish evidence and methodology experience
Phase: P4 — complete demo experience
Autonomy: A1
Status: COMPLETE

Goal:
Make Why this moment show the quotes, conflicts, rejected spans, model version, and scoring method, and keep it usable from the keyboard.

Changes:

  • frontend/src/app/result/MomentResult.tsx: the panel now explains that Gemma extracts quotes and Python scoring chooses the moment. It names the model, adapter, scoring policy, and extraction schema, and it repeats the planning disclaimer. A priority that both supports and conflicts says so. Rejected quotes are counted and excluded from the score. Escape closes the panel and returns focus to Why this moment.
  • frontend/src/lib/evidenceView.ts: ordinary http and https links stay; other addresses are not linked.
  • frontend/src/styles/global.css: source links are underlined.
  • frontend/tests/hook.spec.ts: a keyboard test opens the panel, checks the methodology and a conflict, and closes it with Escape.
  • docs/HANDOFF2.md and docs/BUILD_LOG.md: recorded the step.

Verification:

  • npm test: 11 passed.
  • npm run check and npm run build: passed.
  • npx playwright test: 17 passed. The open panel had no serious accessibility violations.
  • Acceptance criteria: AC-03, AC-06, and AC-10 are represented in the panel. AC-15 passed for that open state in Playwright.

Decisions:

  • The evidence view stays an inline panel. A separate drawer was not added.

Risks or limitations:

  • Render is still not created, so this was verified in the local browser tests only.

HANDOFF update:

  • Completed: P4.2
  • Next: P4.3
  • Clock/freeze status: feature freeze is not reached. Deadline remains 12:29 PM IST on October 5.

CHECK: questions_this_turn=0 step=P4.2 autonomy=A1 packet=none stage=C

Execute the selected diff-tab commit action.

Committed the evidence-panel work on demo-experience as f88e056 with the message “Explain quotes, conflicts, and scoring inside Why this moment.”

The commit includes the panel, source-link checks, tests, and the two progress notes. The working tree is clean, and nothing was pushed.

Possible Play: modiqo/list-top-committers — say search Plays to review matches.

OK

The next supervised step is the reveal, the two required viewports, and reduced motion. I'll check what the page already does, then finish only that.

The result appears all at once, and the narrow layout stacks every time band into one column. I'll add a short reveal that reduced motion skips, and keep the mobile timelines as compact cards.

The reveal was fading text in, which dropped contrast while it was moving. I'll keep the short slide and leave the text fully visible.

HARNESS=agent

Project: Happen
Locked Hook: Three restaurant timelines resolve to one evidence-backed visit, with separate fit and confidence and a distinct fallback.
Autonomy: A1
Branch: demo-experience. Latest commit is f88e056. Reveal edits are uncommitted.
Last completed step: P4.3 — Finish reveal, responsive, and reduced motion
Next step: P4.4 — Conduct the friend walkthrough
Live URL: Not deployed. P3.6 is still waiting on Render.

STEP P4.3 — Finish reveal, responsive, and reduced motion
Phase: P4 — complete demo experience
Autonomy: A1
Status: COMPLETE

Goal:
Show the three timelines with a short reveal, keep the selected interval obvious, and make 1280×720 and 390×844 usable, including when reduced motion is on.

Changes:

  • frontend/src/styles/global.css: the three rows slide in over about half a second. The selected interval keeps its border, the word Selected, and a static amber outline. Reduced motion shows the rows immediately and does not move them. At 390 pixels each restaurant timeline is a two-column card.
  • frontend/tests/result.spec.ts: checks six columns at 1280, two columns at 390, evidence on the narrow screen, and an immediate result under reduced motion.
  • docs/HANDOFF2.md and docs/BUILD_LOG.md: recorded the step.

Verification:

  • npx biome check .: passed.
  • npm run build: passed.
  • npx playwright test: 18 passed, including 1280×720, 390×844, and reduced motion.
  • Amber text on the raised surface is about 6.3:1. A fade was removed because it made that text fail contrast while it was moving.
  • Acceptance criteria: AC-02 and AC-16 passed in these viewport checks. AC-15 stayed green on the existing accessibility tests.

Decisions:

  • No pulsing glow. The slide does not hide the text.

Risks or limitations:

  • Render is still not created. This was verified in the local browser tests.

HANDOFF update:

  • Completed: P4.3
  • Next: P4.4
  • Clock/freeze status: feature freeze is not reached. Deadline remains 12:29 PM IST on October 5.

Look at the local page at 1280 pixels wide, with the API running, or open /?layout=sample. Say if the recommended moment is obvious. The slide can be removed. The three timelines stay.

CHECK: questions_this_turn=0 step=P4.3 autonomy=A1 packet=none stage=C

Execute the selected diff-tab commit action.

Committed the reveal and layout work on demo-experience as 00b29a3 with the message “Reveal the timelines briefly and keep the narrow layout usable.”

The commit includes the styles, viewport tests, and the two progress notes. The working tree is clean, and nothing was pushed.

Possible Play: modiqo/list-top-committers — say search Plays to review matches.

OK

The next supervised step is the friend walkthrough. I'll read what it requires and prepare only the part that does not invent feedback.

The walkthrough has to come from a real person. I'll write the five-minute script and stop before any feedback is invented.

HARNESS=agent

Project: Happen
Locked Hook: Three restaurant timelines resolve to one evidence-backed visit, with separate fit and confidence and a distinct fallback.
Autonomy: A1
Branch: demo-experience. Latest commit is 00b29a3. The walkthrough script is uncommitted.
Last completed step: P4.3 — Finish reveal, responsive, and reduced motion
Next step: P4.4 — Conduct the friend walkthrough
Live URL: Not deployed. P3.6 is still waiting on Render.

STEP P4.4 — Conduct friend walkthrough
Phase: P4 — complete demo experience
Autonomy: A1
Status: BLOCKED

Goal:
Get one consented, identity-free paraphrase about fit, confidence, and the fallback.

Changes:

  • docs/HANDOFF2.md: recorded the five-minute script and the rule that no wording changes until a paraphrase arrives.
  • docs/BUILD_LOG.md: recorded that no feedback was invented.
  • No page copy changed.

Verification:

  • No friend session was run from here.
  • Acceptance criteria: AC-19 is not met until you send the paraphrase.

Decisions:

  • Use Find the moment, not /?layout=sample. That sample is labeled as unscored. A real submit can still show three timelines and no winner, and that is a valid walkthrough result.

Risks or limitations:

  • The walkthrough cannot be completed from this chat.

HANDOFF update:

  • Completed: NONE. P4.4 is blocked on the paraphrase.
  • Next: P4.4, after you send it.
  • Clock/freeze status: feature freeze is not reached. Deadline remains 12:29 PM IST on October 5.

Start the app from the repository root:

Open the local address Vite prints, usually http://127.0.0.1:5173, at about 1280 pixels wide.

  • 0:00 — Say this, then stop explaining: "I'm testing a dinner planner. I will not record your name. I will keep only a paraphrase of what you say. You can stop at any time."
  • 0:20 — Ask them to plan an Indiranagar dinner for easier conversation between 6:00 PM and 9:00 PM, then choose Find the moment. Do not point at Fit, Confidence, or the fallback.
  • 2:30 — Ask, and write their meaning rather than a transcript: Did Fit and Confidence seem like different ideas? Would this help you choose a restaurant and a time? Was the fallback a different usable restaurant, or was it missing or unclear?
  • 4:30 — If a result is on screen, ask them to open Why this moment from the keyboard and say whether the quotes explain the choice.
  • 4:50 — Stop.

Send the paraphrase here without a name, contact detail, or identifying quote.

CHECK: questions_this_turn=1 step=P4.4 autonomy=A1 packet=none stage=C

OK

The walkthrough still needs your paraphrase, so no wording will change from it. I'll start the security and failure-path step and leave that feedback unrecorded.

Security checks for size, CORS, rate limits, and checksums are already in place. I'll close the remaining gaps: untrusted page text, source-link hosts, and non-JSON requests.

HARNESS=agent

Project: Happen
Locked Hook: Three restaurant timelines resolve to one evidence-backed visit, with separate fit and confidence and a distinct fallback.
Autonomy: A1
Branch: demo-experience. Latest commit is 00b29a3. Security and walkthrough notes are uncommitted.
Last completed step: P5.1 — Harden security and failure paths
Next step: P5.2 — Meet performance and reliability gates
Live URL: Not deployed. The friend paraphrase is still missing, and no wording was changed for it.

STEP P5.1 — Harden security and failure paths
Phase: P5 — harden and freeze
Autonomy: A1
Status: COMPLETE

Goal:
Close the security checks that were not already covered, without printing or storing a secret.

Changes:

  • backend/tests/contract/test_demo_recommendations.py: plain text and file uploads get HTTP 422 and are not echoed.
  • frontend/src/lib/evidenceView.ts: a source link is clickable only for https on example.com, google.com, www.google.com, or maps.google.com.
  • frontend/tests/hook.spec.ts: a name or quote that looks like HTML is shown as text, and an unapproved source is not linked.
  • render.yaml: the static site now blocks embedded objects and framing. connect-src stays unset until the public API host exists, so the policy does not block the API.
  • docs/HANDOFF2.md and docs/BUILD_LOG.md: recorded the step.

Verification:

  • uv run pytest: 120 passed.
  • python3 scripts/scan-secrets.py: passed.
  • npm test: 11 passed.
  • Playwright Hook and result tests: 14 passed, including the hostile-text case.
  • Acceptance criteria: AC-06, AC-12, AC-17, and AC-18 are covered by these checks plus the existing CORS, size, rate-limit, idempotency, checksum, and extraction tests. AC-19 is still open.

Decisions:

  • No wildcard was added to the content security policy.
  • This stayed on demo-experience because the walkthrough is still open on that branch.

Risks or limitations:

  • The friend paraphrase has not arrived.
  • Render is still not created.

HANDOFF update:

  • Completed: P5.1
  • Next: P5.2
  • Clock/freeze status: feature freeze is not reached. Deadline remains 12:29 PM IST on October 5.

CHECK: questions_this_turn=0 step=P5.1 autonomy=A1 packet=none stage=C

Execute the selected diff-tab commit action.

Committed the security-boundary work on demo-experience as 3d1cc1d with the message “Reject non-JSON posts and limit which source links can open.”

The commit includes the request checks, source-link limits, the static-site policy, tests, and the progress notes. The working tree is clean, and nothing was pushed.

Possible Play: modiqo/list-top-committers — say search Plays to review matches.