Fix SerpAPI Result Ordering in Global Planning

Read AGENTS.md, docs/HANDOFF.md section 30, docs/HANDOFF2.md, and the audit findings before editing.
Confirm the current branch is global-live-experience and the working tree is clean. Stop if either is false. Do not push, merge, deploy, or publish. Preserve unrelated changes and never print secrets.
Fix the critical problem where SerpAPI result ordering chooses the recommendation.
Requirements:
- In global planning discovery, stop retaining only
_first_place()for each intent. - Normalize a bounded candidate pool of up to five valid places per intent from Maps results.
- Deduplicate candidates using place_id first, data_id second, and a conservative normalized name/address fallback.
- Reject candidates that are clearly outside the resolved destination.
- Keep candidates with missing optional fields, but represent those fields as unknown.
- Perform deterministic preliminary ranking using only provider evidence already present in the search response.
- Keep at least two viable candidates per intent when available so Python can compare alternatives.
- Enrich candidates in rank order and remain inside the eight-request allowance:
- At most two intent searches.
- Details only when required for feasibility or final comparison.
- Reviews only when they can affect an explicitly requested constraint.
- Supporting web search only for critical missing evidence.
- If a provisional finalist is definitely closed, continue to the next candidate while budget remains.
- Final selection must be made in Python after comparing eligible candidates. Provider ordering must only be a stable final tie-breaker, never the main score.
- Preserve deterministic behavior and stable tie handling.
- Update cache types so the normalized candidate pool, not only already-selected places, can be reused.
- Do not introduce another provider, model API, database, or paid service.
Add regression tests proving:
- The first provider result is not selected when a later candidate has stronger verified fit.
- A definitely closed first result is skipped.
- Two candidates with equal evidence use a documented stable tie-breaker.
- Up to two intents still produce at most two stops.
- Discovery cannot exceed eight billed calls.
Update docs/HANDOFF2.md and docs/BUILD_LOG.md with factual results. Run backend formatting, lint, tests, and the secret scanner.
Commit exactly:
Rank a bounded set of live candidates.
Stop after the commit and report the hash, changed files, test totals, and git status.
I'll start by reading the handoff docs and audit findings, then confirm the branch and working tree before changing anything.
The sandbox cannot run here, so I’ll recheck git with full permissions and read section 30 plus the audit findings.
The branch is global-live-experience and the tree is clean. Next I’ll locate the audit findings and the discovery code that currently picks the first result.
SerpApi order is still the recommendation because discovery keeps only the first place. I’ll rank a bounded candidate pool and let Python choose, with provider order only as a tie-breaker.
The plan tests were returning a pin far from Tokyo, so the new destination check correctly dropped it. I’ll put the scripted results on the searched coordinates.
The commit is c82f8c0b3da0c569e7abee4e83c9f1b149fa9215 on global-live-experience. The working tree is clean. Nothing was pushed, merged, deployed, or published.
Python now compares up to five live places per intent. Provider order is only the tie-breaker when the listed evidence is equal; an equal rank then uses the casefolded name and the place id. A pin more than 80 km from the destination is dropped. A finalist whose hours are definitely closed is replaced while the eight-request budget remains. The cache stores that candidate pool.
| Check | Result |
|---|---|
ruff format --check | 75 files already formatted |
ruff check | All checks passed |
pytest | 247 passed |
| Secret scan | Exit 0, no findings |
Changed files:
backend/src/happen_api/api/plans.pybackend/src/happen_api/planning/discovery.pybackend/src/happen_api/planning/itinerary.pybackend/src/happen_api/planning/limits.pybackend/tests/contract/test_plans.pybackend/tests/unit/test_discovery.pybackend/tests/unit/test_itinerary.pybackend/tests/unit/test_limits.pydocs/BUILD_LOG.mddocs/HANDOFF2.md
git status reports global-live-experience with nothing to commit.
The Hacktoberfest Weekend Challenge, Build for a Friend, is open until October 5, 2026 at 6:59 AM UTC. I can draft a DEV post about this ranking fix if you want one.
Community Wisdom: Hybrid Search and Re-Ranking: The Cheapest Quality Win
Source: internals_decoded Tags:
ai,hybridsearch,reranking,bm25A first-stage search should return a bounded candidate set, and a second stage should re-score that set. The post treats the original retrieval order as incomplete. I did not read the comment thread.
Community Wisdom: How to Test Search Relevance Before You Ship a Ranking Change
Source: libme Tags:
postgres,database,testing,sqlA ranking change needs fixed cases that fail when the wrong result rises, including a case where the first hit is not the right one. The new tests cover a stronger later place, a closed first result, an equal-evidence tie, two stops, and the eight-call cap. I did not read the comment thread.
Read the project instructions and current progress first. Confirm the branch is global-live-experience and the tree is clean. Do not push or deploy.
Connect the editable planning brief to deterministic eligibility and scoring. The current application collects constraints but discards most of them.
Requirements:
- Extend the v2 plan request and matching frontend Zod/TypeScript contracts to carry:
- party size
- budget
- preferences
- accessibility needs
- Pass the complete validated constraint set through discovery and itinerary assembly.
- Introduce a typed constraint assessment for every selected stop:
- met
- unmet
- unknown
- not_applicable
- Every assessment must name the constraint and point to supporting evidence when status is met or unmet.
- Unknown evidence must never become a positive score.
- Accessibility requirements are safety-sensitive:
- Recommend a place as verified accessible only when provider evidence supports it.
- If accessibility is requested but unknown for every candidate, return an explicit insufficient-evidence or partial result rather than claiming fit.
- Dietary requirements such as vegetarian or vegan must follow the same evidence rule.
- Budget amount must not be compared with nonnumeric price symbols as though they were equivalent.
- A provider price tier may be compared only with a requested budget tier.
- A numeric amount may be evaluated only against a compatible numeric price with the same currency.
- Party size should be retained and shown. If no capacity/reservation evidence exists, mark suitability unknown.
- Preferences such as quiet, romantic, outdoor seating, casual, and live music may influence scoring only through deterministic, source-backed signals.
- Add transparent deterministic scoring components. Do not expose a raw internal score to the customer.
- Hours feasibility must remain more important than soft preference fit.
- Python remains the sole selector. Gemma remains disabled because its quality gate failed.
Update the result contracts and UI to explain which important constraints were verified and which remain unknown.
Add tests proving:
- A verified quiet candidate can outrank a louder or unknown candidate.
- Unknown quietness does not receive a positive score.
- An unmet accessibility requirement cannot be presented as a fit.
- Budget, party size, accessibility, and preferences survive frontend-to-backend serialization.
- Changing a constraint can change the selected candidate without changing the raw provider evidence.
- Unsupported constraints produce honest unknown states.
Run backend and frontend unit checks, type/build checks, and the secret scanner.
Commit exactly:
Apply planning constraints to deterministic scoring.
Stop and report the commit, scoring rules, test totals, and git status.
Commit 83f56a0c4f86e33e3fe200d02d5fd860600ee190 on global-live-experience: Apply planning constraints to deterministic scoring. The working tree is clean. Nothing was pushed or deployed.
Party size, budget, preferences, and accessibility needs now travel with the plan request. Python scores them from retrieved text. Gemma is still unused.
Scoring rules
- Open hours count as 100. Unknown hours count as 20. A closed place is dropped before selection.
- A listed website adds 4, a Maps link adds 2, and a rating adds the rating times 10, rounded.
- A constraint adds 4 only when retrieved text verifies it. The constraint total is capped at 16, so a preference cannot outrank hours.
- Unknown evidence adds 0. An unmet soft preference, such as a loud room when quiet was requested, also adds 0.
- Wheelchair access, step-free access, a hearing loop, vegetarian, and vegan are hard. An explicit contradiction is ineligible. If a requested one is not verified on the selected stop, the outcome is
insufficient_evidenceand the stop is not marked verified. - A price symbol is compared only with a requested budget tier. A numeric amount is compared only with a numeric price in the same currency: exact, at most, or about (within 25%). A symbol against an amount, or a different currency, stays unknown.
- Party size is returned and shown. Capacity language such as “seats 8” can be met or unmet. With no capacity evidence, suitability stays unknown and adds 0.
- A preference with no deterministic signal, such as “dog friendly,” stays unknown and adds 0.
- Equal evidence uses the earlier provider result, then the casefolded name, then the place id.
- The customer payload names each check as supports, unknown, blocks, or neutral. It does not include a numeric score.
Checks
- Backend: ruff format and ruff check passed. Pytest: 258 passed.
- Frontend: Biome passed. Vitest: 41 passed.
tsc --noEmitand the Vite build passed. - Browser: the landing suite passed, 9 tests, including a quiet preference change that rescores without resolving the destination again and shows that quiet was not verified.
- Secret scan: exit 0, no findings.
Community Wisdom: A Confidence Score Is Not a Probability: Act, Ask, or Abstain
Source: raju_dandigam Tags:
ai,machinelearning,testing,architectureA missing required fact should make the system ask or abstain. An uncalibrated number is not a reason to act, and an unknown outcome should stay unknown rather than be counted as success.
Community Wisdom: Memory Plus Rules: Designing Trustworthy Decision Analysis Without an LLM
Source: dikshith__00117766f65 Tags:
ai,webdev,productivity,softwaredevelopmentTransparent rules should name the source of each finding. Silence from a rule is not evidence that the requirement is satisfied.
Those two posts agree on the rule used here: a missing check is not support. I did not read the comment threads.
The Hacktoberfest Weekend Challenge, Build for a Friend, includes a SerpApi category. Submissions were due 2026-10-05 06:59 UTC.
I can draft a DEV post about this scoring change if you want one. I will not publish it unless you ask.