Overhaul Matching Scorer with Profile Signals

ONLY-12: Matches feel random. Make matching smarter.
Three users said their top match made no sense. The scorer in src/match/score.ts is language overlap plus a hash of the two handles, and the hash is doing most of the work. Make scores come from real signals in the profiles: stack, working hours and timezone, review tone, PR size, pairing, indentation. Every match must explain itself with top reasons and one deal-breaker. npm run eval must put at least 17 of 20 labelled couples in the expected band, and npm test must pass. Evaluate at least two different approaches before choosing one, and say why you chose it. Commit when done.
Done. ONLY-12 is committed on main in three commits, the last being b7b1562. The scorer no longer touches handles, the eval is 20 of 20, and all 13 tests pass.
Results
| Check | Before | After |
|---|---|---|
npm run eval in band | 8/20 | 20/20 |
| Mean band error | 0.75 | 0 |
npm test | 8 pass | 13 pass |
What changed in src/match/score.ts
- Eight signals, each read from the profiles and scored 0 to 1: shared stack, working hours (local peak commit hour blended with chronotype), timezone gap, review tone, PR size ratio, pairing, indentation, and commit cadence. A shared favourite error adds a small bonus.
- Hard clashes carry a flat cost on top of the lost weight: tabs versus spaces, opposite chronotypes, and a PR size ratio of 4x or more.
- Every match returns up to three reasons, strongest first, and exactly one deal-breaker: the signal that cost the most points. Even soulmates get one, so the card never shows the "none found" fallback. Scores are symmetric and unchanged if you rename a profile or its handle.
Two approaches evaluated
- Additive (chosen): score is the weighted sum of signal similarities minus clash costs. It placed all 20 labelled couples in band, and each reason maps to a fixed number of points, so the card can explain it honestly.
- Gated: soft signals (stack, tone, PR size, pairing) multiplied by an hours-and-timezone gate. It plateaued at 17 or 18 of 20. The multiplicative gate sinks priya and yara, who are 11.5h apart but labelled soulmates, and it makes each reason's contribution depend on every other signal, which is harder to explain.
Weights were chosen by a constrained search where every signal the ticket names keeps a meaningful weight, then rounded to integers summing to 100. The comment at the top of the scorer records the comparison.
Worth knowing
- Three eval couples sit right on a band edge: noor and priya at 75, yara and ines at 74, kofi and mei at 50. The labels themselves force this, since yara and ines beat noor and priya on nearly every signal except indentation and distance. Any future weight change should be checked against the eval.
- Timezone gap is a modest weight (7) and carries no flat cost, because priya and yara at 11.5h are labelled soulmates. A gap over 8h is already phrased as the deal-breaker when it is the weakest signal, so ONLY-13 has a hook to build on but will need to reconcile that label.
- The README still shows the v0 eval JSON as its example. I left it, since it documents the format rather than the numbers.