Amazon ML Challenge 2026: the Number That Mattered Was 1.887
- Machine Learning
- Entity Resolution
- LightGBM
- Competition
Three sources describe the same businesses in the US, India and France, and they share no identifiers. One source is a clean reference list; the other two hold noisy copies with transliterated names, swapped legal forms and reordered addresses. Mixed in are decoys: records that copy a real business and change one detail, a house number or a legal form, and match nothing at all.
For every reference business, list the records that describe it. About 10 million records against 1.7 million businesses, scored with F0.5 per business:
F0.5 = 1.25 * TP / (1.25 * TP + 0.25 * FN + FP)
A wrong merge counts with weight 1 and a missed match with 0.25, so precision is worth four times recall. That single fact shaped every decision we made.
Team Inno8 started 28 hours into the 72-hour window and finished at 0.988642 public F0.5 across 8,269 teams, up from 0.940 on our first full-scale upload. All of it ran on one Apple M2 laptop with 16 GB of memory plus free Kaggle T4 and P100 GPUs.
The measuring stick was the real problem
Our first upload scored 0.9537 locally and 0.940 on the board. The second widened the gap to 0.0214. Local validation was rewarding something the test set did not have.
The answer was in the data. Train has 4.6765 records per business and test has 5.7543, with the same number of true matches. Unmatched records per business therefore rise by a factor of 1.887, and about 39.8% of test records match nothing against 26.0% in train. The test set is far more decoy-heavy, decoys cause wrong merges, and wrong merges are what F0.5 punishes.
Reweighting false positives by 1.887 in local validation is the change that made local gains and leaderboard gains move together. After that, every number we measured meant something.
The pipeline
- Normalize. A transliteration dictionary learned from the training pairs, 521 entries like
praivetto private. Exact agreement of transliterated Indian names with their reference name went from 15.3% to 89.5%. - Blocking in DuckDB on IDF-weighted rare keys, ten candidates per record, 19.2 million test pairs. The true entity survives for 98.13% of pairs and ranks first for 97.99%.
- Two LightGBM stages, the first on 60 pair features, the second adding record and entity context plus 28 odd-one-out features.
- Cross-encoders on the uncertain band only: four fine-tuned multilingual E5 models over 902,470 pairs, combined by a LightGBM stacker.
- Decision rules. One owner per record, a learned decoy-word veto, and expected-F0.5 selection per business.
- France, which appears only in test, with no labels, another language and another address format: a two-model guard, a fixed threshold and rules gated by hand-read audits.
Seven things worth keeping
- Measure what the test measures. The most useful number of the whole challenge was 1.887.
- Never train on synthetic exact duplicates. A model learns whatever separates synthetic rows from real ones. Our twin copies looked like +0.003 and were really -0.0017.
- Spend the expensive model only where it matters. The cross-encoders read 902,470 uncertain pairs, not 19.2 million, which is what made them affordable on a laptop.
- A weaker model can still lift a stack. Every cross-encoder scored below stage 2 alone, and each one we kept still raised the stacked AUC, 0.9714 to 0.9834.
- Recall was the bottleneck, and blocking is its ceiling. Missed matches were worth several times the false positives, so rescue passes and reverse blocking went in to recover part of what blocking dropped.
- Without labels, be conservative and read the data. The one French change we could model came in at about a third of its central estimate.
- Late gains need honest error bars. With the last upload deciding the private ranking, the final changes were scored on the half of the slice that had not picked their thresholds. The prediction was +0.00008 and the result was +0.000093.
What did not work
Per-country quantile normalization: no gain. Self-training on confident predictions: no gain. A 95-feature stacker instead of 8: +0.00004. A fifth cross-encoder: slightly worse than four. Broader French change sets: failed their audits.
The largest remaining loss is records with no address whose name is shared by many businesses, chain branches for instance. That is about 0.005 of local F0.5, and nothing in the data separates the branches.
The habit that carried over from solver work is the same one that mattered here: when your local measurement and reality disagree, stop tuning and go find out why they disagree.