Understanding F1 Scores in the Amazon ML Challenge
The mathematical foundations of the competition evaluation metric: precision, recall, harmonic mean, and how leaderboard scores were computed.
Why the F1-Score?
In automated product feature extraction, class imbalance is a primary challenge. A simple accuracy metric would heavily reward models that predict only the majority class.
The F1 Score represents the harmonic mean of Precision and Recall:
$$\text{F1} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}$$
Precision vs. Recall Trade-Off
- Precision: Of all entities predicted by the model, what proportion was actually correct?
- Recall: Of all ground-truth entities present in the catalog dataset, what proportion did the model successfully identify?
A model that hallucinates extra attributes suffers in precision; a model that is overly conservative misses valid attributes and suffers in recall.
Impact on Leaderboard Dynamics
Throughout the competition, many teams experienced substantial score increases (+0.300 to +0.500) simply by tuning probability classification thresholds rather than modifying model weights.
By centering the decision threshold to maximize harmonic balance, teams climbed hundreds of leaderboard ranks.
Read more about competition data handling on our Methodology Page.