Amazon MLChallenge
Explorer
Technical Documentation & Data Pipeline

Dataset & Ranking Methodology

How competition records were collected, normalized, and processed into institutional metrics and inter-college analysis.

1. Data Collection & Completeness Verification

The underlying data was collected from the public competition leaderboard endpoint of the Amazon ML Challenge (Challenge ID: 290344) on Unstop:

  • Pagination Scope: All 114 pages were exhaustively fetched and archived (100% complete coverage).
  • Total Records: 4,334 student teams, 14,480 registered team member profiles, and 1,546 participating academic institutions.
  • Immutability: The dataset is frozen as a historical competition snapshot, ensuring consistency across all query and ranking operations.

2. College & Organization Normalization

Self-reported college names entered during registration often contain spelling variations, differing acronyms, or punctuation variances. The platform applies an automated normalization pipeline:

1. Lowercase conversion and Unicode character stripping.

2. Punctuation removal and whitespace consolidation.

3. Canonical slug generation (e.g., indian-institute-of-technology-iit-madras).

4. Inverted index mapping each canonical slug to its associated teams and student rosters.

3. College Leaderboard Metrics & Calculations

The College Leaderboard calculates multi-dimensional performance benchmarks across participating universities:

Overall Composite Score

Weighted composite formula considering the institution's best team rank (40%), top-5 average rank (30%), top-10 team placements (20%), and participation scale (10%).

Top-5 Average Rank

Arithmetic mean of the top 5 highest-ranking teams affiliated with the college, rewarding institutions with consistent top-tier depth.

Best Team Rank (Top-1)

The highest single leaderboard standing achieved by any team from the college.

Score Distribution

Statistical quartiles, median score, and histogram buckets across all teams affiliated with the college.

4. Inter-College Team Detection

A team is classified as an Inter-College Team if its registered student members report 2 or more distinct normalized academic institutions. The algorithm:

  1. Iterates over the player roster of each team.
  2. Extracts and normalizes each member's organization string.
  3. Counts the cardinality of distinct organizations. Teams with cardinality >= 2 are cataloged into the inter-college directory.

Raw Data vs Derived Metrics

Raw competition records (team names, member rosters, competition scores, submission timestamps, and official team ranks) are exact representations from the public contest snapshot. College rankings and inter-college classifications are derived analytical metrics computed specifically for this exploration tool.