Now live · free & open

Football data, built for machine learning.

Gamblistics is a curated dataset of football (soccer) matches, teams, lineups and per-player statistics - designed from the ground up to train models that predict prediction-market outcomes.

Public. Community-first. Continuously growing.

What's in the dataset

Match-level and player-level data collected from live football matches, joined by stable UUIDs.

Match identity & results

Teams, kickoff time, league, tournament, stage, status, final and half-time scores.

Team match statistics

Possession, shots (on / off / blocked), corners, xG, passes, cards, offsides, fouls, saves - split by overall, first half, second half.

Lineups

Starting XI, substitutes, formations, jersey numbers, tactical positions, minutes on/off, captains, goalkeepers and coaches.

Per-player statistics

~100 metrics per player per match - xG, shots, passing accuracy, duels, ratings and more. Top leagues, 2024–25 season onward.

Built for ML, not for scraping

Most public football data is raw, messy and hard to join. Gamblistics fixes that.

ML-ready Parquet

Delivered as columnar Parquet files with a documented schema. Load straight into pandas, Polars, DuckDB or Spark.

Deterministic joins

One canonical UUID per entity - match, team, player, league, tournament. Joins across tables are trivial and stable over time.

Multi-granularity

Match-level, team-level (per period) and per-player metrics in one coherent, aligned dataset.

prediction-market targets in mind

Designed to support calibrated probability models for final score (Poisson λhome/λaway), 1H score, total corners, over/under 1.5 & 2.5 goals, time of first goal.

Live & growing

Continuously refreshed from live matches through our collection pipeline. New matches, players and tournaments land regularly.

Open to the community

Free to download and use. We want researchers, students and quants building on top of it - no gatekeeping.

Scale today

Approximate figures - growing every week.

1,100+matches
1,300+players
110+tournaments
750+matches with team stats
440+matches with per-player stats

A note on coverage. Per-player statistics are biased toward top leagues from the 2024–25 season onward. Match, team and lineup coverage is broader. We're being upfront so you can plan your modeling accordingly.

Schema at a glance

Every entity has a stable UUID. Every table joins on those UUIDs.

Reference entities

  • countries
  • leagues
  • tournaments
  • seasons
  • stages
  • teams
  • players

Match tables

  • matches - identity, kickoff, score, status
  • match_team_stats - per period (overall / 1H / 2H)
  • match_lineups - XI, subs, formations, minutes
  • match_player_stats - ~100 metrics / player / match

Join keys

  • match_id
  • team_id
  • player_id
  • league_id / tournament_id

Start building

The dataset is public. Download it, join it, model it - and let us know what you build.

# Python · Polars
import polars as pl
matches = pl.read_parquet("hf://datasets/gamblistics-lab/sport-statistics/matches.parquet")
stats   = pl.read_parquet("hf://datasets/gamblistics-lab/sport-statistics/match_team_stats.parquet")
df = matches.join(stats, on="match_id", how="left")