Skip to main content
CARLA HQ
TOURNAMENT BENCHMARKAdaptive Repeated 3-Fold1,000 Bootstrap Resamples21 Open Datasets

Tabular Foundation Model Leaderboard & Matrix

Empirical evaluation of TabICLv2, TabPFN v3, EXAONE Tabular, and Google TabFM across 21 diverse tabular datasets and 489 persisted evaluation splits. Standardized via Bradley-Terry Elo ratings (with 95% bootstrap confidence intervals) and scale-invariant Improvability.

GLOBAL TABULAR FOUNDATION MODEL LEADERBOARD

Tournament Standings & Error Bounds

Evaluated across 21 open benchmarks • adaptive repeated 3-fold IID splits • 1 - AUC, Log-Loss, and RMSE error losses (lower is better).

RankFoundation Model
Elo Rating (95% CI)?Bradley-Terry stationary Elo rating computed across all evaluation-split matchups with 1,000 bootstrap resamples. Centered at baseline 1000. Higher is better.
Improvability ↓?Scale-invariant normalized regret relative to empirical suite ceilings across 1 - AUC, Log-Loss, and RMSE. Lower is better (0.0% is optimal).
Mean Latency?Average execution time in seconds per persisted evaluation split across all 21 benchmark datasets. Lower is faster.
In-Browser Runtime
#1
Google Research • Columnar
1111 (±87)0.85%742.64sServer only
#2
LG AI Research • Multi-Task
1021 (±77)0.85%158.67sServer only
#3
Prior Labs • Prior-Data
936 (±63)1.69%12.52sPyTorch only
#4
TabICLv2 ↗CARLA ENGINE
Inria SODA • In-Context
932 (±61)2.11%4.26sWebGPU (WASM)
PAIRWISE HEAD-TO-HEAD MATRIX

Cross-Dataset Matchup Grid

Direct pairwise dataset win-loss tallies across all 4 foundation models. Click any cell to inspect the complete 21-dataset breakdown.

Model MatchupTabFMEXAONE TabularTabPFNTabICLv2
TabFMGoogle Research[15 : 4][14 : 5][16 : 2]
EXAONE TabularLG AI Research[4 : 15][11 : 9][13 : 6]
TabPFNPrior Labs[5 : 14][9 : 11][11 : 9]
TabICLv2Inria SODA[2 : 16][6 : 13][9 : 11]
BROWSER DEPLOYMENT DEEP DIVE

Which Tabular Foundation Model Can Actually Run in the Browser?

Detailed technical evaluation comparing TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter sizes, RAM/VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.

EVALUATION FRAMEWORK

Rigorous Statistical Benchmark Methodology

Learn how stationary Bradley-Terry Elo ratings, 1,000 bootstrap confidence intervals (±CI), and TabArena scale-invariant Improvability are calculated across all 21 benchmark datasets.