Tabular Foundation Model Leaderboard & Matrix
Empirical evaluation of TabICLv2, TabPFN v3, EXAONE Tabular, and Google TabFM across 21 diverse tabular datasets and 489 persisted evaluation splits. Standardized via Bradley-Terry Elo ratings (with 95% bootstrap confidence intervals) and scale-invariant Improvability.
Tournament Standings & Error Bounds
Evaluated across 21 open benchmarks • adaptive repeated 3-fold IID splits • 1 - AUC, Log-Loss, and RMSE error losses (lower is better).
| Rank | Foundation Model | Elo Rating (95% CI)?Bradley-Terry stationary Elo rating computed across all evaluation-split matchups with 1,000 bootstrap resamples. Centered at baseline 1000. Higher is better. | Improvability ↓?Scale-invariant normalized regret relative to empirical suite ceilings across 1 - AUC, Log-Loss, and RMSE. Lower is better (0.0% is optimal). | Mean Latency?Average execution time in seconds per persisted evaluation split across all 21 benchmark datasets. Lower is faster. | In-Browser Runtime |
|---|---|---|---|---|---|
| #1 | Google Research • Columnar | 1111 (±87) | 0.85% | 742.64s | Server only |
| #2 | LG AI Research • Multi-Task | 1021 (±77) | 0.85% | 158.67s | Server only |
| #3 | Prior Labs • Prior-Data | 936 (±63) | 1.69% | 12.52s | PyTorch only |
| #4 | TabICLv2 ↗CARLA ENGINE Inria SODA • In-Context | 932 (±61) | 2.11% | 4.26s | WebGPU (WASM) |
Cross-Dataset Matchup Grid
Direct pairwise dataset win-loss tallies across all 4 foundation models. Click any cell to inspect the complete 21-dataset breakdown.
Select a Pairwise Benchmark
Detailed breakdown of dataset wins, predictive accuracies, AUC-ROC rankings, and in-browser runtime capabilities.
TabICLv2 vs TabPFN
Inria SODA vs Prior Labs
TabICLv2 vs EXAONE Tabular
Inria SODA vs LG AI Research
TabICLv2 vs TabFM
Inria SODA vs Google Research
TabPFN vs EXAONE Tabular
Prior Labs vs LG AI Research
TabPFN vs TabFM
Prior Labs vs Google Research
EXAONE Tabular vs TabFM
LG AI Research vs Google Research
Which Tabular Foundation Model Can Actually Run in the Browser?
Detailed technical evaluation comparing TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter sizes, RAM/VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.
Rigorous Statistical Benchmark Methodology
Learn how stationary Bradley-Terry Elo ratings, 1,000 bootstrap confidence intervals (±CI), and TabArena scale-invariant Improvability are calculated across all 21 benchmark datasets.