TabPFN v3 (Prior Labs)
Pioneering Prior-Data Fitted Network that approximates full Bayesian posterior inference over synthetic structural priors in a single forward pass without iterative training.
TabPFN Dataset-by-Dataset Telemetry
Results on deterministic IID or non-IID evaluation splits, secondary metrics, suite ceilings, and per-dataset Improvability regret across all 21 benchmark datasets.
| Dataset | Domain | Task | Rows | Cols | Splits?Repeats × folds from the dataset's persisted evaluation protocol. The same row indices are used for every model. | Score?Primary benchmark metric: 1 - AUC error loss for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all three). Evaluated out-of-sample on persisted, model-identical split indices. | Ceiling?Best foundation model benchmark score: lowest 1 - AUC error loss for binary classification, lowest Log-Loss for multiclass, lowest RMSE for regression achieved across the suite. | Improv.?TabArena normalized error regret: (Loss - Ceiling) / (Dummy - Ceiling). 0.0% = Holds suite ceiling. | Latency?Mean inference time per persisted split running out-of-sample inference. | Sheet |
|---|---|---|---|---|---|---|---|---|---|---|
| Abalone | Agriculture, Forestry & Fishing | REG | 4,177 | 8 | 3×3 | 2.03 | 2.02 | +0.2% | 4.5s | |
| Adult | Economics & Public Policy | BIN | 48,842 | 14 | 3×3 | 0.0799 | 0.0677 | +2.8% | 64.1s | |
| Airfoil Self Noise | UCI | REG | 1,503 | 5 | 10×3 | 1.02 | 1.02 | 0.0% | 2.4s | |
| Amazon Employee Access | Information Technology & Enterprise Security | BIN | 32,769 | 9 | 3×3 | 0.1542 | 0.1232 | +8.2% | 31.6s | |
| Bank Marketing | Finance & Banking | BIN | 45,211 | 15 | 3×3 | 0.1883 | 0.1840 | +1.4% | 56.4s | |
| Blood Transfusion Service Center | Healthcare & Biomedicine | BIN | 748 | 4 | 20×3 | 0.2462 | 0.2445 | +0.6% | 1.0s | |
| Breast W | Healthcare & Life Sciences | BIN | 699 | 9 | 20×3 | 0.0056 | 0.0052 | +0.1% | 1.0s | |
| Car | Automotive & Fleet Management | MULTI | 1,728 | 6 | 10×3 | 0.0238 | 0.0108 | +0.9% | 1.4s | |
| Compas Two Years | Legal & Public Safety | BIN | 5,278 | 13 | 3×3 | 0.2615 | 0.2612 | +0.1% | 2.9s | |
| Credit G | Finance & Banking | BIN | 1,000 | 20 | 10×3 | 0.2031 | 0.1938 | +3.0% | 1.2s | |
| Diabetes | Healthcare & Biomedicine | BIN | 768 | 8 | 10×3 | 0.1607 | 0.1607 | 0.0% | 1.1s | |
| Employee Salaries | Human Resources & Workforce Analytics | REG | 9,228 | 10 | 3×3 | 8,308 | 7,427 | +4.1% | 10.7s | |
| Fitness Club | Fitness, Sports & Recreation | BIN | 1,500 | 6 | 10×3 | 0.1792 | 0.1792 | 0.0% | 1.3s | |
| House Sales | Real Estate & Property Valuation | REG | 21,613 | 20 | 3×3 | 102,039 | 98,887 | +1.2% | 38.3s | |
| Houses | Real Estate & Urban Planning | REG | 20,640 | 8 | 3×3 | 39,523 | 37,481 | +2.6% | 30.9s | |
| Monks Problems 2 | Data Science & Artificial Intelligence | BIN | 601 | 6 | 20×3 | 0.0000 | 0.0000 | 0.0% | 1.1s | |
| Phoneme | Speech Processing & Acoustic Engineering | BIN | 5,404 | 5 | 3×3 | 0.0274 | 0.0202 | +1.5% | 2.6s | |
| Spambase | Cybersecurity & IT Infrastructure | BIN | 4,601 | 57 | 3×3 | 0.0085 | 0.0066 | +0.4% | 3.8s | |
| Telco Customer Churn | Telecommunications & Subscription Services | BIN | 7,043 | 20 | 3×3 | 0.1486 | 0.1474 | +0.3% | 4.2s | |
| Titanic | Maritime Safety & Actuarial Risk Analysis | BIN | 1,309 | 11 | 10×3 | 0.1360 | 0.1146 | +5.5% | 1.2s | |
| Vehicle | Automotive Engineering & Computer Vision | MULTI | 846 | 18 | 10×3 | 0.2263 | 0.1986 | +2.3% | 1.1s |
Which Tabular Foundation Model Can Actually Run in the Browser?
Evaluating TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter size, VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.
How Elo Ratings & Improvability are Computed
Explore the mathematical formulations behind our scale-invariant TabArena Improvability regret metrics, stationary Bradley-Terry Elo tournament ratings, and 1,000-sample bootstrap confidence intervals.
Evaluate Tabular Foundation Models In Your Spreadsheets
Carla HQ brings TabICLv2 directly into Google Sheets via in-browser WebGPU inference. Run zero-trust predictions on private data without API keys or cloud server uploads.