Google TabFM v1.0 (Tabular Foundation Model)
Google Research's tabular foundation model architecture focusing on robust columnar tokenization and zero-shot in-context transfer across diverse tabular benchmarks.
TabFM Dataset-by-Dataset Telemetry
Results on deterministic IID or non-IID evaluation splits, secondary metrics, suite ceilings, and per-dataset Improvability regret across all 21 benchmark datasets.
| Dataset | Domain | Task | Rows | Cols | Splits?Repeats × folds from the dataset's persisted evaluation protocol. The same row indices are used for every model. | Score?Primary benchmark metric: 1 - AUC error loss for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all three). Evaluated out-of-sample on persisted, model-identical split indices. | Ceiling?Best foundation model benchmark score: lowest 1 - AUC error loss for binary classification, lowest Log-Loss for multiclass, lowest RMSE for regression achieved across the suite. | Improv.?TabArena normalized error regret: (Loss - Ceiling) / (Dummy - Ceiling). 0.0% = Holds suite ceiling. | Latency?Mean inference time per persisted split running out-of-sample inference. | Sheet |
|---|---|---|---|---|---|---|---|---|---|---|
| Abalone | Agriculture, Forestry & Fishing | REG | 4,177 | 8 | 3×3 | 2.02 | 2.02 | 0.0% | 147.9s | |
| Adult | Economics & Public Policy | BIN | 48,842 | 14 | 3×3 | 0.0677 | 0.0677 | 0.0% | 4421.6s | |
| Airfoil Self Noise | UCI | REG | 1,503 | 5 | 10×3 | 1.11 | 1.02 | +1.4% | 43.1s | |
| Amazon Employee Access | Information Technology & Enterprise Security | BIN | 32,769 | 9 | 3×3 | 0.1273 | 0.1232 | +1.1% | 3389.9s | |
| Bank Marketing | Finance & Banking | BIN | 45,211 | 15 | 3×3 | 0.1840 | 0.1840 | 0.0% | 3095.8s | |
| Blood Transfusion Service Center | Healthcare & Biomedicine | BIN | 748 | 4 | 20×3 | 0.2445 | 0.2445 | 0.0% | 10.4s | |
| Breast W | Healthcare & Life Sciences | BIN | 699 | 9 | 20×3 | 0.0055 | 0.0052 | +0.1% | 10.9s | |
| Car | Automotive & Fleet Management | MULTI | 1,728 | 6 | 10×3 | 0.0108 | 0.0108 | 0.0% | 25.3s | |
| Compas Two Years | Legal & Public Safety | BIN | 5,278 | 13 | 3×3 | 0.2619 | 0.2612 | +0.3% | 107.9s | |
| Credit G | Finance & Banking | BIN | 1,000 | 20 | 10×3 | 0.1938 | 0.1938 | 0.0% | 19.3s | |
| Diabetes | Healthcare & Biomedicine | BIN | 768 | 8 | 10×3 | 0.1623 | 0.1607 | +0.5% | 11.5s | |
| Employee Salaries | Human Resources & Workforce Analytics | REG | 9,228 | 10 | 3×3 | 10,435 | 7,427 | +13.9% | 466.8s | |
| Fitness Club | Fitness, Sports & Recreation | BIN | 1,500 | 6 | 10×3 | 0.1795 | 0.1792 | +0.1% | 22.1s | |
| House Sales | Real Estate & Property Valuation | REG | 21,613 | 20 | 3×3 | 98,887 | 98,887 | 0.0% | 1839.0s | |
| Houses | Real Estate & Urban Planning | REG | 20,640 | 8 | 3×3 | 37,481 | 37,481 | 0.0% | 1513.5s | |
| Monks Problems 2 | Data Science & Artificial Intelligence | BIN | 601 | 6 | 20×3 | 0.0000 | 0.0000 | 0.0% | 9.2s | |
| Phoneme | Speech Processing & Acoustic Engineering | BIN | 5,404 | 5 | 3×3 | 0.0202 | 0.0202 | 0.0% | 97.7s | |
| Spambase | Cybersecurity & IT Infrastructure | BIN | 4,601 | 57 | 3×3 | 0.0066 | 0.0066 | 0.0% | 154.1s | |
| Telco Customer Churn | Telecommunications & Subscription Services | BIN | 7,043 | 20 | 3×3 | 0.1474 | 0.1474 | 0.0% | 171.5s | |
| Titanic | Maritime Safety & Actuarial Risk Analysis | BIN | 1,309 | 11 | 10×3 | 0.1166 | 0.1146 | +0.5% | 22.1s | |
| Vehicle | Automotive Engineering & Computer Vision | MULTI | 846 | 18 | 10×3 | 0.1986 | 0.1986 | 0.0% | 16.0s |
Which Tabular Foundation Model Can Actually Run in the Browser?
Evaluating TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter size, VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.
How Elo Ratings & Improvability are Computed
Explore the mathematical formulations behind our scale-invariant TabArena Improvability regret metrics, stationary Bradley-Terry Elo tournament ratings, and 1,000-sample bootstrap confidence intervals.
Evaluate Tabular Foundation Models In Your Spreadsheets
Carla HQ brings TabICLv2 directly into Google Sheets via in-browser WebGPU inference. Run zero-trust predictions on private data without API keys or cloud server uploads.