Skip to main content
CARLA HQ
MODEL PROFILE & BENCHMARKSRank #1 Global EloOpen Source (Apache 2.0)

Google TabFM v1.0 (Tabular Foundation Model)

Google Research's tabular foundation model architecture focusing on robust columnar tokenization and zero-shot in-context transfer across diverse tabular benchmarks.

Tournament Elo Rating
1111±89 (95% CI)
Global Rank #1 of 4
Overall Improvability
0.85%
Class: 0.16% • Regr: 3.07%
Tournament Matchup Record
32W - 8L(23T)
Across all pairwise combinations
Mean Latency & Runtime
742.64s
Server-side only (Requires Python / JAX / PyTorch environment)
EMPIRICAL BENCHMARK MATRIX

TabFM Dataset-by-Dataset Telemetry

Results on deterministic IID or non-IID evaluation splits, secondary metrics, suite ceilings, and per-dataset Improvability regret across all 21 benchmark datasets.

DatasetDomainTaskRowsCols
Splits?Repeats × folds from the dataset's persisted evaluation protocol. The same row indices are used for every model.
Score?Primary benchmark metric: 1 - AUC error loss for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all three). Evaluated out-of-sample on persisted, model-identical split indices.
Ceiling?Best foundation model benchmark score: lowest 1 - AUC error loss for binary classification, lowest Log-Loss for multiclass, lowest RMSE for regression achieved across the suite.
Improv.?TabArena normalized error regret: (Loss - Ceiling) / (Dummy - Ceiling). 0.0% = Holds suite ceiling.
Latency?Mean inference time per persisted split running out-of-sample inference.
Sheet
AbaloneAgriculture, Forestry & FishingREG4,17783×32.022.020.0%147.9s
AdultEconomics & Public PolicyBIN48,842143×30.06770.06770.0%4421.6s
Airfoil Self NoiseUCIREG1,503510×31.111.02+1.4%43.1s
Amazon Employee AccessInformation Technology & Enterprise SecurityBIN32,76993×30.12730.1232+1.1%3389.9s
Bank MarketingFinance & BankingBIN45,211153×30.18400.18400.0%3095.8s
Blood Transfusion Service CenterHealthcare & BiomedicineBIN748420×30.24450.24450.0%10.4s
Breast WHealthcare & Life SciencesBIN699920×30.00550.0052+0.1%10.9s
CarAutomotive & Fleet ManagementMULTI1,728610×30.01080.01080.0%25.3s
Compas Two YearsLegal & Public SafetyBIN5,278133×30.26190.2612+0.3%107.9s
Credit GFinance & BankingBIN1,0002010×30.19380.19380.0%19.3s
DiabetesHealthcare & BiomedicineBIN768810×30.16230.1607+0.5%11.5s
Employee SalariesHuman Resources & Workforce AnalyticsREG9,228103×310,4357,427+13.9%466.8s
Fitness ClubFitness, Sports & RecreationBIN1,500610×30.17950.1792+0.1%22.1s
House SalesReal Estate & Property ValuationREG21,613203×398,88798,8870.0%1839.0s
HousesReal Estate & Urban PlanningREG20,64083×337,48137,4810.0%1513.5s
Monks Problems 2Data Science & Artificial IntelligenceBIN601620×30.00000.00000.0%9.2s
PhonemeSpeech Processing & Acoustic EngineeringBIN5,40453×30.02020.02020.0%97.7s
SpambaseCybersecurity & IT InfrastructureBIN4,601573×30.00660.00660.0%154.1s
Telco Customer ChurnTelecommunications & Subscription ServicesBIN7,043203×30.14740.14740.0%171.5s
TitanicMaritime Safety & Actuarial Risk AnalysisBIN1,3091110×30.11660.1146+0.5%22.1s
VehicleAutomotive Engineering & Computer VisionMULTI8461810×30.19860.19860.0%16.0s
BROWSER DEPLOYMENT FEASIBILITY

Which Tabular Foundation Model Can Actually Run in the Browser?

Evaluating TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter size, VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.

STATISTICAL METHODOLOGY

How Elo Ratings & Improvability are Computed

Explore the mathematical formulations behind our scale-invariant TabArena Improvability regret metrics, stationary Bradley-Terry Elo tournament ratings, and 1,000-sample bootstrap confidence intervals.

LOCAL WEBGPU MACHINE LEARNING

Evaluate Tabular Foundation Models In Your Spreadsheets

Carla HQ brings TabICLv2 directly into Google Sheets via in-browser WebGPU inference. Run zero-trust predictions on private data without API keys or cloud server uploads.