Skip to main content
CARLA HQ
MODEL PROFILE & BENCHMARKSRank #3 Global EloResearch & Non-Commercial

TabPFN v3 (Prior Labs)

Pioneering Prior-Data Fitted Network that approximates full Bayesian posterior inference over synthetic structural priors in a single forward pass without iterative training.

Tournament Elo Rating
936±63 (95% CI)
Global Rank #3 of 4
Overall Improvability
1.69%
Class: 1.71% • Regr: 1.61%
Tournament Matchup Record
19W - 25L(19T)
Across all pairwise combinations
Mean Latency & Runtime
12.52s
Server-side only (Requires Python runtime / CUDA / PyTorch)
EMPIRICAL BENCHMARK MATRIX

TabPFN Dataset-by-Dataset Telemetry

Results on deterministic IID or non-IID evaluation splits, secondary metrics, suite ceilings, and per-dataset Improvability regret across all 21 benchmark datasets.

DatasetDomainTaskRowsCols
Splits?Repeats × folds from the dataset's persisted evaluation protocol. The same row indices are used for every model.
Score?Primary benchmark metric: 1 - AUC error loss for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all three). Evaluated out-of-sample on persisted, model-identical split indices.
Ceiling?Best foundation model benchmark score: lowest 1 - AUC error loss for binary classification, lowest Log-Loss for multiclass, lowest RMSE for regression achieved across the suite.
Improv.?TabArena normalized error regret: (Loss - Ceiling) / (Dummy - Ceiling). 0.0% = Holds suite ceiling.
Latency?Mean inference time per persisted split running out-of-sample inference.
Sheet
AbaloneAgriculture, Forestry & FishingREG4,17783×32.032.02+0.2%4.5s
AdultEconomics & Public PolicyBIN48,842143×30.07990.0677+2.8%64.1s
Airfoil Self NoiseUCIREG1,503510×31.021.020.0%2.4s
Amazon Employee AccessInformation Technology & Enterprise SecurityBIN32,76993×30.15420.1232+8.2%31.6s
Bank MarketingFinance & BankingBIN45,211153×30.18830.1840+1.4%56.4s
Blood Transfusion Service CenterHealthcare & BiomedicineBIN748420×30.24620.2445+0.6%1.0s
Breast WHealthcare & Life SciencesBIN699920×30.00560.0052+0.1%1.0s
CarAutomotive & Fleet ManagementMULTI1,728610×30.02380.0108+0.9%1.4s
Compas Two YearsLegal & Public SafetyBIN5,278133×30.26150.2612+0.1%2.9s
Credit GFinance & BankingBIN1,0002010×30.20310.1938+3.0%1.2s
DiabetesHealthcare & BiomedicineBIN768810×30.16070.16070.0%1.1s
Employee SalariesHuman Resources & Workforce AnalyticsREG9,228103×38,3087,427+4.1%10.7s
Fitness ClubFitness, Sports & RecreationBIN1,500610×30.17920.17920.0%1.3s
House SalesReal Estate & Property ValuationREG21,613203×3102,03998,887+1.2%38.3s
HousesReal Estate & Urban PlanningREG20,64083×339,52337,481+2.6%30.9s
Monks Problems 2Data Science & Artificial IntelligenceBIN601620×30.00000.00000.0%1.1s
PhonemeSpeech Processing & Acoustic EngineeringBIN5,40453×30.02740.0202+1.5%2.6s
SpambaseCybersecurity & IT InfrastructureBIN4,601573×30.00850.0066+0.4%3.8s
Telco Customer ChurnTelecommunications & Subscription ServicesBIN7,043203×30.14860.1474+0.3%4.2s
TitanicMaritime Safety & Actuarial Risk AnalysisBIN1,3091110×30.13600.1146+5.5%1.2s
VehicleAutomotive Engineering & Computer VisionMULTI8461810×30.22630.1986+2.3%1.1s
BROWSER DEPLOYMENT FEASIBILITY

Which Tabular Foundation Model Can Actually Run in the Browser?

Evaluating TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter size, VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.

STATISTICAL METHODOLOGY

How Elo Ratings & Improvability are Computed

Explore the mathematical formulations behind our scale-invariant TabArena Improvability regret metrics, stationary Bradley-Terry Elo tournament ratings, and 1,000-sample bootstrap confidence intervals.

LOCAL WEBGPU MACHINE LEARNING

Evaluate Tabular Foundation Models In Your Spreadsheets

Carla HQ brings TabICLv2 directly into Google Sheets via in-browser WebGPU inference. Run zero-trust predictions on private data without API keys or cloud server uploads.