Skip to main content
CARLA HQ
PAIRWISE BENCHMARKAdaptive Repeated 3-Fold21 Datasets

TabPFN vs TabFM Benchmark

Empirical head-to-head evaluation of TabPFN v3 (Prior Labs) versus Google TabFM (Google Research) on identical, precomputed IID or non-IID splits across 21 diverse tabular datasets.

HEAD-TO-HEAD WINS?Direct pairwise dataset matchup wins where the model achieved lower out-of-sample error loss (1 - AUC for binary, Log-Loss for multiclass, RMSE for regression).
TabPFN5
:
14TabFM
(2 ties) across 21 benchmarks
TOURNAMENT ELO?Global Bradley-Terry Elo rating computed across all evaluation-split matchups with 1,000 bootstrap resamples. Higher is better (baseline 1000).
TabPFN936 (±61)
vs
TabFM1111 (±86)
Higher is better • 95% Bootstrap CI
MACRO IMPROVABILITY?Normalized average error regret relative to empirical suite ceilings across all 21 datasets. Lower is better (0.0% is optimal ceiling).
TabPFN1.69%
vs
TabFM0.85%
Lower is better • Error regret vs ceiling
MEAN SPLIT LATENCY?Average execution latency in seconds per persisted IID or non-IID evaluation split across all benchmarks. Lower is faster.
TabPFN12.52s
vs
TabFM742.64s
Lower is faster • adaptive split protocol
ARCHITECTURAL BLUEPRINTS

Model Specifications & Design Trade-offs

Comparing transformer architecture, Bayesian priors, sequence context windows, and browser execution runtime.

Prior Labs

TabPFN v3

GitHub / Paper ↗

Pioneering Prior-Data Fitted Network that approximates full Bayesian posterior inference over synthetic structural priors in a single forward pass without iterative training.

ArchitecturePrior-Data Fitted Network (PFN) with Causally Invariant Embeddings
Prior / PretrainingBayesian structural equation models and synthetic Gaussian priors
Context WindowFixed context window (typically 1k–10k training samples)
Runtime PlatformServer-side only (Requires Python runtime / CUDA / PyTorch)
LicenseResearch & Non-Commercial
Google Research

Google TabFM

GitHub / Paper ↗

Google Research's tabular foundation model architecture focusing on robust columnar tokenization and zero-shot in-context transfer across diverse tabular benchmarks.

ArchitectureColumnar Embedding & Self-Supervised In-Context Transformer
Prior / PretrainingGenerative synthetic distributions and masked column modeling
Context WindowColumnar token budget per batch
Runtime PlatformServer-side only (Requires Python / JAX / PyTorch environment)
LicenseOpen Source (Apache 2.0)
DEEP DIVE REPORTON-DEVICE RUNTIME ANALYSIS

TabICLv2 vs TabPFN-3 vs TabFM vs EXAONE: Which Model Can Run in the Browser?

Evaluating model weight size (MB vs GB), VRAM requirements, Python server infrastructure overhead, licensing terms, and WebGPU client-side execution compatibility.

HEAD-TO-HEAD BREAKDOWN

21-Dataset Metric Evaluation Matrix

Empirical performance on deterministic, application-appropriate IID or non-IID splits: 1 - AUC for binary classification, Log-Loss for multiclass, and RMSE for regression (all metrics: lower is better).

DatasetDomainTaskRows
Splits?Repeats × folds. IID datasets use three folds with size-aware repeats, and the same row indices are used for every model.
Metric?Primary evaluation error loss: 1 - AUC for binary classification, Log-Loss for multiclass, RMSE for regression. All metrics: lower is better.
TabPFN
Loss?TabPFN v3 primary out-of-sample error loss (lower is better).
TabFM
Loss?Google TabFM primary out-of-sample error loss (lower is better).
Margin (Δ)?Error loss difference. Negative (green) indicates TabPFN lead; positive (red) indicates TabFM lead.
Winner?Model achieving lower out-of-sample error loss on this benchmark split.
TabPFN
Latency?Mean inference runtime in seconds per evaluation split for TabPFN.
TabFM
Latency?Mean inference runtime in seconds per evaluation split for TabFM.
Abalone ↗Agriculture, Forestry & FishingREG4,1773×3RMSE2.032.02+0.00TabFM4.50s147.89s
Adult ↗Economics & Public PolicyBIN48,8423×31 - AUC0.07990.0677+0.0122TabFM64.05s4421.56s
Airfoil Self Noise ↗UCIREG1,50310×3RMSE1.021.11-0.08TabPFN2.44s43.06s
Amazon Employee Access ↗Information Technology & Enterprise SecurityBIN32,7693×31 - AUC0.15420.1273+0.0269TabFM31.56s3389.93s
Bank Marketing ↗Finance & BankingBIN45,2113×31 - AUC0.18830.1840+0.0043TabFM56.36s3095.77s
Blood Transfusion Service Center ↗Healthcare & BiomedicineBIN74820×31 - AUC0.24620.2445+0.0017TabFM1.05s10.38s
Breast W ↗Healthcare & Life SciencesBIN69920×31 - AUC0.00560.0055±0.00Tie1.04s10.88s
Car ↗Automotive & Fleet ManagementMULTI1,72810×3Log-Loss0.02380.0108+0.0130TabFM1.38s25.35s
Compas Two Years ↗Legal & Public SafetyBIN5,2783×31 - AUC0.26150.2619-0.0004TabPFN2.89s107.91s
Credit G ↗Finance & BankingBIN1,00010×31 - AUC0.20310.1938+0.0093TabFM1.22s19.29s
Diabetes ↗Healthcare & BiomedicineBIN76810×31 - AUC0.16070.1623-0.0016TabPFN1.06s11.50s
Employee Salaries ↗Human Resources & Workforce AnalyticsREG9,2283×3RMSE8,30810,435-2,127TabPFN10.70s466.77s
Fitness Club ↗Fitness, Sports & RecreationBIN1,50010×31 - AUC0.17920.1795-0.0004TabPFN1.30s22.13s
House Sales ↗Real Estate & Property ValuationREG21,6133×3RMSE102,03998,887+3,153TabFM38.26s1839.00s
Houses ↗Real Estate & Urban PlanningREG20,6403×3RMSE39,52337,481+2,042TabFM30.93s1513.45s
Monks Problems 2 ↗Data Science & Artificial IntelligenceBIN60120×31 - AUC0.00000.0000±0.00Tie1.15s9.21s
Phoneme ↗Speech Processing & Acoustic EngineeringBIN5,4043×31 - AUC0.02740.0202+0.0071TabFM2.64s97.68s
Spambase ↗Cybersecurity & IT InfrastructureBIN4,6013×31 - AUC0.00850.0066+0.0019TabFM3.81s154.06s
Telco Customer Churn ↗Telecommunications & Subscription ServicesBIN7,0433×31 - AUC0.14860.1474+0.0012TabFM4.22s171.54s
Titanic ↗Maritime Safety & Actuarial Risk AnalysisBIN1,30910×31 - AUC0.13600.1166+0.0194TabFM1.25s22.09s
Vehicle ↗Automotive Engineering & Computer VisionMULTI84610×3Log-Loss0.22630.1986+0.0277TabFM1.11s15.97s
ZERO-SERVER CLIENT-SIDE EXECUTION

Evaluate Foundation Models In Your Spreadsheets

Carla HQ brings TabICLv2 foundation models directly into Google Sheets via WebGPU. Run zero-trust predictions on private data without API keys or cloud pipelines.

FREQUENTLY ASKED QUESTIONS

Frequently Asked Questions: TabPFN vs TabFM

Common architectural and benchmarking questions comparing TabPFN v3 and Google TabFM.

How does TabPFN compare to TabFM on tabular benchmarks?

Across the Carla benchmark suite, TabPFN v3 and Google TabFM are evaluated head-to-head on identical, precomputed IID or non-IID splits using standardized error losses (1 - AUC for binary classification, Log-Loss for multiclass, and RMSE for regression). IID datasets use adaptive repeated 3-fold evaluation to reduce variance. TabICLv2 enables 100% private in-browser WebGPU execution, whereas server-side models like TabFM require heavy PyTorch infrastructure and cloud GPU backend servers.

Can TabPFN or TabFM run in the browser without server infrastructure?

Only TabICLv2 (via Carla HQ and nanotabicl ONNX Runtime WASM/WebGPU) is engineered to execute entirely client-side inside web browsers like Google Chrome. TabPFN v3 requires full Python and PyTorch server infrastructure with high memory allocations.

What are the primary architectural differences between TabPFN and TabFM?

TabPFN v3 uses prior-data fitted network (pfn) with causally invariant embeddings, whereas Google TabFM leverages columnar embedding & self-supervised in-context transformer. TabICLv2 specifically utilizes prefix attention with KV cache reuse for real-time tabular evaluation in spreadsheets.