TabPFN vs TabFM Benchmark
Empirical head-to-head evaluation of TabPFN v3 (Prior Labs) versus Google TabFM (Google Research) on identical, precomputed IID or non-IID splits across 21 diverse tabular datasets.
Model Specifications & Design Trade-offs
Comparing transformer architecture, Bayesian priors, sequence context windows, and browser execution runtime.
TabPFN v3
Pioneering Prior-Data Fitted Network that approximates full Bayesian posterior inference over synthetic structural priors in a single forward pass without iterative training.
Google TabFM
Google Research's tabular foundation model architecture focusing on robust columnar tokenization and zero-shot in-context transfer across diverse tabular benchmarks.
TabICLv2 vs TabPFN-3 vs TabFM vs EXAONE: Which Model Can Run in the Browser?
Evaluating model weight size (MB vs GB), VRAM requirements, Python server infrastructure overhead, licensing terms, and WebGPU client-side execution compatibility.
21-Dataset Metric Evaluation Matrix
Empirical performance on deterministic, application-appropriate IID or non-IID splits: 1 - AUC for binary classification, Log-Loss for multiclass, and RMSE for regression (all metrics: lower is better).
| Dataset | Domain | Task | Rows | Splits?Repeats × folds. IID datasets use three folds with size-aware repeats, and the same row indices are used for every model. | Metric?Primary evaluation error loss: 1 - AUC for binary classification, Log-Loss for multiclass, RMSE for regression. All metrics: lower is better. | TabPFN Loss?TabPFN v3 primary out-of-sample error loss (lower is better). | TabFM Loss?Google TabFM primary out-of-sample error loss (lower is better). | Margin (Δ)?Error loss difference. Negative (green) indicates TabPFN lead; positive (red) indicates TabFM lead. | Winner?Model achieving lower out-of-sample error loss on this benchmark split. | TabPFN Latency?Mean inference runtime in seconds per evaluation split for TabPFN. | TabFM Latency?Mean inference runtime in seconds per evaluation split for TabFM. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Abalone ↗ | Agriculture, Forestry & Fishing | REG | 4,177 | 3×3 | RMSE | 2.03 | 2.02 | +0.00 | TabFM | 4.50s | 147.89s |
| Adult ↗ | Economics & Public Policy | BIN | 48,842 | 3×3 | 1 - AUC | 0.0799 | 0.0677 | +0.0122 | TabFM | 64.05s | 4421.56s |
| Airfoil Self Noise ↗ | UCI | REG | 1,503 | 10×3 | RMSE | 1.02 | 1.11 | -0.08 | TabPFN | 2.44s | 43.06s |
| Amazon Employee Access ↗ | Information Technology & Enterprise Security | BIN | 32,769 | 3×3 | 1 - AUC | 0.1542 | 0.1273 | +0.0269 | TabFM | 31.56s | 3389.93s |
| Bank Marketing ↗ | Finance & Banking | BIN | 45,211 | 3×3 | 1 - AUC | 0.1883 | 0.1840 | +0.0043 | TabFM | 56.36s | 3095.77s |
| Blood Transfusion Service Center ↗ | Healthcare & Biomedicine | BIN | 748 | 20×3 | 1 - AUC | 0.2462 | 0.2445 | +0.0017 | TabFM | 1.05s | 10.38s |
| Breast W ↗ | Healthcare & Life Sciences | BIN | 699 | 20×3 | 1 - AUC | 0.0056 | 0.0055 | ±0.00 | Tie | 1.04s | 10.88s |
| Car ↗ | Automotive & Fleet Management | MULTI | 1,728 | 10×3 | Log-Loss | 0.0238 | 0.0108 | +0.0130 | TabFM | 1.38s | 25.35s |
| Compas Two Years ↗ | Legal & Public Safety | BIN | 5,278 | 3×3 | 1 - AUC | 0.2615 | 0.2619 | -0.0004 | TabPFN | 2.89s | 107.91s |
| Credit G ↗ | Finance & Banking | BIN | 1,000 | 10×3 | 1 - AUC | 0.2031 | 0.1938 | +0.0093 | TabFM | 1.22s | 19.29s |
| Diabetes ↗ | Healthcare & Biomedicine | BIN | 768 | 10×3 | 1 - AUC | 0.1607 | 0.1623 | -0.0016 | TabPFN | 1.06s | 11.50s |
| Employee Salaries ↗ | Human Resources & Workforce Analytics | REG | 9,228 | 3×3 | RMSE | 8,308 | 10,435 | -2,127 | TabPFN | 10.70s | 466.77s |
| Fitness Club ↗ | Fitness, Sports & Recreation | BIN | 1,500 | 10×3 | 1 - AUC | 0.1792 | 0.1795 | -0.0004 | TabPFN | 1.30s | 22.13s |
| House Sales ↗ | Real Estate & Property Valuation | REG | 21,613 | 3×3 | RMSE | 102,039 | 98,887 | +3,153 | TabFM | 38.26s | 1839.00s |
| Houses ↗ | Real Estate & Urban Planning | REG | 20,640 | 3×3 | RMSE | 39,523 | 37,481 | +2,042 | TabFM | 30.93s | 1513.45s |
| Monks Problems 2 ↗ | Data Science & Artificial Intelligence | BIN | 601 | 20×3 | 1 - AUC | 0.0000 | 0.0000 | ±0.00 | Tie | 1.15s | 9.21s |
| Phoneme ↗ | Speech Processing & Acoustic Engineering | BIN | 5,404 | 3×3 | 1 - AUC | 0.0274 | 0.0202 | +0.0071 | TabFM | 2.64s | 97.68s |
| Spambase ↗ | Cybersecurity & IT Infrastructure | BIN | 4,601 | 3×3 | 1 - AUC | 0.0085 | 0.0066 | +0.0019 | TabFM | 3.81s | 154.06s |
| Telco Customer Churn ↗ | Telecommunications & Subscription Services | BIN | 7,043 | 3×3 | 1 - AUC | 0.1486 | 0.1474 | +0.0012 | TabFM | 4.22s | 171.54s |
| Titanic ↗ | Maritime Safety & Actuarial Risk Analysis | BIN | 1,309 | 10×3 | 1 - AUC | 0.1360 | 0.1166 | +0.0194 | TabFM | 1.25s | 22.09s |
| Vehicle ↗ | Automotive Engineering & Computer Vision | MULTI | 846 | 10×3 | Log-Loss | 0.2263 | 0.1986 | +0.0277 | TabFM | 1.11s | 15.97s |
Evaluate Foundation Models In Your Spreadsheets
Carla HQ brings TabICLv2 foundation models directly into Google Sheets via WebGPU. Run zero-trust predictions on private data without API keys or cloud pipelines.
Frequently Asked Questions: TabPFN vs TabFM
Common architectural and benchmarking questions comparing TabPFN v3 and Google TabFM.
How does TabPFN compare to TabFM on tabular benchmarks?
Across the Carla benchmark suite, TabPFN v3 and Google TabFM are evaluated head-to-head on identical, precomputed IID or non-IID splits using standardized error losses (1 - AUC for binary classification, Log-Loss for multiclass, and RMSE for regression). IID datasets use adaptive repeated 3-fold evaluation to reduce variance. TabICLv2 enables 100% private in-browser WebGPU execution, whereas server-side models like TabFM require heavy PyTorch infrastructure and cloud GPU backend servers.
Can TabPFN or TabFM run in the browser without server infrastructure?
Only TabICLv2 (via Carla HQ and nanotabicl ONNX Runtime WASM/WebGPU) is engineered to execute entirely client-side inside web browsers like Google Chrome. TabPFN v3 requires full Python and PyTorch server infrastructure with high memory allocations.
What are the primary architectural differences between TabPFN and TabFM?
TabPFN v3 uses prior-data fitted network (pfn) with causally invariant embeddings, whereas Google TabFM leverages columnar embedding & self-supervised in-context transformer. TabICLv2 specifically utilizes prefix attention with KV cache reuse for real-time tabular evaluation in spreadsheets.