TabICLv2 vs EXAONE Tabular Benchmark
Empirical head-to-head evaluation of TabICLv2 (Inria SODA) versus EXAONE Tabular (LG AI Research) on identical, precomputed IID or non-IID splits across 21 diverse tabular datasets.
Model Specifications & Design Trade-offs
Comparing transformer architecture, Bayesian priors, sequence context windows, and browser execution runtime.
TabICLv2
Inria SODA's state-of-the-art tabular foundation model engineered for instantaneous in-context learning. TabICLv2 runs zero-shot inference without iterative gradient descent and is fully optimized for client-side WebGPU execution in Google Sheets via Carla HQ.
EXAONE Tabular
Large-scale tabular foundation model by LG AI Research designed for high-capacity tabular representation learning, handling mixed categorical and continuous variables across tabular domains.
TabICLv2 vs TabPFN-3 vs TabFM vs EXAONE: Which Model Can Run in the Browser?
Evaluating model weight size (MB vs GB), VRAM requirements, Python server infrastructure overhead, licensing terms, and WebGPU client-side execution compatibility.
21-Dataset Metric Evaluation Matrix
Empirical performance on deterministic, application-appropriate IID or non-IID splits: 1 - AUC for binary classification, Log-Loss for multiclass, and RMSE for regression (all metrics: lower is better).
| Dataset | Domain | Task | Rows | Splits?Repeats × folds. IID datasets use three folds with size-aware repeats, and the same row indices are used for every model. | Metric?Primary evaluation error loss: 1 - AUC for binary classification, Log-Loss for multiclass, RMSE for regression. All metrics: lower is better. | TabICLv2 Loss?TabICLv2 primary out-of-sample error loss (lower is better). | EXAONE Tabular Loss?EXAONE Tabular primary out-of-sample error loss (lower is better). | Margin (Δ)?Error loss difference. Negative (green) indicates TabICLv2 lead; positive (red) indicates EXAONE Tabular lead. | Winner?Model achieving lower out-of-sample error loss on this benchmark split. | TabICLv2 Latency?Mean inference runtime in seconds per evaluation split for TabICLv2. | EXAONE Tabular Latency?Mean inference runtime in seconds per evaluation split for EXAONE Tabular. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Abalone ↗ | Agriculture, Forestry & Fishing | REG | 4,177 | 3×3 | RMSE | 2.03 | 2.03 | +0.00 | EXAONE Tabular | 1.07s | 13.78s |
| Adult ↗ | Economics & Public Policy | BIN | 48,842 | 3×3 | 1 - AUC | 0.0793 | 0.0680 | +0.0113 | EXAONE Tabular | 23.84s | 859.66s |
| Airfoil Self Noise ↗ | UCI | REG | 1,503 | 10×3 | RMSE | 1.12 | 1.16 | -0.04 | TabICLv2 | 0.65s | 2.34s |
| Amazon Employee Access ↗ | Information Technology & Enterprise Security | BIN | 32,769 | 3×3 | 1 - AUC | 0.1483 | 0.1232 | +0.0251 | EXAONE Tabular | 11.93s | 203.54s |
| Bank Marketing ↗ | Finance & Banking | BIN | 45,211 | 3×3 | 1 - AUC | 0.1970 | 0.1876 | +0.0094 | EXAONE Tabular | 21.49s | 749.39s |
| Blood Transfusion Service Center ↗ | Healthcare & Biomedicine | BIN | 748 | 20×3 | 1 - AUC | 0.2446 | 0.2480 | -0.0035 | TabICLv2 | 0.61s | 0.24s |
| Breast W ↗ | Healthcare & Life Sciences | BIN | 699 | 20×3 | 1 - AUC | 0.0052 | 0.0054 | -0.0002 | TabICLv2 | 0.67s | 0.30s |
| Car ↗ | Automotive & Fleet Management | MULTI | 1,728 | 10×3 | Log-Loss | 0.0231 | 0.0346 | -0.0115 | TabICLv2 | 0.74s | 0.56s |
| Compas Two Years ↗ | Legal & Public Safety | BIN | 5,278 | 3×3 | 1 - AUC | 0.2701 | 0.2612 | +0.0089 | EXAONE Tabular | 1.73s | 4.33s |
| Credit G ↗ | Finance & Banking | BIN | 1,000 | 10×3 | 1 - AUC | 0.2016 | 0.1964 | +0.0053 | EXAONE Tabular | 0.89s | 0.69s |
| Diabetes ↗ | Healthcare & Biomedicine | BIN | 768 | 10×3 | 1 - AUC | 0.1632 | 0.1638 | -0.0005 | TabICLv2 | 0.65s | 0.30s |
| Employee Salaries ↗ | Human Resources & Workforce Analytics | REG | 9,228 | 3×3 | RMSE | 8,126 | 7,427 | +699 | EXAONE Tabular | 2.13s | 77.76s |
| Fitness Club ↗ | Fitness, Sports & Recreation | BIN | 1,500 | 10×3 | 1 - AUC | 0.1795 | 0.1797 | -0.0002 | TabICLv2 | 0.70s | 0.49s |
| House Sales ↗ | Real Estate & Property Valuation | REG | 21,613 | 3×3 | RMSE | 111,524 | 99,520 | +12,004 | EXAONE Tabular | 7.52s | 931.00s |
| Houses ↗ | Real Estate & Urban Planning | REG | 20,640 | 3×3 | RMSE | 40,971 | 39,971 | +1,000 | EXAONE Tabular | 5.04s | 459.14s |
| Monks Problems 2 ↗ | Data Science & Artificial Intelligence | BIN | 601 | 20×3 | 1 - AUC | 0.0000 | 0.0000 | ±0.00 | Tie | 0.60s | 0.24s |
| Phoneme ↗ | Speech Processing & Acoustic Engineering | BIN | 5,404 | 3×3 | 1 - AUC | 0.0278 | 0.0276 | +0.0002 | EXAONE Tabular | 1.37s | 2.54s |
| Spambase ↗ | Cybersecurity & IT Infrastructure | BIN | 4,601 | 3×3 | 1 - AUC | 0.0075 | 0.0076 | ±0.00 | Tie | 3.66s | 12.63s |
| Telco Customer Churn ↗ | Telecommunications & Subscription Services | BIN | 7,043 | 3×3 | 1 - AUC | 0.1488 | 0.1484 | +0.0005 | EXAONE Tabular | 2.63s | 11.95s |
| Titanic ↗ | Maritime Safety & Actuarial Risk Analysis | BIN | 1,309 | 10×3 | 1 - AUC | 0.1229 | 0.1146 | +0.0083 | EXAONE Tabular | 0.78s | 0.58s |
| Vehicle ↗ | Automotive Engineering & Computer Vision | MULTI | 846 | 10×3 | Log-Loss | 0.2401 | 0.2393 | +0.0008 | EXAONE Tabular | 0.82s | 0.53s |
Evaluate Foundation Models In Your Spreadsheets
Carla HQ brings TabICLv2 foundation models directly into Google Sheets via WebGPU. Run zero-trust predictions on private data without API keys or cloud pipelines.
Frequently Asked Questions: TabICLv2 vs EXAONE Tabular
Common architectural and benchmarking questions comparing TabICLv2 and EXAONE Tabular.
How does TabICLv2 compare to EXAONE Tabular on tabular benchmarks?
Across the Carla benchmark suite, TabICLv2 and EXAONE Tabular are evaluated head-to-head on identical, precomputed IID or non-IID splits using standardized error losses (1 - AUC for binary classification, Log-Loss for multiclass, and RMSE for regression). IID datasets use adaptive repeated 3-fold evaluation to reduce variance. TabICLv2 enables 100% private in-browser WebGPU execution, whereas server-side models like EXAONE Tabular require heavy PyTorch infrastructure and cloud GPU backend servers.
Can TabICLv2 or EXAONE Tabular run in the browser without server infrastructure?
Only TabICLv2 (via Carla HQ and nanotabicl ONNX Runtime WASM/WebGPU) is engineered to execute entirely client-side inside web browsers like Google Chrome. EXAONE Tabular requires full Python and PyTorch server infrastructure with high memory allocations.
What are the primary architectural differences between TabICLv2 and EXAONE Tabular?
TabICLv2 uses in-context tabular transformer with prefix attention & kv caching, whereas EXAONE Tabular leverages multi-task columnar autoregressive tabular transformer. TabICLv2 specifically utilizes prefix attention with KV cache reuse for real-time tabular evaluation in spreadsheets.