Skip to main content
CARLA HQ
MODEL PROFILE & BENCHMARKSRank #4 Global EloOpen Source (BSD-3-Clause)WebGPU Client-Side Runtime

TabICLv2 (Inria SODA)

Inria SODA's state-of-the-art tabular foundation model engineered for instantaneous in-context learning. TabICLv2 runs zero-shot inference without iterative gradient descent and is fully optimized for client-side WebGPU execution in Google Sheets via Carla HQ.

Tournament Elo Rating
932±58 (95% CI)
Global Rank #4 of 4
Overall Improvability
2.11%
Class: 1.83% • Regr: 2.99%
Tournament Matchup Record
9W - 30L(24T)
Across all pairwise combinations
Mean Latency & Runtime
4.26s
Native WebGPU Browser Runtime (Carla Engine / ONNX Runtime WASM)
EMPIRICAL BENCHMARK MATRIX

TabICLv2 Dataset-by-Dataset Telemetry

Results on deterministic IID or non-IID evaluation splits, secondary metrics, suite ceilings, and per-dataset Improvability regret across all 21 benchmark datasets.

DatasetDomainTaskRowsCols
Splits?Repeats × folds from the dataset's persisted evaluation protocol. The same row indices are used for every model.
Score?Primary benchmark metric: 1 - AUC error loss for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all three). Evaluated out-of-sample on persisted, model-identical split indices.
Ceiling?Best foundation model benchmark score: lowest 1 - AUC error loss for binary classification, lowest Log-Loss for multiclass, lowest RMSE for regression achieved across the suite.
Improv.?TabArena normalized error regret: (Loss - Ceiling) / (Dummy - Ceiling). 0.0% = Holds suite ceiling.
Latency?Mean inference time per persisted split running out-of-sample inference.
Sheet
AbaloneAgriculture, Forestry & FishingREG4,17783×32.032.02+0.8%1.1s
AdultEconomics & Public PolicyBIN48,842143×30.07930.0677+2.7%23.8s
Airfoil Self NoiseUCIREG1,503510×31.121.02+1.7%652ms
Amazon Employee AccessInformation Technology & Enterprise SecurityBIN32,76993×30.14830.1232+6.7%11.9s
Bank MarketingFinance & BankingBIN45,211153×30.19700.1840+4.1%21.5s
Blood Transfusion Service CenterHealthcare & BiomedicineBIN748420×30.24460.24450.0%610ms
Breast WHealthcare & Life SciencesBIN699920×30.00520.00520.0%672ms
CarAutomotive & Fleet ManagementMULTI1,728610×30.02310.0108+0.9%737ms
Compas Two YearsLegal & Public SafetyBIN5,278133×30.27010.2612+3.7%1.7s
Credit GFinance & BankingBIN1,0002010×30.20160.1938+2.6%888ms
DiabetesHealthcare & BiomedicineBIN768810×30.16320.1607+0.7%651ms
Employee SalariesHuman Resources & Workforce AnalyticsREG9,228103×38,1267,427+3.2%2.1s
Fitness ClubFitness, Sports & RecreationBIN1,500610×30.17950.1792+0.1%703ms
House SalesReal Estate & Property ValuationREG21,613203×3111,52498,887+4.7%7.5s
HousesReal Estate & Urban PlanningREG20,64083×340,97137,481+4.5%5.0s
Monks Problems 2Data Science & Artificial IntelligenceBIN601620×30.00000.00000.0%598ms
PhonemeSpeech Processing & Acoustic EngineeringBIN5,40453×30.02780.0202+1.6%1.4s
SpambaseCybersecurity & IT InfrastructureBIN4,601573×30.00750.0066+0.2%3.7s
Telco Customer ChurnTelecommunications & Subscription ServicesBIN7,043203×30.14880.1474+0.4%2.6s
TitanicMaritime Safety & Actuarial Risk AnalysisBIN1,3091110×30.12290.1146+2.2%776ms
VehicleAutomotive Engineering & Computer VisionMULTI8461810×30.24010.1986+3.5%818ms
BROWSER DEPLOYMENT FEASIBILITY

Which Tabular Foundation Model Can Actually Run in the Browser?

Evaluating TabICLv2, TabPFN-3, Google TabFM, and EXAONE-Tabular on parameter size, VRAM footprint, licensing terms, and WebGPU client-side execution compatibility.

STATISTICAL METHODOLOGY

How Elo Ratings & Improvability are Computed

Explore the mathematical formulations behind our scale-invariant TabArena Improvability regret metrics, stationary Bradley-Terry Elo tournament ratings, and 1,000-sample bootstrap confidence intervals.

LOCAL WEBGPU MACHINE LEARNING

Evaluate Tabular Foundation Models In Your Spreadsheets

Carla HQ brings TabICLv2 directly into Google Sheets via in-browser WebGPU inference. Run zero-trust predictions on private data without API keys or cloud server uploads.