Skip to main content
CARLA HQ
TABULAR BENCHMARK SUITE

21 Real-World Datasets & In-Browser Benchmarks

Comprehensive performance evaluation of TabICLv2 executing 100% client-side inside Google Sheets via WebGPU tensor kernels. Zero cloud roundtrips, zero preprocessing, and native integration with LLMs via Model Context Protocol (MCP).

21Evaluation Datasets
14 / 2 / 5Bin / Multi / Reg
100%In-Browser WebGPU
$0.00Cloud Compute Incurred
0 bytesData Exfiltrated
BENCHMARK MATRIX

Benchmark Results & Dataset Directory

Precomputed, deterministic IID and non-IID splits evaluate in-browser TabICLv2 latency, throughput, and out-of-sample accuracy against reference baselines.

DatasetDomainTaskRowsCols
Splits?Adaptive repeated 3-fold IID evaluation. Smaller datasets receive more repeats to reduce variance; every split is precomputed for identical Python and browser evaluation.Explore Split Protocol →
Score?Primary benchmark error loss: 1 - AUC for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all tasks). Evaluated on the dataset's persisted split indices.Explore Methodology Guide →
OpenML?Top-ranked baseline score from traditional tree ensembles (XGBoost, Random Forest) on OpenML.Explore OpenML Reference →
Latency?Mean inference time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Learn about Latency →
Sheet
AbaloneAgriculture, Forestry & FishingREG4,17783×32.0311.7s
AdultEconomics & Public PolicyBIN48,842143×30.07930.1067273.2s
Airfoil Self NoiseUCIREG1,503510×31.124.1s
Amazon Employee AccessInformation Technology & Enterprise SecurityBIN32,76993×30.14830.3362129.2s
Bank MarketingFinance & BankingBIN45,211153×30.19700.0642235.4s
Blood Transfusion Service CenterHealthcare & BiomedicineBIN748420×30.24460.24351.1s
Breast WHealthcare & Life SciencesBIN699920×30.00520.02821.1s
CarAutomotive & Fleet ManagementMULTI1,728610×30.02310.00001.5s
Compas Two YearsLegal & Public SafetyBIN5,278133×30.27015.8s
Credit GFinance & BankingBIN1,0002010×30.20160.22191.4s
DiabetesHealthcare & BiomedicineBIN768810×30.16320.18101.2s
Employee SalariesHuman Resources & Workforce AnalyticsREG9,228103×38,12636.3s
Fitness ClubFitness, Sports & RecreationBIN1,500610×30.17951.3s
House SalesReal Estate & Property ValuationREG21,613203×3111,52462.7s
HousesReal Estate & Urban PlanningREG20,64083×340,97150.3s
Monks Problems 2Data Science & Artificial IntelligenceBIN601620×30.00000.00001.0s
PhonemeSpeech Processing & Acoustic EngineeringBIN5,40453×30.02780.03265.1s
SpambaseCybersecurity & IT InfrastructureBIN4,601573×30.00750.011810.9s
Telco Customer ChurnTelecommunications & Subscription ServicesBIN7,043203×30.148810.2s
TitanicMaritime Safety & Actuarial Risk AnalysisBIN1,3091110×30.12291.3s
VehicleAutomotive Engineering & Computer VisionMULTI8461810×30.24010.03971.3s
TESTING PROTOCOL

Benchmark Methodology & Runtime Specs

How TabICLv2 in-browser foundation models are evaluated under zero-trust, client-side constraints.

Local WebGPU & WASM Runtime

Inference runs directly on device hardware using WebGPU compute shaders and WebAssembly SIMD kernels. Tabular attention passes execute locally without sending row data to an external server.

IID & Non-IID Evaluation

IID datasets use size-aware repeated 3-fold evaluation: 3, 10, or 20 repeats reduce variance where it matters most. Split indices are persisted once and shared by Python and browser runners for exact, leakage-safe comparisons.

Zero Hyperparameter Tuning

Unlike gradient-boosted trees (XGBoost/LightGBM) that require extensive grid searches and feature encodings, TabICLv2 ingests raw tabular rows in-context with zero manual preprocessing.

Deterministic Local Execution

Benchmark runs execute 100% client-side via WebGPU shader pipelines with fixed random seeds, ensuring reproducible performance within standard browser constraints.

OPEN DATASETS & MODELS

Export, Reproduce & Train Locally

All 21 demo datasets are published as versioned CuratedContainers in the Hugging Face Hub repository carlahq/demo-tabular-benchmark-containers, with Parquet data, task metadata, and persisted evaluation splits.