21 Real-World Datasets & In-Browser Benchmarks
Comprehensive performance evaluation of TabICLv2 executing 100% client-side inside Google Sheets via WebGPU tensor kernels. Zero cloud roundtrips, zero preprocessing, and native integration with LLMs via Model Context Protocol (MCP).
Machine learning use cases by domain
Start with a business question. Each guide connects it to useful data, real benchmark results, and a Google Sheets workflow you can try.
Benchmark Results & Dataset Directory
Precomputed, deterministic IID and non-IID splits evaluate in-browser TabICLv2 latency, throughput, and out-of-sample accuracy against reference baselines.
| Dataset | Domain | Task | Rows | Cols | Splits?Adaptive repeated 3-fold IID evaluation. Smaller datasets receive more repeats to reduce variance; every split is precomputed for identical Python and browser evaluation.Explore Split Protocol → | Score?Primary benchmark error loss: 1 - AUC for binary classification, Log-Loss for multiclass, RMSE for regression (lower is better for all tasks). Evaluated on the dataset's persisted split indices.Explore Methodology Guide → | OpenML?Top-ranked baseline score from traditional tree ensembles (XGBoost, Random Forest) on OpenML.Explore OpenML Reference → | Latency?Mean inference time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Learn about Latency → | Sheet |
|---|---|---|---|---|---|---|---|---|---|
| Abalone | Agriculture, Forestry & Fishing | REG | 4,177 | 8 | 3×3 | 2.03 | — | 11.7s | |
| Adult | Economics & Public Policy | BIN | 48,842 | 14 | 3×3 | 0.0793 | 0.1067 | 273.2s | |
| Airfoil Self Noise | UCI | REG | 1,503 | 5 | 10×3 | 1.12 | — | 4.1s | |
| Amazon Employee Access | Information Technology & Enterprise Security | BIN | 32,769 | 9 | 3×3 | 0.1483 | 0.3362 | 129.2s | |
| Bank Marketing | Finance & Banking | BIN | 45,211 | 15 | 3×3 | 0.1970 | 0.0642 | 235.4s | |
| Blood Transfusion Service Center | Healthcare & Biomedicine | BIN | 748 | 4 | 20×3 | 0.2446 | 0.2435 | 1.1s | |
| Breast W | Healthcare & Life Sciences | BIN | 699 | 9 | 20×3 | 0.0052 | 0.0282 | 1.1s | |
| Car | Automotive & Fleet Management | MULTI | 1,728 | 6 | 10×3 | 0.0231 | 0.0000 | 1.5s | |
| Compas Two Years | Legal & Public Safety | BIN | 5,278 | 13 | 3×3 | 0.2701 | — | 5.8s | |
| Credit G | Finance & Banking | BIN | 1,000 | 20 | 10×3 | 0.2016 | 0.2219 | 1.4s | |
| Diabetes | Healthcare & Biomedicine | BIN | 768 | 8 | 10×3 | 0.1632 | 0.1810 | 1.2s | |
| Employee Salaries | Human Resources & Workforce Analytics | REG | 9,228 | 10 | 3×3 | 8,126 | — | 36.3s | |
| Fitness Club | Fitness, Sports & Recreation | BIN | 1,500 | 6 | 10×3 | 0.1795 | — | 1.3s | |
| House Sales | Real Estate & Property Valuation | REG | 21,613 | 20 | 3×3 | 111,524 | — | 62.7s | |
| Houses | Real Estate & Urban Planning | REG | 20,640 | 8 | 3×3 | 40,971 | — | 50.3s | |
| Monks Problems 2 | Data Science & Artificial Intelligence | BIN | 601 | 6 | 20×3 | 0.0000 | 0.0000 | 1.0s | |
| Phoneme | Speech Processing & Acoustic Engineering | BIN | 5,404 | 5 | 3×3 | 0.0278 | 0.0326 | 5.1s | |
| Spambase | Cybersecurity & IT Infrastructure | BIN | 4,601 | 57 | 3×3 | 0.0075 | 0.0118 | 10.9s | |
| Telco Customer Churn | Telecommunications & Subscription Services | BIN | 7,043 | 20 | 3×3 | 0.1488 | — | 10.2s | |
| Titanic | Maritime Safety & Actuarial Risk Analysis | BIN | 1,309 | 11 | 10×3 | 0.1229 | — | 1.3s | |
| Vehicle | Automotive Engineering & Computer Vision | MULTI | 846 | 18 | 10×3 | 0.2401 | 0.0397 | 1.3s |
Benchmark Methodology & Runtime Specs
How TabICLv2 in-browser foundation models are evaluated under zero-trust, client-side constraints.
Local WebGPU & WASM Runtime
Inference runs directly on device hardware using WebGPU compute shaders and WebAssembly SIMD kernels. Tabular attention passes execute locally without sending row data to an external server.
IID & Non-IID Evaluation
IID datasets use size-aware repeated 3-fold evaluation: 3, 10, or 20 repeats reduce variance where it matters most. Split indices are persisted once and shared by Python and browser runners for exact, leakage-safe comparisons.
Zero Hyperparameter Tuning
Unlike gradient-boosted trees (XGBoost/LightGBM) that require extensive grid searches and feature encodings, TabICLv2 ingests raw tabular rows in-context with zero manual preprocessing.
Deterministic Local Execution
Benchmark runs execute 100% client-side via WebGPU shader pipelines with fixed random seeds, ensuring reproducible performance within standard browser constraints.
Export, Reproduce & Train Locally
All 21 demo datasets are published as versioned CuratedContainers in the Hugging Face Hub repository carlahq/demo-tabular-benchmark-containers, with Parquet data, task metadata, and persisted evaluation splits.