Skip to main content
CARLA HQ
Healthcare & Biomedicine Healthcare & Life Sciences use casesNACE Q86Binary ClassificationTarget: diabetes_diagnosisOpenML #37 ↗

Diabetes

Clinical diagnostic benchmark comprising 768 patient records designed to predict the onset of diabetes mellitus in high-risk adult females based on diagnostic and metabolic biomarkers.

✨ Try in CarlaInstall Chrome Extension ↗100% In-Browser WebGPU • Zero Cloud Upload
768Records (Rows)
8Predictive Features
8 / 1Numeric / Categorical
0.0%Missing Value Ratio
0.16321 - AUC Error
1.16sWebGPU Mean Split
BUSINESS CONTEXT & OBJECTIVEClinical Diagnostics & Preventative Healthcare (Q86)

Business Objective: Diabetes

01Business Context

Primary care physicians, metabolic health clinics, and health insurance providers seek early-stage risk stratifications to identify individuals at elevated risk of Type 2 diabetes before irreversible vascular or organ damage occurs.

02Analytical Objective

Accurately predict binary diabetes onset (tested_positive vs. tested_negative) utilizing non-invasive demographic factors and routine clinical lab measurements.

03Economic & Decision Impact

Enables proactive preventative interventions (lifestyle programs, Metformin therapy) reducing long-term chronic care expenditure by up to 60%. Mitigates severe false negatives (delayed diagnosis leading to neuropathy, nephropathy, and cardiovascular events) while controlling false positives to avoid unnecessary clinical anxiety and second-line diagnostic costs.

ML Benchmark Narrative

On this classic 768-row tabular benchmark, optimized tree ensembles (LightGBM, XGBoost, Random Forest) traditionally achieve ROC-AUC scores between 0.82 and 0.84, often hampered by noisy zero-value physiological readings. TabICL and Carla in-context foundation models achieve competitive zero-shot discrimination (AUC ~0.83) instantly without hyperparameter tuning, handling sparse physiological representations directly in-context.

Source Origin:National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) / UCI Machine Learning Repository / OpenML
Creator:Vincent Sigillito (Johns Hopkins University Applied Physics Laboratory) & NIDDK researchers (1990)
License:Open Data / Public Domain (CC BY 4.0 compatible)
FEATURE SPECIFICATION & DATA DICTIONARY

Target & Feature Column Definitions

Exhaustive business definitions, measurement units, roles, and target variables across all 9 columns.

ColVariable / Feature NameRoleTypeUnit / ScaleDescription & Business Meaning
Adiabetes_diagnosisDiabetes Diagnosis OutcomeTARGETCAT
Binary diagnostic indicator reflecting whether the patient was tested positive ('tested_positive') or negative ('tested_negative') for diabetes according to WHO criteria.
Bpregnancy_countNumber of PregnanciesFeatureNUMcount
Total number of times the patient has been pregnant; associated with gestational metabolic strain and insulin resistance.
Cglucose_concentration_2h2-Hour Plasma GlucoseFeatureNUMmg/dL
Plasma glucose concentration measured 2 hours after an oral glucose tolerance test (OGTT). Key diagnostic indicator where values >= 200 mg/dL indicate diabetes.
Ddiastolic_blood_pressureDiastolic Blood PressureFeatureNUMmm Hg
Diastolic blood pressure reading; elevated values are correlated with metabolic syndrome and systemic vascular stiffness.
Etriceps_skinfold_thicknessTriceps Skinfold ThicknessFeatureNUMmm
Measurement of the fold of skin and subcutaneous fat over the triceps muscle, serving as an anthropometric proxy for peripheral body fat percentage.
Fserum_insulin_2h2-Hour Serum InsulinFeatureNUMmu U/ml
Serum insulin level measured 2 hours after glucose loading; assesses pancreatic beta-cell response and peripheral insulin sensitivity.
Gbody_mass_indexBody Mass Index (BMI)FeatureNUMkg/m²
Body mass index calculated as weight in kilograms divided by height in meters squared (kg/m²), indicating obesity levels.
Hdiabetes_pedigree_functionDiabetes Pedigree FunctionFeatureNUMscore
Continuous genetic scoring function synthesizing family history of diabetes to estimate genetic risk predisposition.
Ipatient_age_yearsPatient AgeFeatureNUMyears
Chronological age of the patient in years (all subjects are females >= 21 years old).
DATASET PREVIEW

Interactive Data Table

Explore rows, feature values, and target labels for Diabetes.

Loading dataset...
LOCAL EXECUTION TELEMETRY

TabICLv2 WebGPU Benchmark Results

Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.

Mean Split Latency?Mean wall-clock execution time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Explore Latency Guide →
1.16s30 splits • 8 ensembles/split
Total Evaluation Runtime?Total wall-clock duration across all selected splits, including in-context encoding, WebGPU shader execution, and result aggregation.Explore Split Protocol →
34.86sIncludes warmup & sync
Evaluation Protocol?Precomputed row indices ensure the browser and Python runners evaluate identical out-of-sample observations without leakage.Explore Split Protocol →
10×3 SplitsRepeated IID • 30 total
TabICLv2 Score?Primary task performance score (1 - AUC Error) achieved by TabICLv2 across out-of-sample evaluation splits.Explore Metric Formula →
0.16321 - AUC Error
Evaluation SplitTrain RowsTest Rows
Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more →
Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details →
F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more →
Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details →
Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details →
R1 / F15122561.17s78.13%0.86590.65430.73610.5889
R1 / F25122561.21s76.95%0.83400.62420.72060.5506
R1 / F35122561.19s75.39%0.81870.64410.64770.6404
R2 / F15122561.17s75.00%0.82750.59490.69120.5222
R2 / F25122561.24s78.91%0.87540.65380.76120.5730
R2 / F35122561.15s73.44%0.81510.59520.63290.5618
R3 / F15122561.15s75.78%0.82400.65170.65910.6444
R3 / F25122561.15s75.78%0.84830.60260.70150.5281
R3 / F35122561.21s75.78%0.83650.59210.71430.5056
R4 / F15122561.14s76.56%0.84010.62960.70830.5667
R4 / F25122561.16s71.88%0.81160.55000.61970.4944
R4 / F35122561.21s78.13%0.84980.66670.70890.6292
R5 / F15122561.15s73.05%0.81380.59650.62960.5667
R5 / F25122561.15s78.13%0.85170.66270.71430.6180
R5 / F35122561.14s78.13%0.85410.64560.73910.5730
R6 / F15122561.14s78.13%0.86620.65000.74290.5778
R6 / F25122561.19s73.44%0.80180.59520.63290.5618
R6 / F35122561.13s75.39%0.83210.59870.69120.5281
R7 / F15122561.14s76.95%0.84070.65090.69620.6111
R7 / F25122561.14s76.56%0.82760.62960.69860.5730
R7 / F35122561.15s76.56%0.85930.61040.72310.5281
R8 / F15122561.17s75.00%0.82050.61450.67110.5667
R8 / F25122561.15s76.95%0.84560.63350.70830.5730
R8 / F35122561.16s77.34%0.84580.61840.74600.5281
R9 / F15122561.15s76.17%0.82740.65140.67060.6333
R9 / F25122561.15s75.78%0.84950.59740.70770.5169
R9 / F35122561.15s74.61%0.84180.59630.66670.5393
R10 / F15122561.17s76.95%0.84200.63350.71830.5667
R10 / F25122561.15s72.66%0.80220.57320.62670.5281
R10 / F35122561.15s76.17%0.84230.63470.67950.5955
Mean ± Std1.16s75.99% ± 1.75%0.837 ± 0.0180.622 ± 0.0280.6920.566
UNIFIED BENCHMARK LEADERBOARD

Foundation Models vs Traditional ML Baselines

Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.

RankAlgorithm / ModelModel FamilyRuntime
Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more →
F-Measure?Harmonic mean of precision and recall.Learn more →
Reference / Repo
#1AttributeSelectedClassifier HoeffdingTreeDecision TreePython77.34%0.81900.7665OpenML #577230 ↗
#2ClassificationViaRegression M5PDecision TreePython77.34%0.83040.7656OpenML #578402 ↗
#3AttributeSelectedClassifier SMO PolyKernelSupport Vector Machine (SVM)Python77.08%0.71320.7592OpenML #575789 ↗
#4AttributeSelectedClassifier Bagging JRipBagging EnsemblePython77.08%0.79840.7615OpenML #576370 ↗
#5AttributeSelectedClassifier Bagging LMTBagging EnsemblePython77.08%0.82480.7666OpenML #576534 ↗
#6AttributeSelectedClassifier LMTMachine Learning ModelPython77.08%0.82120.7622OpenML #577469 ↗
#7RandomForestRandom ForestPython76.95%0.83120.7679OpenML #573702 ↗
#8Bagging REPTreeDecision TreePython76.95%0.82700.7648OpenML #574537 ↗
#9KernelLogisticRegression RBFKernelLogistic / Linear ModelPython76.82%0.83810.7618OpenML #323386 ↗
#10SMO PolyKernelSupport Vector Machine (SVM)Python76.82%0.71290.7577OpenML #550234 ↗
#11TabPFN v3Tabular Foundation ModelPyTorch/Python76.73%0.83930.6354PriorLabs/tabpfn ↗
#12EXAONE TabularTabular Foundation ModelPyTorch/Python76.50%0.83620.6332LGAI-Research/EXAONE-Tabular ↗
#13Google TabFM v1.0Tabular Foundation ModelPyTorch/Python76.42%0.83770.6335google-research/tabfm ↗
#14nanotabicl VanillaTabular Foundation ModelPyTorch/Python76.20%0.83680.6252soda-inria/nanotabicl ↗
#15Streaming nanotabiclTabular Foundation ModelPyTorch/Python76.16%0.83680.6248soda-inria/nanotabicl ↗
#16Full TabICLv2 (PyTorch Reference)Tabular Foundation ModelPyTorch/Python76.16%0.83680.6248soda-inria/tabicl ↗
#17Carla Engine (TabICLv2 WebGPU)Tabular Foundation ModelBrowser75.99%0.83700.6217100% In-Browser
STATISTICAL DISTRIBUTIONS

Variable Schema & Summary Distributions

Observed numerical ranges, category cardinalities, missing rates, and sample values across 768 rows.

ColVariable NameTypeMissingDistinctSummary Stats / DistributionSample Values
Adiabetes_diagnosisTARGETCAT022 cats: tested_negative, tested_positivetested_positivetested_negativetested_positive
Bpregnancy_countNUM017[0, 17] μ=3.8 σ=3.4618
Cglucose_concentration_2hNUM0136[0, 199] μ=120.9 σ=32.014885183
Ddiastolic_blood_pressureNUM047[0, 122] μ=69.1 σ=19.3726664
Etriceps_skinfold_thicknessNUM051[0, 99] μ=20.5 σ=15.935290
Fserum_insulin_2hNUM0186[0, 846] μ=79.8 σ=115.2000
Gbody_mass_indexNUM0248[0, 67.1] μ=32.0 σ=7.933.626.623.3
Hdiabetes_pedigree_functionNUM0517[0.078, 2.42] μ=0.5 σ=0.30.6270.3510.672
Ipatient_age_yearsNUM052[21, 81] μ=33.2 σ=11.8503132
FREQUENTLY ASKED QUESTIONS

Frequently Asked Questions: Diabetes

Common questions regarding the Diabetes dataset, machine learning task formulations, and in-browser tabular inference.

What is the Diabetes (Pima Indians) dataset used for?

The Diabetes dataset is a standard diagnostic machine learning benchmark containing 768 patient cases used to train and evaluate classification algorithms for predicting the onset of Type 2 diabetes based on 8 clinical and metabolic features.

Can I run zero-server predictions on this dataset in Google Sheets?

Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models, allowing instant scoring without manual model training or backend infrastructure.

What machine learning models perform best on the Diabetes dataset?

Gradient boosting frameworks like LightGBM, XGBoost, and CatBoost alongside in-context foundation models like TabICL typically achieve the strongest classification performance, yielding ROC-AUC scores in the 0.82-0.84 range.

LIVE EVALUATION IN GOOGLE SHEETS

Test TabICLv2 on Diabetes Yourself

Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.

Step 1

Launch Carla

Click once to open the spreadsheet and Carla side panel together.

Step 2

Review the Setup

Carla selects the dataset target and task from this page automatically.

Step 3

Evaluate & Predict

Run predictions and compute metrics with zero server uploads.

Provenance & Attribution

https://www.jair.org/index.php/jair/article/view/10129

License: PublicData Source: OpenMLOpenML Page: https://www.openml.org/d/37