Skip to main content
CARLA HQ
Economics & Public Policy Public Services & Risk use casesNACE M72Binary ClassificationTarget: income_bracketOpenML #1590 ↗

Adult

Benchmark demographic dataset consisting of 48,842 records extracted from the 1994 US Census Bureau database, designed to predict whether an individual's annual income exceeds $50,000.

✨ Try in CarlaInstall Chrome Extension ↗100% In-Browser WebGPU • Zero Cloud Upload
48,842Records (Rows)
14Predictive Features
6 / 9Numeric / Categorical
0.9%Missing Value Ratio
0.07931 - AUC Error
273.20sWebGPU Mean Split
BUSINESS CONTEXT & OBJECTIVEMarket Research & Socioeconomic Analytics (M72)

Business Objective: Adult

01Business Context

Financial institutions, marketers, and public policy researchers require granular socioeconomic indicators to evaluate creditworthiness, segment consumer demographics, and assess workforce equity across regions without intrusive direct income disclosures.

02Analytical Objective

Predict whether an individual earns more than $50,000 annually based on demographic, occupational, and capital investment characteristics.

03Economic & Decision Impact

Enables precise customer lifetime value estimation, targeted financial product marketing, and automated wealth tiering while avoiding costly underwriting errors and mitigating demographic model bias.

ML Benchmark Narrative

As one of the most established tabular benchmarks in machine learning, the Adult dataset achieves ~87% accuracy with tuned gradient boosted trees (LightGBM/XGBoost). Modern zero-shot tabular foundation models (TabICL / Carla) match or exceed tree-based baselines with zero parameter tuning directly in-context.

Source Origin:UCI Machine Learning Repository / US Census Bureau
Creator:Ronny Kohavi and Barry Becker (1996)
License:Creative Commons Attribution 4.0 International (CC BY 4.0)
FEATURE SPECIFICATION & DATA DICTIONARY

Target & Feature Column Definitions

Exhaustive business definitions, measurement units, roles, and target variables across all 15 columns.

ColVariable / Feature NameRoleTypeUnit / ScaleDescription & Business Meaning
Aincome_bracketIncome BracketTARGETCAT
Annual income tier classification indicating whether total income is <=50K or >50K.
Bage_yearsAgeFeatureNUMyears
Age of the individual in years.
Cemployment_sectorEmployment SectorFeatureCAT
General sector and type of employer (e.g., Private, Self-emp-not-inc, Self-emp-inc, Federal-gov, Local-gov, State-gov, Without-pay, Never-worked).
Ddemographic_weightFinal Demographic WeightFeatureNUMratio
Statistical sampling weight assigned by the US Census Bureau representing population control totals across demographic segments within each state.
Eeducation_levelHighest Education LevelFeatureCAT
Highest level of formal education completed (e.g., Bachelors, Some-college, 11th, HS-grad, Prof-school, Assoc-acdm, Assoc-voc, 9th, 7th-8th, 12th, Masters, 1st-4th, 10th, Doctorate, 5th-6th, Preschool).
Feducation_yearsEducation Level (Numeric)FeatureNUMyears
Total completed years/stages of formal education represented on an ordinal numerical scale (1-16).
Gmarital_statusMarital StatusFeatureCAT
Marital status of the individual (e.g., Married-civ-spouse, Divorced, Never-married, Separated, Widowed, Married-spouse-absent, Married-AF-spouse).
Hjob_occupationOccupationFeatureCAT
Primary job role or professional category (e.g., Tech-support, Craft-repair, Other-service, Sales, Exec-managerial, Prof-specialty, Handlers-cleaners, Machine-op-inspct, Adm-clerical, Farming-fishing, Transport-moving, Priv-house-serv, Protective-serv, Armed-Forces).
Ihousehold_relationshipHousehold RelationshipFeatureCAT
Role of the individual within their household (e.g., Wife, Own-child, Husband, Not-in-family, Other-relative, Unmarried).
Jethnicity_raceRace / EthnicityFeatureCAT
Self-reported racial background (e.g., White, Asian-Pac-Islander, Amer-Indian-Eskimo, Other, Black).
KgenderGenderFeatureCAT
Biological sex of the individual (Male, Female).
Lannual_capital_gain_usdCapital GainsFeatureNUMUSD
Recorded annual income gains derived from investments and capital asset sales.
Mannual_capital_loss_usdCapital LossesFeatureNUMUSD
Recorded annual capital losses incurred from investments and asset sales.
Nweekly_work_hoursWeekly Work HoursFeatureNUMhours
Reported average working hours per week.
Ocountry_of_originCountry of OriginFeatureCAT
Country of birth or nationality of the individual.
DATASET PREVIEW

Interactive Data Table

Explore rows, feature values, and target labels for Adult.

Loading dataset...
LOCAL EXECUTION TELEMETRY

TabICLv2 WebGPU Benchmark Results

Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.

Mean Split Latency?Mean wall-clock execution time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Explore Latency Guide →
273.20s9 splits • 8 ensembles/split
Total Evaluation Runtime?Total wall-clock duration across all selected splits, including in-context encoding, WebGPU shader execution, and result aggregation.Explore Split Protocol →
2458.80sIncludes warmup & sync
Evaluation Protocol?Precomputed row indices ensure the browser and Python runners evaluate identical out-of-sample observations without leakage.Explore Split Protocol →
3×3 SplitsRepeated IID • 9 total
TabICLv2 Score?Primary task performance score (1 - AUC Error) achieved by TabICLv2 across out-of-sample evaluation splits.Explore Metric Formula →
0.07931 - AUC Error
Evaluation SplitTrain RowsTest Rows
Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more →
Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details →
F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more →
Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details →
Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details →
R1 / F132,56116,281274.34s86.76%0.92330.69240.78000.6224
R1 / F232,56116,281274.10s86.95%0.92600.70570.76690.6535
R1 / F332,56216,280272.09s86.62%0.92220.69790.75860.6462
R2 / F132,56116,281273.47s87.06%0.92550.70430.77700.6440
R2 / F232,56116,281273.45s86.38%0.92240.67710.78210.5970
R2 / F332,56216,280272.00s83.21%0.91040.68630.62060.7677
R3 / F132,56116,281273.66s86.28%0.92130.68370.76250.6196
R3 / F232,56116,281273.55s86.93%0.92620.69940.77760.6355
R3 / F332,56216,280272.14s86.82%0.92410.69970.76940.6416
Mean ± Std273.20s86.33% ± 1.13%0.922 ± 0.0050.694 ± 0.0090.7550.648
UNIFIED BENCHMARK LEADERBOARD

Foundation Models vs Traditional ML Baselines

Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.

RankAlgorithm / ModelModel FamilyRuntime
Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more →
F-Measure?Harmonic mean of precision and recall.Learn more →
Reference / Repo
#1Google TabFM v1.0Tabular Foundation ModelPyTorch/Python87.60%0.93230.7162google-research/tabfm ↗
#2EXAONE TabularTabular Foundation ModelPyTorch/Python87.58%0.93200.7171LGAI-Research/EXAONE-Tabular ↗
#3nanotabicl VanillaTabular Foundation ModelPyTorch/Python86.83%0.92370.6992soda-inria/nanotabicl ↗
#4FilteredClassifier J48Decision TreePython86.72%0.89330.8615OpenML #1785495 ↗
#5FilteredClassifier MultiSearch LogitBoost REPTreeLogistic / Linear ModelPython86.63%0.91990.8603OpenML #1838085 ↗
#6DTNBMachine Learning ModelPython86.43%0.91780.8617OpenML #1810088 ↗
#7SimpleCartMachine Learning ModelPython86.42%0.89010.8573OpenML #1666057 ↗
#8TabPFN v3Tabular Foundation ModelPyTorch/Python86.38%0.92010.6864PriorLabs/tabpfn ↗
#9Carla Engine (TabICLv2 WebGPU)Tabular Foundation ModelBrowser86.33%0.92240.6941100% In-Browser
#10NBTreeDecision TreePython86.32%0.90780.8599OpenML #1666083 ↗
#11FilteredClassifier MultilayerPerceptronNeural Network (MLP)Python86.30%0.88930.8566OpenML #1799965 ↗
#12Full TabICLv2 (PyTorch Reference)Tabular Foundation ModelPyTorch/Python86.23%0.92070.6875soda-inria/tabicl ↗
#13Streaming nanotabiclTabular Foundation ModelPyTorch/Python86.18%0.91780.6817soda-inria/nanotabicl ↗
#14J48Decision TreePython86.17%0.89330.8553OpenML #1687985 ↗
#15END ND J48Decision TreePython86.11%0.88920.8552OpenML #1815759 ↗
#16A2DEMachine Learning ModelPython86.07%0.92510.8628OpenML #1664564 ↗
#17Bagging J48Decision TreePython86.00%0.90790.8547OpenML #1666377 ↗
STATISTICAL DISTRIBUTIONS

Variable Schema & Summary Distributions

Observed numerical ranges, category cardinalities, missing rates, and sample values across 48,842 rows.

ColVariable NameTypeMissingDistinctSummary Stats / DistributionSample Values
Aincome_bracketTARGETCAT022 cats: <=50K, >50K<=50K<=50K>50K
Bage_yearsNUM074[17, 90] μ=38.6 σ=13.7253828
Cemployment_sectorCAT2,799 (5.73%)88 cats: Private, Self-emp-not-inc, Local-gov +5 morePrivatePrivateLocal-gov
Ddemographic_weightNUM028,523[12285, 1490400] μ=189664.1 σ=105602.922680289814336951
Eeducation_levelCAT01610 cats: HS-grad, Some-college, Bachelors +7 more11thHS-gradAssoc-acdm
Feducation_yearsNUM016[1, 16] μ=10.1 σ=2.67912
Gmarital_statusCAT077 cats: Married-civ-spouse, Never-married, Divorced +4 moreNever-marriedMarried-civ-spouseMarried-civ-spouse
Hjob_occupationCAT2,809 (5.75%)1410 cats: Prof-specialty, Craft-repair, Exec-managerial +7 moreMachine-op-inspctFarming-fishingProtective-serv
Ihousehold_relationshipCAT066 cats: Husband, Not-in-family, Own-child +3 moreOwn-childHusbandHusband
Jethnicity_raceCAT055 cats: White, Black, Asian-Pac-Islander +2 moreBlackWhiteWhite
KgenderCAT022 cats: Male, FemaleMaleMaleMale
Lannual_capital_gain_usdNUM0123[0, 99999] μ=1079.1 σ=7451.9000
Mannual_capital_loss_usdNUM099[0, 4356] μ=87.5 σ=403.0000
Nweekly_work_hoursNUM096[1, 99] μ=40.4 σ=12.4405040
Ocountry_of_originCAT857 (1.75%)4110 cats: United-States, Mexico, Philippines +7 moreUnited-StatesUnited-StatesUnited-States
FREQUENTLY ASKED QUESTIONS

Frequently Asked Questions: Adult

Common questions regarding the Adult dataset, machine learning task formulations, and in-browser tabular inference.

What is the Adult dataset used for?

The Adult dataset is an established tabular machine learning benchmark extracted from the 1994 US Census, primarily used to evaluate classification algorithms on predicting whether an individual earns more than $50,000 annually.

Can I run zero-server predictions on this dataset in Google Sheets?

Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models.

What machine learning models perform best on the Adult dataset?

Gradient boosted decision trees such as XGBoost and LightGBM alongside modern tabular foundation models like TabICL routinely achieve top-tier performance (~87% accuracy / 0.92+ ROC-AUC) on this benchmark.

LIVE EVALUATION IN GOOGLE SHEETS

Test TabICLv2 on Adult Yourself

Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.

Step 1

Launch Carla

Click once to open the spreadsheet and Carla side panel together.

Step 2

Review the Setup

Carla selects the dataset target and task from this page automatically.

Step 3

Evaluate & Predict

Run predictions and compute metrics with zero server uploads.