Skip to main content
CARLA HQ
Cybersecurity & IT Infrastructure Commerce & Industry use casesNACE J62.09Binary ClassificationTarget: is_spamOpenML #44 ↗

Spambase

The Spambase dataset contains 4,601 email records characterized by word, character, and capital letter run-length frequencies to classify incoming messages as spam or legitimate ham.

✨ Try in CarlaInstall Chrome Extension ↗100% In-Browser WebGPU • Zero Cloud Upload
4,601Records (Rows)
57Predictive Features
58 / 0Numeric / Categorical
0.0%Missing Value Ratio
0.00751 - AUC Error
10.85sWebGPU Mean Split
BUSINESS CONTEXT & OBJECTIVEEnterprise Software & Cybersecurity (J62.09)

Business Objective: Spambase

01Business Context

Modern enterprise email systems face high volumes of unsolicited commercial emails, phishing attempts, and bulk advertisements that congest mail servers, drain employee productivity, and expose corporate infrastructure to security risks.

02Analytical Objective

Classify incoming emails in real-time as spam (1) or non-spam/ham (0) based on token frequencies, character distributions, and capitalization patterns.

03Economic & Decision Impact

Effective automated filtering saves thousands of knowledge worker hours while mitigating cyber risk; however, false positives (flagging critical business communications as spam) create immediate business disruption, requiring calibrated high-precision decision thresholds.

ML Benchmark Narrative

Classic gradient-boosted trees (XGBoost, LightGBM) and Random Forests typically reach 94%–95.5% accuracy on Spambase by exploiting non-linear interactions across capital run-lengths and spam keywords. Zero-shot in-context tabular foundation models (TabICL / Carla) achieve competitive discrimination out-of-the-box directly within spreadsheet workflows without manual feature scaling or threshold tuning.

Source Origin:UCI Machine Learning Repository & OpenML
Creator:Mark Hopkins, Erik Reeber, George Forman, Jaap Suermondt (Hewlett-Packard Laboratories) (1999)
License:Creative Commons Attribution 4.0 International (CC BY 4.0)
FEATURE SPECIFICATION & DATA DICTIONARY

Target & Feature Column Definitions

Exhaustive business definitions, measurement units, roles, and target variables across all 58 columns.

ColVariable / Feature NameRoleTypeUnit / ScaleDescription & Business Meaning
Ais_spamSpam ClassificationTARGETNUM
Binary ground-truth email classification: 1 for spam (unsolicited commercial email), 0 for legitimate email (ham).
Bword_freq_makeWord Frequency: 'make'FeatureNUM%
Percentage of words in the email matching 'make' (100 * count / total words).
Cword_freq_addressWord Frequency: 'address'FeatureNUM%
Percentage of words in the email matching 'address'.
Dword_freq_allWord Frequency: 'all'FeatureNUM%
Percentage of words in the email matching 'all'.
Eword_freq_3dWord Frequency: '3d'FeatureNUM%
Percentage of words in the email matching '3d'.
Fword_freq_ourWord Frequency: 'our'FeatureNUM%
Percentage of words in the email matching 'our'.
Gword_freq_overWord Frequency: 'over'FeatureNUM%
Percentage of words in the email matching 'over'.
Hword_freq_removeWord Frequency: 'remove'FeatureNUM%
Percentage of words in the email matching 'remove' (common unsubscribe spam indicator).
Iword_freq_internetWord Frequency: 'internet'FeatureNUM%
Percentage of words in the email matching 'internet'.
Jword_freq_orderWord Frequency: 'order'FeatureNUM%
Percentage of words in the email matching 'order'.
Kword_freq_mailWord Frequency: 'mail'FeatureNUM%
Percentage of words in the email matching 'mail'.
Lword_freq_receiveWord Frequency: 'receive'FeatureNUM%
Percentage of words in the email matching 'receive'.
Mword_freq_willWord Frequency: 'will'FeatureNUM%
Percentage of words in the email matching 'will'.
Nword_freq_peopleWord Frequency: 'people'FeatureNUM%
Percentage of words in the email matching 'people'.
Oword_freq_reportWord Frequency: 'report'FeatureNUM%
Percentage of words in the email matching 'report'.
Pword_freq_addressesWord Frequency: 'addresses'FeatureNUM%
Percentage of words in the email matching 'addresses'.
Qword_freq_freeWord Frequency: 'free'FeatureNUM%
Percentage of words in the email matching 'free' (frequent promotional indicator).
Rword_freq_businessWord Frequency: 'business'FeatureNUM%
Percentage of words in the email matching 'business'.
Sword_freq_emailWord Frequency: 'email'FeatureNUM%
Percentage of words in the email matching 'email'.
Tword_freq_youWord Frequency: 'you'FeatureNUM%
Percentage of words in the email matching 'you'.
Uword_freq_creditWord Frequency: 'credit'FeatureNUM%
Percentage of words in the email matching 'credit'.
Vword_freq_yourWord Frequency: 'your'FeatureNUM%
Percentage of words in the email matching 'your'.
Wword_freq_fontWord Frequency: 'font'FeatureNUM%
Percentage of words in the email matching 'font' (HTML formatting artifact).
Xword_freq_000Word Frequency: '000'FeatureNUM%
Percentage of words in the email matching '000' (large monetary amount indicator).
Yword_freq_moneyWord Frequency: 'money'FeatureNUM%
Percentage of words in the email matching 'money'.
Zword_freq_hpWord Frequency: 'hp'FeatureNUM%
Percentage of words in the email matching 'hp' (Hewlett-Packard workplace indicator).
AAword_freq_hplWord Frequency: 'hpl'FeatureNUM%
Percentage of words in the email matching 'hpl' (HP Labs workplace indicator).
ABword_freq_georgeWord Frequency: 'george'FeatureNUM%
Percentage of words in the email matching 'george' (researcher personal identifier).
ACword_freq_650Word Frequency: '650'FeatureNUM%
Percentage of words in the email matching '650' (Palo Alto / Bay Area area code).
ADword_freq_labWord Frequency: 'lab'FeatureNUM%
Percentage of words in the email matching 'lab'.
AEword_freq_labsWord Frequency: 'labs'FeatureNUM%
Percentage of words in the email matching 'labs'.
AFword_freq_telnetWord Frequency: 'telnet'FeatureNUM%
Percentage of words in the email matching 'telnet'.
AGword_freq_857Word Frequency: '857'FeatureNUM%
Percentage of words in the email matching '857'.
AHword_freq_dataWord Frequency: 'data'FeatureNUM%
Percentage of words in the email matching 'data'.
AIword_freq_415Word Frequency: '415'FeatureNUM%
Percentage of words in the email matching '415' (San Francisco area code).
AJword_freq_85Word Frequency: '85'FeatureNUM%
Percentage of words in the email matching '85'.
AKword_freq_technologyWord Frequency: 'technology'FeatureNUM%
Percentage of words in the email matching 'technology'.
ALword_freq_1999Word Frequency: '1999'FeatureNUM%
Percentage of words in the email matching '1999'.
AMword_freq_partsWord Frequency: 'parts'FeatureNUM%
Percentage of words in the email matching 'parts'.
ANword_freq_pmWord Frequency: 'pm'FeatureNUM%
Percentage of words in the email matching 'pm'.
AOword_freq_directWord Frequency: 'direct'FeatureNUM%
Percentage of words in the email matching 'direct'.
APword_freq_csWord Frequency: 'cs'FeatureNUM%
Percentage of words in the email matching 'cs' (Computer Science indicator).
AQword_freq_meetingWord Frequency: 'meeting'FeatureNUM%
Percentage of words in the email matching 'meeting'.
ARword_freq_originalWord Frequency: 'original'FeatureNUM%
Percentage of words in the email matching 'original'.
ASword_freq_projectWord Frequency: 'project'FeatureNUM%
Percentage of words in the email matching 'project'.
ATword_freq_reWord Frequency: 're'FeatureNUM%
Percentage of words in the email matching 're' (email reply prefix).
AUword_freq_eduWord Frequency: 'edu'FeatureNUM%
Percentage of words in the email matching 'edu' (academic domain suffix).
AVword_freq_tableWord Frequency: 'table'FeatureNUM%
Percentage of words in the email matching 'table'.
AWword_freq_conferenceWord Frequency: 'conference'FeatureNUM%
Percentage of words in the email matching 'conference'.
AXchar_freq_semicolonCharacter Frequency: ';'FeatureNUM%
Percentage of characters in the email matching semicolon ';'.
AYchar_freq_parenthesis_leftCharacter Frequency: '('FeatureNUM%
Percentage of characters in the email matching opening parenthesis '('.
AZchar_freq_bracket_leftCharacter Frequency: '['FeatureNUM%
Percentage of characters in the email matching opening square bracket '['.
BAchar_freq_exclamationCharacter Frequency: '!'FeatureNUM%
Percentage of characters in the email matching exclamation mark '!' (promotional enthusiasm marker).
BBchar_freq_dollarCharacter Frequency: '$'FeatureNUM%
Percentage of characters in the email matching dollar sign '$' (financial incentive marker).
BCchar_freq_hashCharacter Frequency: '#'FeatureNUM%
Percentage of characters in the email matching hash '#'.
BDcapital_run_length_averageCapital Run Length: AverageFeatureNUMcharacters
Average length of uninterrupted sequences of capital letters.
BEcapital_run_length_longestCapital Run Length: LongestFeatureNUMcharacters
Length of the longest uninterrupted sequence of capital letters.
BFcapital_run_length_totalCapital Run Length: TotalFeatureNUMcharacters
Total sum of all uppercase characters occurring across consecutive capital sequences.
DATASET PREVIEW

Interactive Data Table

Explore rows, feature values, and target labels for Spambase.

Loading dataset...
LOCAL EXECUTION TELEMETRY

TabICLv2 WebGPU Benchmark Results

Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.

Mean Split Latency?Mean wall-clock execution time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Explore Latency Guide →
10.85s9 splits • 8 ensembles/split
Total Evaluation Runtime?Total wall-clock duration across all selected splits, including in-context encoding, WebGPU shader execution, and result aggregation.Explore Split Protocol →
97.69sIncludes warmup & sync
Evaluation Protocol?Precomputed row indices ensure the browser and Python runners evaluate identical out-of-sample observations without leakage.Explore Split Protocol →
3×3 SplitsRepeated IID • 9 total
TabICLv2 Score?Primary task performance score (1 - AUC Error) achieved by TabICLv2 across out-of-sample evaluation splits.Explore Metric Formula →
0.00751 - AUC Error
Evaluation SplitTrain RowsTest Rows
Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more →
Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details →
F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more →
Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details →
Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details →
R1 / F13,0671,53411.00s96.15%0.99190.95100.95650.9455
R1 / F23,0671,53410.92s97.26%0.99520.96530.96220.9685
R1 / F33,0681,53310.66s96.15%0.99020.95120.95040.9520
R2 / F13,0671,53410.88s96.68%0.99300.95750.96480.9504
R2 / F23,0671,53410.85s96.54%0.99360.95590.95990.9520
R2 / F33,0681,53310.68s96.22%0.99140.95230.94610.9586
R3 / F13,0671,53411.03s96.35%0.99230.95390.94930.9587
R3 / F23,0671,53411.04s95.96%0.98960.94820.95780.9387
R3 / F33,0681,53310.62s97.26%0.99550.96510.96830.9619
Mean ± Std10.85s96.51% ± 0.45%0.993 ± 0.0020.956 ± 0.0060.9570.954
UNIFIED BENCHMARK LEADERBOARD

Foundation Models vs Traditional ML Baselines

Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.

RankAlgorithm / ModelModel FamilyRuntime
Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more →
F-Measure?Harmonic mean of precision and recall.Learn more →
Reference / Repo
#1Google TabFM v1.0Tabular Foundation ModelPyTorch/Python96.65%0.99340.9576google-research/tabfm ↗
#2EXAONE TabularTabular Foundation ModelPyTorch/Python96.52%0.99240.9556LGAI-Research/EXAONE-Tabular ↗
#3Carla Engine (TabICLv2 WebGPU)Tabular Foundation ModelBrowser96.51%0.99250.9556100% In-Browser
#4Full TabICLv2 (PyTorch Reference)Tabular Foundation ModelPyTorch/Python96.46%0.99250.9550soda-inria/tabicl ↗
#5nanotabicl VanillaTabular Foundation ModelPyTorch/Python96.46%0.99240.9549soda-inria/nanotabicl ↗
#6Streaming nanotabiclTabular Foundation ModelPyTorch/Python96.39%0.99240.9542soda-inria/nanotabicl ↗
#7TabPFN v3Tabular Foundation ModelPyTorch/Python96.25%0.99150.9525PriorLabs/tabpfn ↗
#8RandomForestRandom ForestPython95.61%0.98820.9560OpenML #573705 ↗
#9AdaBoostM1 J48AdaBoostPython95.07%0.98110.9506OpenML #574463 ↗
#10RandomRulesMachine Learning ModelPython94.72%0.98270.9471OpenML #54776 ↗
#11RandomSubSpace REPTreeDecision TreePython94.39%0.98120.9437OpenML #578502 ↗
#12AdaBoostM1 REPTreeAdaBoostPython94.37%0.98280.9437OpenML #574502 ↗
#13AttributeSelectedClassifier RandomForestRandom ForestPython94.28%0.97820.9428OpenML #575720 ↗
#14Bagging JRipBagging EnsemblePython94.20%0.97260.9418OpenML #574620 ↗
#15Bagging J48Decision TreePython94.18%0.97740.9416OpenML #574581 ↗
#16Bagging REPTreeDecision TreePython94.15%0.98060.9414OpenML #604 ↗
#17Bagging REPTreeDecision TreePython94.15%0.98060.9414OpenML #574540 ↗
STATISTICAL DISTRIBUTIONS

Variable Schema & Summary Distributions

Observed numerical ranges, category cardinalities, missing rates, and sample values across 4,601 rows.

ColVariable NameTypeMissingDistinctSummary Stats / DistributionSample Values
Ais_spamTARGETNUM02[0, 1] μ=0.4 σ=0.5111
Bword_freq_makeNUM0142[0, 4.54] μ=0.1 σ=0.300.210.06
Cword_freq_addressNUM0171[0, 14.28] μ=0.2 σ=1.30.640.280
Dword_freq_allNUM0214[0, 5.1] μ=0.3 σ=0.50.640.50.71
Eword_freq_3dNUM043[0, 42.81] μ=0.1 σ=1.4000
Fword_freq_ourNUM0255[0, 10] μ=0.3 σ=0.70.320.141.23
Gword_freq_overNUM0141[0, 5.88] μ=0.1 σ=0.300.280.19
Hword_freq_removeNUM0173[0, 7.27] μ=0.1 σ=0.400.210.19
Iword_freq_internetNUM0170[0, 11.11] μ=0.1 σ=0.400.070.12
Jword_freq_orderNUM0144[0, 5.26] μ=0.1 σ=0.3000.64
Kword_freq_mailNUM0245[0, 18.18] μ=0.2 σ=0.600.940.25
Lword_freq_receiveNUM0113[0, 2.61] μ=0.1 σ=0.200.210.38
Mword_freq_willNUM0316[0, 9.67] μ=0.5 σ=0.90.640.790.45
Nword_freq_peopleNUM0158[0, 5.55] μ=0.1 σ=0.300.650.12
Oword_freq_reportNUM0133[0, 10] μ=0.1 σ=0.300.210
Pword_freq_addressesNUM0118[0, 4.41] μ=0.0 σ=0.300.141.75
Qword_freq_freeNUM0253[0, 20] μ=0.2 σ=0.80.320.140.06
Rword_freq_businessNUM0197[0, 7.14] μ=0.1 σ=0.400.070.06
Sword_freq_emailNUM0229[0, 9.09] μ=0.2 σ=0.51.290.281.03
Tword_freq_youNUM0575[0, 18.75] μ=1.7 σ=1.81.933.471.36
Uword_freq_creditNUM0148[0, 18.18] μ=0.1 σ=0.5000.32
Vword_freq_yourNUM0401[0, 11.11] μ=0.8 σ=1.20.961.590.51
Wword_freq_fontNUM099[0, 17.1] μ=0.1 σ=1.0000
Xword_freq_000NUM0164[0, 5.45] μ=0.1 σ=0.400.431.16
Yword_freq_moneyNUM0143[0, 12.5] μ=0.1 σ=0.400.430.06
Zword_freq_hpNUM0395[0, 20.83] μ=0.5 σ=1.7000
AAword_freq_hplNUM0281[0, 16.66] μ=0.3 σ=0.9000
ABword_freq_georgeNUM0240[0, 33.33] μ=0.8 σ=3.4000
ACword_freq_650NUM0200[0, 9.09] μ=0.1 σ=0.5000
ADword_freq_labNUM0156[0, 14.28] μ=0.1 σ=0.6000
AEword_freq_labsNUM0179[0, 5.88] μ=0.1 σ=0.5000
AFword_freq_telnetNUM0128[0, 12.5] μ=0.1 σ=0.4000
AGword_freq_857NUM0106[0, 4.76] μ=0.0 σ=0.3000
AHword_freq_dataNUM0184[0, 18.18] μ=0.1 σ=0.6000
AIword_freq_415NUM0110[0, 4.76] μ=0.0 σ=0.3000
AJword_freq_85NUM0177[0, 20] μ=0.1 σ=0.5000
AKword_freq_technologyNUM0159[0, 7.69] μ=0.1 σ=0.4000
ALword_freq_1999NUM0188[0, 6.89] μ=0.1 σ=0.400.070
AMword_freq_partsNUM053[0, 8.33] μ=0.0 σ=0.2000
ANword_freq_pmNUM0163[0, 11.11] μ=0.1 σ=0.4000
AOword_freq_directNUM0125[0, 4.76] μ=0.1 σ=0.3000.06
APword_freq_csNUM0108[0, 7.14] μ=0.0 σ=0.4000
AQword_freq_meetingNUM0186[0, 14.28] μ=0.1 σ=0.8000
ARword_freq_originalNUM0136[0, 3.57] μ=0.0 σ=0.2000.12
ASword_freq_projectNUM0160[0, 20] μ=0.1 σ=0.6000
ATword_freq_reNUM0230[0, 21.42] μ=0.3 σ=1.0000.06
AUword_freq_eduNUM0227[0, 22.05] μ=0.2 σ=0.9000.06
AVword_freq_tableNUM038[0, 2.17] μ=0.0 σ=0.1000
AWword_freq_conferenceNUM0106[0, 10] μ=0.0 σ=0.3000
AXchar_freq_semicolonNUM0313[0, 4.385] μ=0.0 σ=0.2000.01
AYchar_freq_parenthesis_leftNUM0641[0, 9.752] μ=0.1 σ=0.300.1320.143
AZchar_freq_bracket_leftNUM0225[0, 4.081] μ=0.0 σ=0.1000
BAchar_freq_exclamationNUM0964[0, 32.478] μ=0.3 σ=0.80.7780.3720.276
BBchar_freq_dollarNUM0504[0, 6.003] μ=0.1 σ=0.200.180.184
BCchar_freq_hashNUM0316[0, 19.829] μ=0.0 σ=0.400.0480.01
BDcapital_run_length_averageNUM02,161[1, 1102.5] μ=5.2 σ=31.73.7565.1149.821
BEcapital_run_length_longestNUM0271[1, 9989] μ=52.2 σ=194.961101485
BFcapital_run_length_totalNUM0919[1, 15841] μ=283.3 σ=606.327810282259
FREQUENTLY ASKED QUESTIONS

Frequently Asked Questions: Spambase

Common questions regarding the Spambase dataset, machine learning task formulations, and in-browser tabular inference.

What is the Spambase dataset used for?

Spambase is a canonical machine learning benchmark containing 4,601 email records designed to train and evaluate binary classification models that distinguish spam from legitimate non-spam emails based on token and character frequencies.

Can I run zero-server predictions on this dataset in Google Sheets?

Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models.

What machine learning models perform best on Spambase?

Gradient-boosted decision trees (XGBoost, LightGBM, CatBoost) and modern in-context foundation models (TabICL / Carla) achieve leading accuracy (94%–95.5%) by capturing non-linear interactions across capital run-length metrics and key token distributions.

LIVE EVALUATION IN GOOGLE SHEETS

Test TabICLv2 on Spambase Yourself

Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.

Step 1

Launch Carla

Click once to open the spreadsheet and Carla side panel together.

Step 2

Review the Setup

Carla selects the dataset target and task from this page automatically.

Step 3

Evaluate & Predict

Run predictions and compute metrics with zero server uploads.