survivedOpenML #40945 ↗Titanic
Canonical 1,309-record benchmark dataset describing Titanic passenger demographics, socio-economic markers, and cabin locations to predict binary survival outcomes.
Business Objective: Titanic
01Business Context
Passenger evacuation protocols and maritime safety logistics during mass-casualty emergency operations, historically governed by triage heuristics such as 'women and children first' and stratified socio-economic access to survival resources.
02Analytical Objective
Predict the binary survival probability of individual passengers based on socio-economic tier, demographic profile, cabin positioning, and family accompaniment.
03Economic & Decision Impact
Enables historical casualty auditing, safety triage simulation, and actuarial disaster loss modeling where false positives (misclassifying deceased passengers as survivors) distort survival resource allocations and emergency egress planning.
Widely utilized as the definitive baseline for tabular binary classification, traditional gradient-boosted decision trees (XGBoost, LightGBM) and Random Forests routinely achieve 0.80–0.84 ROC-AUC on clean non-leaking feature subsets (pclass, sex, age, sibsp, parch, fare, embarked). Foundation models like TabICL and Carla demonstrate robust zero-shot in-context inference directly on raw spreadsheet rows without explicit imputation pipelines or one-hot categorical encoding.
Target & Feature Column Definitions
Exhaustive business definitions, measurement units, roles, and target variables across all 14 columns.
| Col | Variable / Feature Name | Role | Type | Unit / Scale | Description & Business Meaning |
|---|---|---|---|---|---|
| A | survivedSurvival Status | TARGET | NUM | — | Binary passenger outcome indicating whether the individual survived the disaster (0 = Did not survive / Deceased, 1 = Survived). |
| B | pclassPassenger Class | Feature | NUM | — | Ticket class representing socio-economic status (1 = 1st Class / Upper Tier, 2 = 2nd Class / Middle Tier, 3 = 3rd Class / Lower Tier). |
| C | namePassenger Name | Feature | CAT | — | Full formal name and honorific title (e.g. Mr., Mrs., Miss, Master, Dr., Rev.) of the passenger. |
| D | sexSex | Feature | CAT | — | Biological sex / gender of the passenger (female, male). |
| E | ageAge | Feature | NUM | years | Age of the passenger in years (fractional for infants under 1 year old; missing values represent unknown ages). |
| F | sibspSiblings / Spouses Aboard | Feature | NUM | count | Total count of siblings (brother, sister, stepbrother, stepsister) or spouses (husband, wife) traveling with the passenger. |
| G | parchParents / Children Aboard | Feature | NUM | count | Total count of parents (mother, father) or children (daughter, son, stepdaughter, stepson) traveling with the passenger. |
| H | ticketTicket Number | Feature | CAT | — | Alphanumeric ticket booking identifier or serial reference assigned to the passenger. |
| I | farePassenger Fare | Feature | NUM | GBP | Total passenger fare paid in pre-1970 British Pounds (£) for boarding the voyage. |
| J | cabinCabin Number | Feature | CAT | — | Assigned deck and stateroom cabin identifier (e.g. B5, C22 C26). Deck letter reflects vertical proximity to boat deck. |
| K | embarkedPort of Embarkation | Feature | CAT | — | Port where the passenger boarded the vessel (C = Cherbourg, Q = Queenstown, S = Southampton). |
| L | boatLifeboat Identifier | LEAKAGE | CAT | — | Identification number or letter of the rescue lifeboat used by surviving passengers (post-event indicator). TARGET LEAKAGE RISK(post event) The lifeboat identifier indicates which rescue boat a passenger boarded during or after the disaster. This feature is recorded post-event and provides direct leakage of survival status, as having an assigned lifeboat strongly determines that the passenger survived. • Excluded during model training & zero-shot inference to avoid artificially inflated metric scores. |
| L | homedestHome / Destination | Feature | CAT | — | Geographic origin / home city and ultimate travel destination intended by the passenger. |
| M | bodyBody Identification Number | LEAKAGE | NUM | — | Coroner recovery identification index assigned to deceased passenger bodies retrieved from the ocean. TARGET LEAKAGE RISK(post event) The body identification number is assigned by recovery crews and coroners to deceased passengers whose bodies were retrieved from the ocean after the sinking. This is a post-event feature that deterministically indicates the passenger did not survive. • Excluded during model training & zero-shot inference to avoid artificially inflated metric scores. |
Interactive Data Table
Explore rows, feature values, and target labels for Titanic.
TabICLv2 WebGPU Benchmark Results
Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.
| Evaluation Split | Train Rows | Test Rows | Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more → | Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details → | ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details → | F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more → | Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details → | Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details → |
|---|---|---|---|---|---|---|---|---|
| R1 / F1 | 872 | 437 | 1.42s | 79.86% | 0.8519 | 0.7233 | 0.7616 | 0.6886 |
| R1 / F2 | 873 | 436 | 1.41s | 81.65% | 0.8707 | 0.7403 | 0.8085 | 0.6826 |
| R1 / F3 | 873 | 436 | 1.34s | 82.57% | 0.8780 | 0.7610 | 0.7961 | 0.7289 |
| R2 / F1 | 872 | 437 | 1.35s | 79.18% | 0.8708 | 0.7036 | 0.7714 | 0.6467 |
| R2 / F2 | 873 | 436 | 1.34s | 81.19% | 0.8488 | 0.7303 | 0.8102 | 0.6647 |
| R2 / F3 | 873 | 436 | 1.33s | 85.09% | 0.8950 | 0.7855 | 0.8686 | 0.7169 |
| R3 / F1 | 872 | 437 | 1.34s | 85.13% | 0.8981 | 0.7766 | 0.9113 | 0.6766 |
| R3 / F2 | 873 | 436 | 1.34s | 81.65% | 0.8630 | 0.7403 | 0.8085 | 0.6826 |
| R3 / F3 | 873 | 436 | 1.35s | 82.57% | 0.8775 | 0.7669 | 0.7813 | 0.7530 |
| R4 / F1 | 872 | 437 | 1.35s | 83.07% | 0.8560 | 0.7566 | 0.8394 | 0.6886 |
| R4 / F2 | 873 | 436 | 1.35s | 82.34% | 0.8750 | 0.7556 | 0.8041 | 0.7126 |
| R4 / F3 | 873 | 436 | 1.34s | 81.42% | 0.8827 | 0.7344 | 0.8058 | 0.6747 |
| R5 / F1 | 872 | 437 | 1.34s | 79.86% | 0.8564 | 0.6986 | 0.8160 | 0.6108 |
| R5 / F2 | 873 | 436 | 1.37s | 85.32% | 0.8999 | 0.7987 | 0.8411 | 0.7605 |
| R5 / F3 | 873 | 436 | 1.34s | 82.57% | 0.8903 | 0.7580 | 0.8041 | 0.7169 |
| R6 / F1 | 872 | 437 | 1.35s | 81.01% | 0.8645 | 0.7492 | 0.7561 | 0.7425 |
| R6 / F2 | 873 | 436 | 1.35s | 83.26% | 0.8948 | 0.7638 | 0.8310 | 0.7066 |
| R6 / F3 | 873 | 436 | 1.38s | 80.96% | 0.8534 | 0.7167 | 0.8268 | 0.6325 |
| R7 / F1 | 872 | 437 | 1.35s | 83.07% | 0.8867 | 0.7613 | 0.8252 | 0.7066 |
| R7 / F2 | 873 | 436 | 1.34s | 81.42% | 0.8655 | 0.7344 | 0.8116 | 0.6707 |
| R7 / F3 | 873 | 436 | 1.34s | 81.19% | 0.8794 | 0.7172 | 0.8387 | 0.6265 |
| R8 / F1 | 872 | 437 | 1.35s | 78.03% | 0.8302 | 0.6842 | 0.7591 | 0.6228 |
| R8 / F2 | 873 | 436 | 1.34s | 82.34% | 0.8814 | 0.7475 | 0.8261 | 0.6826 |
| R8 / F3 | 873 | 436 | 1.33s | 85.09% | 0.9021 | 0.7962 | 0.8301 | 0.7651 |
| R9 / F1 | 872 | 437 | 1.32s | 81.92% | 0.8710 | 0.7358 | 0.8333 | 0.6587 |
| R9 / F2 | 873 | 436 | 1.34s | 79.59% | 0.8459 | 0.7192 | 0.7600 | 0.6826 |
| R9 / F3 | 873 | 436 | 1.36s | 87.16% | 0.9155 | 0.8170 | 0.8929 | 0.7530 |
| R10 / F1 | 872 | 437 | 1.33s | 81.92% | 0.8781 | 0.7443 | 0.8099 | 0.6886 |
| R10 / F2 | 873 | 436 | 1.35s | 83.49% | 0.8989 | 0.7647 | 0.8417 | 0.7006 |
| R10 / F3 | 873 | 436 | 1.34s | 83.26% | 0.8670 | 0.7509 | 0.8661 | 0.6627 |
| Mean ± Std | — | — | 1.35s | 82.24% ± 1.97% | 0.875 ± 0.019 | 0.748 ± 0.030 | 0.818 | 0.690 |
Foundation Models vs Traditional ML Baselines
Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.
| Rank | Algorithm / Model | Model Family | Runtime | Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details → | ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more → | F-Measure?Harmonic mean of precision and recall.Learn more → | Reference / Repo |
|---|---|---|---|---|---|---|---|
| #1 | EXAONE Tabular | Tabular Foundation Model | PyTorch/Python | 82.96% | 0.8854 | 0.7569 | LGAI-Research/EXAONE-Tabular ↗ |
| #2 | Google TabFM v1.0 | Tabular Foundation Model | PyTorch/Python | 82.47% | 0.8834 | 0.7509 | google-research/tabfm ↗ |
| #3 | Streaming nanotabicl | Tabular Foundation Model | PyTorch/Python | 82.38% | 0.8771 | 0.7498 | soda-inria/nanotabicl ↗ |
| #4 | Full TabICLv2 (PyTorch Reference) | Tabular Foundation Model | PyTorch/Python | 82.38% | 0.8771 | 0.7498 | soda-inria/tabicl ↗ |
| #5 | nanotabicl Vanilla | Tabular Foundation Model | PyTorch/Python | 82.35% | 0.8771 | 0.7494 | soda-inria/nanotabicl ↗ |
| #6 | Carla Engine (TabICLv2 WebGPU) | Tabular Foundation Model | Browser | 82.24% | 0.8749 | 0.7477 | 100% In-Browser |
| #7 | TabPFN v3 | Tabular Foundation Model | PyTorch/Python | 81.13% | 0.8640 | 0.7327 | PriorLabs/tabpfn ↗ |
Variable Schema & Summary Distributions
Observed numerical ranges, category cardinalities, missing rates, and sample values across 1,309 rows.
| Col | Variable Name | Type | Missing | Distinct | Summary Stats / Distribution | Sample Values |
|---|---|---|---|---|---|---|
| A | survivedTARGET | NUM | 0 | 2 | [0, 1] μ=0.4 σ=0.5 | 110 |
| B | pclass | NUM | 0 | 3 | [1, 3] μ=2.3 σ=0.8 | 111 |
| C | name | CAT | 0 | 1,307 | — | Allen, Miss. Elisabeth WaltonAllison, Master. Hudson TrevorAllison, Miss. Helen Loraine |
| D | sex | CAT | 0 | 2 | 2 cats: male, female | femalemalefemale |
| E | age | NUM | 263 (20.09%) | 98 | [0.1667, 80] μ=29.9 σ=14.4 | 290.91672 |
| F | sibsp | NUM | 0 | 7 | [0, 8] μ=0.5 σ=1.0 | 011 |
| G | parch | NUM | 0 | 8 | [0, 9] μ=0.4 σ=0.9 | 022 |
| H | ticket | CAT | 0 | 929 | — | 24160113781113781 |
| I | fare | NUM | 1 (0.08%) | 281 | [0, 512.3292] μ=33.3 σ=51.7 | 211.3375151.55151.55 |
| J | cabin | CAT | 1,014 (77.46%) | 186 | — | B5C22 C26C22 C26 |
| K | embarked | CAT | 2 (0.15%) | 3 | 3 cats: S, C, Q | SSS |
| L | boatLEAKAGE | CAT | 823 (62.87%) | 27 | 10 cats: 13, C, 15 +7 more | 2113 |
| L | homedest | CAT | 564 (43.09%) | 369 | — | St Louis, MOMontreal, PQ / Chesterville, ONMontreal, PQ / Chesterville, ON |
| M | bodyLEAKAGE | NUM | 1,188 (90.76%) | 121 | [1, 328] μ=160.8 σ=97.3 | 13522124 |
Frequently Asked Questions: Titanic
Common questions regarding the Titanic dataset, machine learning task formulations, and in-browser tabular inference.
What is the Titanic dataset used for?
The Titanic dataset is the premier tabular benchmark for binary classification, used to predict passenger survival based on socio-economic tier, demographic profiles, family accompaniment, and travel attributes.
Can I run zero-server predictions on this dataset in Google Sheets?
Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models without requiring backend configuration or feature engineering.
What machine learning models perform best on Titanic?
Standard gradient-boosted trees (LightGBM, XGBoost) and Random Forests typically reach 0.80–0.84 ROC-AUC on clean non-leaking feature sets, while tabular foundation models like TabICL deliver high-accuracy zero-shot inference without prior training.
Test TabICLv2 on Titanic Yourself
Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.
Launch Carla
Click once to open the spreadsheet and Carla side panel together.
Review the Setup
Carla selects the dataset target and task from this page automatically.
Evaluate & Predict
Run predictions and compute metrics with zero server uploads.
Provenance & Attribution
OpenML Dataset 40945