Skip to main content
CARLA HQ
Data Science & Artificial Intelligence Operations & Workforce Analytics use casesNACE M72Binary ClassificationTarget: is_target_concept_metOpenML #334 ↗

Monks Problems 2

The MONK's Problems 2 benchmark dataset comprises 601 records evaluated across 6 discrete categorical features to determine whether an entity satisfies an exact non-linear combinatorial condition. It serves as a foundational machine learning benchmark for evaluating algorithm capability in learning complex logical parity-like interactions.

✨ Try in CarlaInstall Chrome Extension ↗100% In-Browser WebGPU • Zero Cloud Upload
601Records (Rows)
6Predictive Features
7 / 0Numeric / Categorical
0.0%Missing Value Ratio
0.00001 - AUC Error
1.05sWebGPU Mean Split
BUSINESS CONTEXT & OBJECTIVEArtificial Intelligence & Machine Learning Research (M72)

Business Objective: Monks Problems 2

01Business Context

Conceived at the 2nd European Summer School on Machine Learning (Corsendonk Priory) to compare inductive learning paradigms, the MONK-2 task tests an algorithm's capability to learn non-linear logical interactions without manual feature engineering across discrete feature dimensions.

02Analytical Objective

Predict binary classification status ('class' = 1 vs 0) based on whether exactly two of the six discrete attributes equal 1.

03Economic & Decision Impact

Provides definitive diagnostic validation of model inductive biases; standard decision tree ensembles struggle with high-order XOR/parity-like interactions, leading to costly model selection errors and degraded real-world rule extraction.

ML Benchmark Narrative

Standard axis-aligned decision trees (such as basic CART or shallow Random Forests) struggle on MONK-2 due to the combinatorial 'exactly two' condition which demands deep, repeated splits. In contrast, neural networks and modern in-context foundation models (TabICL / Carla) effectively capture full-table relational dependencies and high-order feature intersections zero-shot directly from spreadsheet context.

Source Origin:Carnegie Mellon University & UCI Machine Learning Repository
Creator:Sebastian Thrun et al. (1991)
License:Public Domain / Open Data
FEATURE SPECIFICATION & DATA DICTIONARY

Target & Feature Column Definitions

Exhaustive business definitions, measurement units, roles, and target variables across all 7 columns.

ColVariable / Feature NameRoleTypeUnit / ScaleDescription & Business Meaning
Ais_target_concept_metTarget Concept Match (Class)TARGETNUM
Binary target outcome indicating whether exactly two of the six attributes equal 1 (1 = True, 0 = False).
Battribute_1_head_shapeAttribute 1 (Head Shape)FeatureNUM
Discrete categorical feature representing head shape with 3 discrete values (1: round, 2: square, 3: octagon).
Cattribute_2_body_shapeAttribute 2 (Body Shape)FeatureNUM
Discrete categorical feature representing body shape with 3 discrete values (1: round, 2: square, 3: octagon).
Dattribute_3_is_smilingAttribute 3 (Is Smiling)FeatureNUM
Discrete binary feature indicating facial expression (1: smiling, 2: not smiling).
Eattribute_4_holding_objectAttribute 4 (Holding Object)FeatureNUM
Discrete categorical feature indicating held object (1: sword, 2: balloon, 3: flag).
Fattribute_5_jacket_colorAttribute 5 (Jacket Color)FeatureNUM
Discrete categorical feature denoting jacket color (1: red, 2: yellow, 3: green, 4: blue).
Gattribute_6_has_tieAttribute 6 (Has Tie)FeatureNUM
Discrete binary feature denoting tie presence (1: yes, 2: no).
DATASET PREVIEW

Interactive Data Table

Explore rows, feature values, and target labels for Monks Problems 2.

Loading dataset...
LOCAL EXECUTION TELEMETRY

TabICLv2 WebGPU Benchmark Results

Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.

Mean Split Latency?Mean wall-clock execution time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Explore Latency Guide →
1.05s60 splits • 8 ensembles/split
Total Evaluation Runtime?Total wall-clock duration across all selected splits, including in-context encoding, WebGPU shader execution, and result aggregation.Explore Split Protocol →
62.97sIncludes warmup & sync
Evaluation Protocol?Precomputed row indices ensure the browser and Python runners evaluate identical out-of-sample observations without leakage.Explore Split Protocol →
20×3 SplitsRepeated IID • 60 total
TabICLv2 Score?Primary task performance score (1 - AUC Error) achieved by TabICLv2 across out-of-sample evaluation splits.Explore Metric Formula →
0.00001 - AUC Error
Evaluation SplitTrain RowsTest Rows
Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more →
Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details →
F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more →
Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details →
Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details →
R1 / F14002011.06s100.00%1.00001.00001.00001.0000
R1 / F24012001.04s100.00%1.00001.00001.00001.0000
R1 / F34012001.04s100.00%1.00001.00001.00001.0000
R2 / F14002011.03s100.00%1.00001.00001.00001.0000
R2 / F24012001.04s100.00%1.00001.00001.00001.0000
R2 / F34012001.06s100.00%1.00001.00001.00001.0000
R3 / F14002011.03s100.00%1.00001.00001.00001.0000
R3 / F24012001.06s100.00%1.00001.00001.00001.0000
R3 / F34012001.05s100.00%1.00001.00001.00001.0000
R4 / F14002011.05s100.00%1.00001.00001.00001.0000
R4 / F24012001.05s100.00%1.00001.00001.00001.0000
R4 / F34012001.05s100.00%1.00001.00001.00001.0000
R5 / F14002011.04s100.00%1.00001.00001.00001.0000
R5 / F24012001.05s100.00%1.00001.00001.00001.0000
R5 / F34012001.05s100.00%1.00001.00001.00001.0000
R6 / F14002011.05s100.00%1.00001.00001.00001.0000
R6 / F24012001.04s100.00%1.00001.00001.00001.0000
R6 / F34012001.05s100.00%1.00001.00001.00001.0000
R7 / F14002011.04s100.00%1.00001.00001.00001.0000
R7 / F24012001.04s100.00%1.00001.00001.00001.0000
R7 / F34012001.05s100.00%1.00001.00001.00001.0000
R8 / F14002011.04s100.00%1.00001.00001.00001.0000
R8 / F24012001.05s100.00%1.00001.00001.00001.0000
R8 / F34012001.04s100.00%1.00001.00001.00001.0000
R9 / F14002011.03s100.00%1.00001.00001.00001.0000
R9 / F24012001.05s100.00%1.00001.00001.00001.0000
R9 / F34012001.07s100.00%1.00001.00001.00001.0000
R10 / F14002011.05s100.00%1.00001.00001.00001.0000
R10 / F24012001.06s100.00%1.00001.00001.00001.0000
R10 / F34012001.05s100.00%1.00001.00001.00001.0000
R11 / F14002011.04s100.00%1.00001.00001.00001.0000
R11 / F24012001.05s100.00%1.00001.00001.00001.0000
R11 / F34012001.03s100.00%1.00001.00001.00001.0000
R12 / F14002011.03s100.00%1.00001.00001.00001.0000
R12 / F24012001.04s100.00%1.00001.00001.00001.0000
R12 / F34012001.04s100.00%1.00001.00001.00001.0000
R13 / F14002011.20s100.00%1.00001.00001.00001.0000
R13 / F24012001.05s100.00%1.00001.00001.00001.0000
R13 / F34012001.05s100.00%1.00001.00001.00001.0000
R14 / F14002011.04s100.00%1.00001.00001.00001.0000
R14 / F24012001.08s100.00%1.00001.00001.00001.0000
R14 / F34012001.05s100.00%1.00001.00001.00001.0000
R15 / F14002011.04s100.00%1.00001.00001.00001.0000
R15 / F24012001.07s100.00%1.00001.00001.00001.0000
R15 / F34012001.04s100.00%1.00001.00001.00001.0000
R16 / F14002011.04s100.00%1.00001.00001.00001.0000
R16 / F24012001.04s100.00%1.00001.00001.00001.0000
R16 / F34012001.05s100.00%1.00001.00001.00001.0000
R17 / F14002011.04s100.00%1.00001.00001.00001.0000
R17 / F24012001.04s100.00%1.00001.00001.00001.0000
R17 / F34012001.09s100.00%1.00001.00001.00001.0000
R18 / F14002011.04s100.00%1.00001.00001.00001.0000
R18 / F24012001.04s100.00%1.00001.00001.00001.0000
R18 / F34012001.05s100.00%1.00001.00001.00001.0000
R19 / F14002011.05s100.00%1.00001.00001.00001.0000
R19 / F24012001.05s100.00%1.00001.00001.00001.0000
R19 / F34012001.05s100.00%1.00001.00001.00001.0000
R20 / F14002011.05s100.00%1.00001.00001.00001.0000
R20 / F24012001.05s100.00%1.00001.00001.00001.0000
R20 / F34012001.05s100.00%1.00001.00001.00001.0000
Mean ± Std1.05s100.00% ± 0.00%1.000 ± 0.0001.000 ± 0.0001.0001.000
UNIFIED BENCHMARK LEADERBOARD

Foundation Models vs Traditional ML Baselines

Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.

RankAlgorithm / ModelModel FamilyRuntime
Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more →
F-Measure?Harmonic mean of precision and recall.Learn more →
Reference / Repo
#1Carla Engine (TabICLv2 WebGPU)Tabular Foundation ModelBrowser100.00%1.00001.0000100% In-Browser
#2Streaming nanotabiclTabular Foundation ModelPyTorch/Python100.00%1.00001.0000soda-inria/nanotabicl ↗
#3nanotabicl VanillaTabular Foundation ModelPyTorch/Python100.00%1.00001.0000soda-inria/nanotabicl ↗
#4Full TabICLv2 (PyTorch Reference)Tabular Foundation ModelPyTorch/Python100.00%1.00001.0000soda-inria/tabicl ↗
#5TabPFN v3Tabular Foundation ModelPyTorch/Python100.00%1.00001.0000PriorLabs/tabpfn ↗
#6Google TabFM v1.0Tabular Foundation ModelPyTorch/Python100.00%1.00001.0000google-research/tabfm ↗
#7FilteredClassifier MultiSearch SMO RBFKernelSupport Vector Machine (SVM)Python100.00%1.00001.0000OpenML #593455 ↗
#8FilteredClassifier MultilayerPerceptronNeural Network (MLP)Python100.00%1.00001.0000OpenML #1762905 ↗
#9MultilayerPerceptronNeural Network (MLP)Python100.00%1.00001.0000OpenML #1810966 ↗
#10Bagging MultilayerPerceptronNeural Network (MLP)Python100.00%1.00001.0000OpenML #1814953 ↗
#11FilteredClassifier MultiSearch MultilayerPerceptronNeural Network (MLP)Python100.00%1.00001.0000OpenML #1846081 ↗
#12EXAONE TabularTabular Foundation ModelPyTorch/Python99.48%1.00000.9926LGAI-Research/EXAONE-Tabular ↗
#13BFTreeDecision TreePython97.50%0.99580.9752OpenML #1666257 ↗
#14SimpleCartMachine Learning ModelPython96.84%0.99860.9686OpenML #1665974 ↗
#15ClassificationViaRegression M5PDecision TreePython95.84%0.98730.9586OpenML #1739056 ↗
#16AdaBoostM1 LMTAdaBoostPython85.36%0.91850.8539OpenML #1673309 ↗
#17IB1Machine Learning ModelPython83.36%0.79680.8298OpenML #1572885 ↗
STATISTICAL DISTRIBUTIONS

Variable Schema & Summary Distributions

Observed numerical ranges, category cardinalities, missing rates, and sample values across 601 rows.

ColVariable NameTypeMissingDistinctSummary Stats / DistributionSample Values
Ais_target_concept_metTARGETNUM02[0, 1] μ=0.3 σ=0.5000
Battribute_1_head_shapeNUM03[1, 3] μ=2.0 σ=0.8111
Cattribute_2_body_shapeNUM03[1, 3] μ=2.0 σ=0.8111
Dattribute_3_is_smilingNUM02[1, 2] μ=1.5 σ=0.5111
Eattribute_4_holding_objectNUM03[1, 3] μ=2.0 σ=0.8112
Fattribute_5_jacket_colorNUM04[1, 4] μ=2.5 σ=1.1241
Gattribute_6_has_tieNUM02[1, 2] μ=1.5 σ=0.5211
FREQUENTLY ASKED QUESTIONS

Frequently Asked Questions: Monks Problems 2

Common questions regarding the Monks Problems 2 dataset, machine learning task formulations, and in-browser tabular inference.

What is the Monks Problems 2 dataset used for?

MONK-2 is a classic machine learning benchmark used to test an algorithm's capability to learn non-linear combinatorial concepts (specifically: exactly two of the six attributes equal 1) across discrete features.

Can I run zero-server predictions on this dataset in Google Sheets?

Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models without dedicated backend infrastructure.

What machine learning models perform best on Monks Problems 2?

Models that effectively represent combinatorial feature interactions—such as multi-layer perceptrons, symbolic rule learners, and in-context tabular foundation models (TabICL / Carla)—substantially outperform unaugmented shallow decision trees on MONK-2.

LIVE EVALUATION IN GOOGLE SHEETS

Test TabICLv2 on Monks Problems 2 Yourself

Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.

Step 1

Launch Carla

Click once to open the spreadsheet and Carla side panel together.

Step 2

Review the Setup

Carla selects the dataset target and task from this page automatically.

Step 3

Evaluate & Predict

Run predictions and compute metrics with zero server uploads.

Provenance & Attribution

https://archive.ics.uci.edu/ml/citation_policy.html

License: PublicData Source: OpenMLOpenML Page: https://www.openml.org/d/334