is_target_concept_metOpenML #334 ↗Monks Problems 2
The MONK's Problems 2 benchmark dataset comprises 601 records evaluated across 6 discrete categorical features to determine whether an entity satisfies an exact non-linear combinatorial condition. It serves as a foundational machine learning benchmark for evaluating algorithm capability in learning complex logical parity-like interactions.
Business Objective: Monks Problems 2
01Business Context
Conceived at the 2nd European Summer School on Machine Learning (Corsendonk Priory) to compare inductive learning paradigms, the MONK-2 task tests an algorithm's capability to learn non-linear logical interactions without manual feature engineering across discrete feature dimensions.
02Analytical Objective
Predict binary classification status ('class' = 1 vs 0) based on whether exactly two of the six discrete attributes equal 1.
03Economic & Decision Impact
Provides definitive diagnostic validation of model inductive biases; standard decision tree ensembles struggle with high-order XOR/parity-like interactions, leading to costly model selection errors and degraded real-world rule extraction.
Standard axis-aligned decision trees (such as basic CART or shallow Random Forests) struggle on MONK-2 due to the combinatorial 'exactly two' condition which demands deep, repeated splits. In contrast, neural networks and modern in-context foundation models (TabICL / Carla) effectively capture full-table relational dependencies and high-order feature intersections zero-shot directly from spreadsheet context.
Target & Feature Column Definitions
Exhaustive business definitions, measurement units, roles, and target variables across all 7 columns.
| Col | Variable / Feature Name | Role | Type | Unit / Scale | Description & Business Meaning |
|---|---|---|---|---|---|
| A | is_target_concept_metTarget Concept Match (Class) | TARGET | NUM | — | Binary target outcome indicating whether exactly two of the six attributes equal 1 (1 = True, 0 = False). |
| B | attribute_1_head_shapeAttribute 1 (Head Shape) | Feature | NUM | — | Discrete categorical feature representing head shape with 3 discrete values (1: round, 2: square, 3: octagon). |
| C | attribute_2_body_shapeAttribute 2 (Body Shape) | Feature | NUM | — | Discrete categorical feature representing body shape with 3 discrete values (1: round, 2: square, 3: octagon). |
| D | attribute_3_is_smilingAttribute 3 (Is Smiling) | Feature | NUM | — | Discrete binary feature indicating facial expression (1: smiling, 2: not smiling). |
| E | attribute_4_holding_objectAttribute 4 (Holding Object) | Feature | NUM | — | Discrete categorical feature indicating held object (1: sword, 2: balloon, 3: flag). |
| F | attribute_5_jacket_colorAttribute 5 (Jacket Color) | Feature | NUM | — | Discrete categorical feature denoting jacket color (1: red, 2: yellow, 3: green, 4: blue). |
| G | attribute_6_has_tieAttribute 6 (Has Tie) | Feature | NUM | — | Discrete binary feature denoting tie presence (1: yes, 2: no). |
Interactive Data Table
Explore rows, feature values, and target labels for Monks Problems 2.
TabICLv2 WebGPU Benchmark Results
Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.
| Evaluation Split | Train Rows | Test Rows | Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more → | Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details → | ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details → | F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more → | Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details → | Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details → |
|---|---|---|---|---|---|---|---|---|
| R1 / F1 | 400 | 201 | 1.06s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R1 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R1 / F3 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R2 / F1 | 400 | 201 | 1.03s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R2 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R2 / F3 | 401 | 200 | 1.06s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R3 / F1 | 400 | 201 | 1.03s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R3 / F2 | 401 | 200 | 1.06s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R3 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R4 / F1 | 400 | 201 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R4 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R4 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R5 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R5 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R5 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R6 / F1 | 400 | 201 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R6 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R6 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R7 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R7 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R7 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R8 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R8 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R8 / F3 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R9 / F1 | 400 | 201 | 1.03s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R9 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R9 / F3 | 401 | 200 | 1.07s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R10 / F1 | 400 | 201 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R10 / F2 | 401 | 200 | 1.06s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R10 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R11 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R11 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R11 / F3 | 401 | 200 | 1.03s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R12 / F1 | 400 | 201 | 1.03s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R12 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R12 / F3 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R13 / F1 | 400 | 201 | 1.20s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R13 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R13 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R14 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R14 / F2 | 401 | 200 | 1.08s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R14 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R15 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R15 / F2 | 401 | 200 | 1.07s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R15 / F3 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R16 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R16 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R16 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R17 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R17 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R17 / F3 | 401 | 200 | 1.09s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R18 / F1 | 400 | 201 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R18 / F2 | 401 | 200 | 1.04s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R18 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R19 / F1 | 400 | 201 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R19 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R19 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R20 / F1 | 400 | 201 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R20 / F2 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| R20 / F3 | 401 | 200 | 1.05s | 100.00% | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| Mean ± Std | — | — | 1.05s | 100.00% ± 0.00% | 1.000 ± 0.000 | 1.000 ± 0.000 | 1.000 | 1.000 |
Foundation Models vs Traditional ML Baselines
Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.
| Rank | Algorithm / Model | Model Family | Runtime | Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details → | ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more → | F-Measure?Harmonic mean of precision and recall.Learn more → | Reference / Repo |
|---|---|---|---|---|---|---|---|
| #1 | Carla Engine (TabICLv2 WebGPU) | Tabular Foundation Model | Browser | 100.00% | 1.0000 | 1.0000 | 100% In-Browser |
| #2 | Streaming nanotabicl | Tabular Foundation Model | PyTorch/Python | 100.00% | 1.0000 | 1.0000 | soda-inria/nanotabicl ↗ |
| #3 | nanotabicl Vanilla | Tabular Foundation Model | PyTorch/Python | 100.00% | 1.0000 | 1.0000 | soda-inria/nanotabicl ↗ |
| #4 | Full TabICLv2 (PyTorch Reference) | Tabular Foundation Model | PyTorch/Python | 100.00% | 1.0000 | 1.0000 | soda-inria/tabicl ↗ |
| #5 | TabPFN v3 | Tabular Foundation Model | PyTorch/Python | 100.00% | 1.0000 | 1.0000 | PriorLabs/tabpfn ↗ |
| #6 | Google TabFM v1.0 | Tabular Foundation Model | PyTorch/Python | 100.00% | 1.0000 | 1.0000 | google-research/tabfm ↗ |
| #7 | FilteredClassifier MultiSearch SMO RBFKernel | Support Vector Machine (SVM) | Python | 100.00% | 1.0000 | 1.0000 | OpenML #593455 ↗ |
| #8 | FilteredClassifier MultilayerPerceptron | Neural Network (MLP) | Python | 100.00% | 1.0000 | 1.0000 | OpenML #1762905 ↗ |
| #9 | MultilayerPerceptron | Neural Network (MLP) | Python | 100.00% | 1.0000 | 1.0000 | OpenML #1810966 ↗ |
| #10 | Bagging MultilayerPerceptron | Neural Network (MLP) | Python | 100.00% | 1.0000 | 1.0000 | OpenML #1814953 ↗ |
| #11 | FilteredClassifier MultiSearch MultilayerPerceptron | Neural Network (MLP) | Python | 100.00% | 1.0000 | 1.0000 | OpenML #1846081 ↗ |
| #12 | EXAONE Tabular | Tabular Foundation Model | PyTorch/Python | 99.48% | 1.0000 | 0.9926 | LGAI-Research/EXAONE-Tabular ↗ |
| #13 | BFTree | Decision Tree | Python | 97.50% | 0.9958 | 0.9752 | OpenML #1666257 ↗ |
| #14 | SimpleCart | Machine Learning Model | Python | 96.84% | 0.9986 | 0.9686 | OpenML #1665974 ↗ |
| #15 | ClassificationViaRegression M5P | Decision Tree | Python | 95.84% | 0.9873 | 0.9586 | OpenML #1739056 ↗ |
| #16 | AdaBoostM1 LMT | AdaBoost | Python | 85.36% | 0.9185 | 0.8539 | OpenML #1673309 ↗ |
| #17 | IB1 | Machine Learning Model | Python | 83.36% | 0.7968 | 0.8298 | OpenML #1572885 ↗ |
Variable Schema & Summary Distributions
Observed numerical ranges, category cardinalities, missing rates, and sample values across 601 rows.
| Col | Variable Name | Type | Missing | Distinct | Summary Stats / Distribution | Sample Values |
|---|---|---|---|---|---|---|
| A | is_target_concept_metTARGET | NUM | 0 | 2 | [0, 1] μ=0.3 σ=0.5 | 000 |
| B | attribute_1_head_shape | NUM | 0 | 3 | [1, 3] μ=2.0 σ=0.8 | 111 |
| C | attribute_2_body_shape | NUM | 0 | 3 | [1, 3] μ=2.0 σ=0.8 | 111 |
| D | attribute_3_is_smiling | NUM | 0 | 2 | [1, 2] μ=1.5 σ=0.5 | 111 |
| E | attribute_4_holding_object | NUM | 0 | 3 | [1, 3] μ=2.0 σ=0.8 | 112 |
| F | attribute_5_jacket_color | NUM | 0 | 4 | [1, 4] μ=2.5 σ=1.1 | 241 |
| G | attribute_6_has_tie | NUM | 0 | 2 | [1, 2] μ=1.5 σ=0.5 | 211 |
Frequently Asked Questions: Monks Problems 2
Common questions regarding the Monks Problems 2 dataset, machine learning task formulations, and in-browser tabular inference.
What is the Monks Problems 2 dataset used for?
MONK-2 is a classic machine learning benchmark used to test an algorithm's capability to learn non-linear combinatorial concepts (specifically: exactly two of the six attributes equal 1) across discrete features.
Can I run zero-server predictions on this dataset in Google Sheets?
Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models without dedicated backend infrastructure.
What machine learning models perform best on Monks Problems 2?
Models that effectively represent combinatorial feature interactions—such as multi-layer perceptrons, symbolic rule learners, and in-context tabular foundation models (TabICL / Carla)—substantially outperform unaugmented shallow decision trees on MONK-2.
Test TabICLv2 on Monks Problems 2 Yourself
Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.
Launch Carla
Click once to open the spreadsheet and Carla side panel together.
Review the Setup
Carla selects the dataset target and task from this page automatically.
Evaluate & Predict
Run predictions and compute metrics with zero server uploads.
Provenance & Attribution
https://archive.ics.uci.edu/ml/citation_policy.html