Skip to main content
CARLA HQ
Fitness, Sports & Recreation Commerce & Industry use casesNACE R93.13Binary ClassificationTarget: attended

Fitness Club

The Fitness Club dataset contains 1,500 booking records from the Canadian gym chain GoalZone to predict class attendance and minimize no-show rates for high-demand workout sessions.

✨ Try in CarlaInstall Chrome Extension ↗100% In-Browser WebGPU • Zero Cloud Upload
1,500Records (Rows)
6Predictive Features
3 / 4Numeric / Categorical
0.2%Missing Value Ratio
0.17951 - AUC Error
1.25sWebGPU Mean Split
BUSINESS CONTEXT & OBJECTIVEFitness & Wellness Operations (R93.13)

Business Objective: Fitness Club

01Business Context

GoalZone operates boutique fitness classes capped at 15 or 25 participants. High booking demand frequently leads to fully booked sessions that paradoxically suffer from elevated no-show rates, leaving equipment underutilized and turning away motivated walk-in members.

02Analytical Objective

Predict whether a booked member will attend their scheduled fitness session based on membership tenure, physical profile, booking lead time, and session metadata.

03Economic & Decision Impact

Accurate attendance forecasting allows dynamic overbooking and waitlist backfilling, maximizing studio capacity utilization while preventing revenue loss. False negatives (predicting no-show for an attendee) risk customer dissatisfaction from overbooking, whereas false positives (predicting attendance for a no-show) leave empty spots on studio floors.

ML Benchmark Narrative

Standard gradient-boosted trees (XGBoost, LightGBM) and Random Forest achieve solid baseline ROC-AUC scores (~0.75-0.79) by heavily leveraging membership tenure and booking lead time. In-context tabular foundation models such as TabICL and Carla match or exceed these tree-based baselines zero-shot, eliminating feature encoding overhead and hyperparameter tuning directly within business spreadsheets.

Source Origin:DataCamp / Kaggle
Creator:Ddosad (2023)
License:Open Data Commons / CC BY 4.0
FEATURE SPECIFICATION & DATA DICTIONARY

Target & Feature Column Definitions

Exhaustive business definitions, measurement units, roles, and target variables across all 7 columns.

ColVariable / Feature NameRoleTypeUnit / ScaleDescription & Business Meaning
GattendedAttended ClassTARGETCAT
Binary ground-truth outcome indicating whether the booked member attended the scheduled class session ('Yes' or 'No').
Amonths_as_memberMembership Tenure (Months)FeatureNUMmonths
Total number of consecutive months the individual has held an active gym membership at GoalZone.
BweightMember WeightFeatureNUMkg
Current recorded body weight of the member in kilograms.
Cdays_beforeBooking Lead TimeFeatureNUMdays
Number of days in advance the member reserved their spot in the class.
Dday_of_weekClass Day of WeekFeatureCAT
The day of the week on which the fitness class is scheduled (e.g. Mon, Tue, Wed, Thu, Fri, Sat, Sun).
EtimeClass Time PeriodFeatureCAT
Scheduled time block of the class, indicating morning (AM) or afternoon/evening (PM) sessions.
FcategoryFitness Class CategoryFeatureCAT
Specific workout discipline or class type (e.g. HIIT, Strength, Cycling, Yoga, Aqua).
DATASET PREVIEW

Interactive Data Table

Explore rows, feature values, and target labels for Fitness Club.

Loading dataset...
LOCAL EXECUTION TELEMETRY

TabICLv2 WebGPU Benchmark Results

Complete performance metrics on persisted IID or non-IID evaluation splits, executed 100% locally in the browser sandbox.

Mean Split Latency?Mean wall-clock execution time per persisted evaluation split running 100% locally via WebGPU inside Chrome.Explore Latency Guide →
1.25s30 splits • 8 ensembles/split
Total Evaluation Runtime?Total wall-clock duration across all selected splits, including in-context encoding, WebGPU shader execution, and result aggregation.Explore Split Protocol →
37.57sIncludes warmup & sync
Evaluation Protocol?Precomputed row indices ensure the browser and Python runners evaluate identical out-of-sample observations without leakage.Explore Split Protocol →
10×3 SplitsRepeated IID • 30 total
TabICLv2 Score?Primary task performance score (1 - AUC Error) achieved by TabICLv2 across out-of-sample evaluation splits.Explore Metric Formula →
0.17951 - AUC Error
Evaluation SplitTrain RowsTest Rows
Split Duration?Wall-clock inference time taken for this evaluation split running 100% locally via WebGPU.Learn more →
Accuracy?Proportion of correct test predictions across all classes.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability across all thresholds.Google ML Guide & Details →
F1 Score?Harmonic mean of Precision and Recall, robust against class imbalance.Learn more →
Precision?Proportion of predicted positives that were actually positive.Google ML Guide & Details →
Recall?Proportion of actual positive cases successfully captured.Google ML Guide & Details →
R1 / F11,0005001.31s78.60%0.81790.59320.70270.5132
R1 / F21,0005001.25s78.60%0.82070.60810.68030.5497
R1 / F31,0005001.27s77.40%0.82160.52720.71590.4172
R2 / F11,0005001.25s76.60%0.80390.56180.65220.4934
R2 / F21,0005001.24s79.60%0.83020.60160.73330.5099
R2 / F31,0005001.24s77.40%0.82550.54620.69390.4503
R3 / F11,0005001.25s77.80%0.83090.60500.65890.5592
R3 / F21,0005001.24s80.40%0.82510.60480.77320.4967
R3 / F31,0005001.25s75.60%0.79850.51590.64360.4305
R4 / F11,0005001.24s77.00%0.79030.58480.64800.5329
R4 / F21,0005001.25s79.20%0.84550.59060.72820.4967
R4 / F31,0005001.26s78.00%0.82530.56350.70300.4702
R5 / F11,0005001.25s79.80%0.81970.58780.77420.4737
R5 / F21,0005001.23s77.00%0.82140.59070.63850.5497
R5 / F31,0005001.24s76.80%0.82040.55040.66360.4702
R6 / F11,0005001.24s79.40%0.83830.60840.72070.5263
R6 / F21,0005001.24s78.00%0.81240.57030.69520.4834
R6 / F31,0005001.23s77.40%0.81270.56700.67270.4901
R7 / F11,0005001.25s73.60%0.78890.51470.58330.4605
R7 / F21,0005001.25s80.40%0.82660.59500.79120.4768
R7 / F31,0005001.24s79.60%0.84150.61650.71300.5430
R8 / F11,0005001.23s78.40%0.81920.57810.71150.4868
R8 / F21,0005001.24s78.00%0.83260.56350.70300.4702
R8 / F31,0005001.24s78.40%0.81440.60870.67200.5563
R9 / F11,0005001.24s78.40%0.84200.58140.70750.4934
R9 / F21,0005001.25s77.20%0.81190.56150.66970.4834
R9 / F31,0005001.30s76.80%0.80470.55380.66060.4768
R10 / F11,0005001.34s81.20%0.86380.62990.78430.5263
R10 / F21,0005001.25s75.20%0.79950.53730.61540.4768
R10 / F31,0005001.25s77.80%0.79550.56470.69230.4768
Mean ± Std1.25s77.99% ± 1.59%0.820 ± 0.0170.576 ± 0.0290.6930.495
UNIFIED BENCHMARK LEADERBOARD

Foundation Models vs Traditional ML Baselines

Side-by-side evaluation of our in-browser Carla engine (WebGPU), open-source tabular foundation models, and OpenML baselines.

RankAlgorithm / ModelModel FamilyRuntime
Predictive Accuracy?Proportion of correct test predictions across cross-validation splits.Google ML Guide & Details →
ROC-AUC?Area under the ROC curve evaluating ranking capability.Learn more →
F-Measure?Harmonic mean of precision and recall.Learn more →
Reference / Repo
#1nanotabicl VanillaTabular Foundation ModelPyTorch/Python78.08%0.82050.5794soda-inria/nanotabicl ↗
#2Streaming nanotabiclTabular Foundation ModelPyTorch/Python78.07%0.82050.5792soda-inria/nanotabicl ↗
#3Full TabICLv2 (PyTorch Reference)Tabular Foundation ModelPyTorch/Python78.07%0.82050.5792soda-inria/tabicl ↗
#4TabPFN v3Tabular Foundation ModelPyTorch/Python78.01%0.82080.5797PriorLabs/tabpfn ↗
#5Carla Engine (TabICLv2 WebGPU)Tabular Foundation ModelBrowser77.99%0.82000.5761100% In-Browser
#6Google TabFM v1.0Tabular Foundation ModelPyTorch/Python77.97%0.82050.5778google-research/tabfm ↗
#7EXAONE TabularTabular Foundation ModelPyTorch/Python77.94%0.82030.5798LGAI-Research/EXAONE-Tabular ↗
STATISTICAL DISTRIBUTIONS

Variable Schema & Summary Distributions

Observed numerical ranges, category cardinalities, missing rates, and sample values across 1,500 rows.

ColVariable NameTypeMissingDistinctSummary Stats / DistributionSample Values
GattendedTARGETCAT02YesYesNo
Amonths_as_memberNUM072[1, 148] μ=15.6 σ=12.9151813
BweightNUM20 (1.33%)1,241[55.41, 170.52] μ=82.6 σ=12.865.4777.8567.26
Cdays_beforeNUM019[1, 29] μ=8.3 σ=4.16810
Dday_of_weekCAT07WedThuFri
EtimeCAT02AMAMAM
FcategoryCAT06HIITStrengthCycling
FREQUENTLY ASKED QUESTIONS

Frequently Asked Questions: Fitness Club

Common questions regarding the Fitness Club dataset, machine learning task formulations, and in-browser tabular inference.

What is the Fitness Club dataset used for?

The Fitness Club dataset is used for binary classification tasks to predict whether a gym member will attend a booked class session based on membership tenure, weight, booking lead time, and session schedule details.

Can I run zero-server predictions on this dataset in Google Sheets?

Yes, Carla HQ enables zero-server in-context tabular prediction directly in spreadsheets using TabICL foundation models without needing Python pipelines or external hosting.

What machine learning models perform best on the Fitness Club dataset?

Gradient boosting algorithms like XGBoost and LightGBM provide strong baseline ROC-AUC performance (~0.77), while tabular foundation models like TabICL and Carla deliver comparable accuracy zero-shot without manual hyperparameter tuning.

LIVE EVALUATION IN GOOGLE SHEETS

Test TabICLv2 on Fitness Club Yourself

Open the pre-loaded Google Sheet and let Carla configure the target and task for local, zero-cloud tabular machine learning.

Step 1

Launch Carla

Click once to open the spreadsheet and Carla side panel together.

Step 2

Review the Setup

Carla selects the dataset target and task from this page automatically.

Step 3

Evaluate & Predict

Run predictions and compute metrics with zero server uploads.

Provenance & Attribution

@misc{ddosad2023fitness, author = {Ddosad}, title = {Fitness Club Dataset for ML Classification}, year = {2023}, howpublished = {\url{https://www.kaggle.com/datasets/ddosad/datacamps-data-science-associate-certification}}, note = {Kaggle dataset} }

License: Public DomainData Source: KaggleOriginal Source: https://www.kaggle.com/datasets/ddosad/datacamps-data-science-associate-certification