Evaluation dashboard

The honest result is the useful result.

The champion ensemble performed well on low-risk cases but struggled with the transitional mid-risk class. These results come from a profile-isolated holdout.

Locked holdout · 93 rows
Accuracy 74.2% 69 of 93 predictions correct
Macro F1 0.600 Equal weight across risk classes
Balanced accuracy 0.596 Mean class recall
Quadratic kappa 0.639 Ordinal agreement beyond chance

Holdout accuracy by model

The diverse ensemble was selected using development data before holdout evaluation.

Champion confusion matrix

Rows are observed classes; columns are predicted classes.

Observed risk
LowMidHigh Low5496.4%00.0%23.6% Mid1473.7%315.8%210.5% High422.2%211.1%1266.7%
Predicted risk
Low recall96.4%
Mid recall15.8%
High recall66.7%

Two benchmarks

Why the lower score carries more evidence.

The project retained a legacy benchmark for sensitivity analysis, then separated it from the leakage-resistant research result.

Profile-isolated evaluation 74.2%

Exact duplicates removed; identical predictor profiles kept in one split or fold. This is the reported research estimate.

versus
Legacy row split 90.1%

79.8% of test rows had a predictor profile also present in training. Useful as an audit, not as clinical or external validation.

Model-agnostic SHAP importance

Mean absolute change in predicted probability, aggregated over the explanation sample.

Systolic BP0.083
Blood glucose0.078
Body temp0.052
Diastolic BP0.020
Heart rate0.014
Age0.014

Permutation importance

Average decrease in holdout accuracy after shuffling each raw input.

Blood glucose0.137
Body temp0.070
Systolic BP0.062
Diastolic BP0.034
Age0.015
Heart rate0.006

Complete model comparison

No candidate dominated every criterion. Random Forest led macro F1; Extra Trees led balanced accuracy and high-risk recall.

ModelAccuracyBalanced accuracyMacro F1QWKMid recallHigh recall
Diverse ensemble74.19%0.59630.59990.63870.15790.6667
Extra Trees73.12%0.61450.60140.61570.21050.7222
Random Forest73.12%0.61350.61290.62310.26320.6667
SVC68.82%0.56550.55540.55310.21050.6111
LightGBM70.97%0.57750.58840.60420.21050.6111
XGBoost70.97%0.56590.57210.57980.15790.6111
CatBoost73.12%0.57970.56100.60670.05260.7222

Interpretation boundary

These results establish an internal, reproducible proof of concept. They do not establish safety, clinical utility, external validity, or performance for child and neonatal outcomes.