Evaluation dashboard
The honest result is the useful result.
The champion ensemble performed well on low-risk cases but struggled with the transitional mid-risk class. These results come from a profile-isolated holdout.
Holdout accuracy by model
The diverse ensemble was selected using development data before holdout evaluation.
Champion confusion matrix
Rows are observed classes; columns are predicted classes.
Two benchmarks
Why the lower score carries more evidence.
The project retained a legacy benchmark for sensitivity analysis, then separated it from the leakage-resistant research result.
Exact duplicates removed; identical predictor profiles kept in one split or fold. This is the reported research estimate.
79.8% of test rows had a predictor profile also present in training. Useful as an audit, not as clinical or external validation.
Model-agnostic SHAP importance
Mean absolute change in predicted probability, aggregated over the explanation sample.
Permutation importance
Average decrease in holdout accuracy after shuffling each raw input.
Complete model comparison
No candidate dominated every criterion. Random Forest led macro F1; Extra Trees led balanced accuracy and high-risk recall.
| Model | Accuracy | Balanced accuracy | Macro F1 | QWK | Mid recall | High recall |
|---|---|---|---|---|---|---|
| Diverse ensemble | 74.19% | 0.5963 | 0.5999 | 0.6387 | 0.1579 | 0.6667 |
| Extra Trees | 73.12% | 0.6145 | 0.6014 | 0.6157 | 0.2105 | 0.7222 |
| Random Forest | 73.12% | 0.6135 | 0.6129 | 0.6231 | 0.2632 | 0.6667 |
| SVC | 68.82% | 0.5655 | 0.5554 | 0.5531 | 0.2105 | 0.6111 |
| LightGBM | 70.97% | 0.5775 | 0.5884 | 0.6042 | 0.2105 | 0.6111 |
| XGBoost | 70.97% | 0.5659 | 0.5721 | 0.5798 | 0.1579 | 0.6111 |
| CatBoost | 73.12% | 0.5797 | 0.5610 | 0.6067 | 0.0526 | 0.7222 |
Interpretation boundary
These results establish an internal, reproducible proof of concept. They do not establish safety, clinical utility, external validity, or performance for child and neonatal outcomes.