Version B evidence
Experiment dashboard
Every value is read from artifacts/results at request time, never hardcoded.
How to read these metrics
- F1
- Balance of precision and recall. The headline detection metric on the test split.
- AUROC
- Probability that a random positive case scores above a random negative one.
- PR-AUC
- Precision-recall area. More informative than AUROC when classes are imbalanced.
- MCC
- Matthews correlation coefficient. A balanced single-number quality score.
- ECE
- Expected calibration error. Gap between predicted probabilities and observed frequencies.
- Brier
- Mean squared error of the predicted probabilities. Lower is better.
- Platt / isotonic
- Post-hoc calibration methods. Platt is the deployable default here.
- SHAP
- Per-feature attribution explaining how the model reached a score.
- McNemar
- Paired significance test between two classifiers on the same test cases.
- Bootstrap CI
- Confidence interval from resampling with replacement.
Corrected Baselines — grouped 5-fold CV, seeds 42/123/456 (B2)
Grouped by item_idx (no leakage); McNemar and bootstrap CIs from b2_statistical_tests.json.
| Model | Det. | Precision | Recall | F1 | AUROC | PR-AUC | MCC | ECE |
|---|---|---|---|---|---|---|---|---|
| Majority | yes | 0.0000 | 0.0000 | 0.0000 | — | — | 0.0000 | — |
| Heuristic (1 − overlap) | yes | 0.9473 | 0.9233 | 0.9352 | 0.9132 | 0.8196 | 0.8723 | 0.3541 |
| Logistic Regression | no | 0.9768 | 0.9553 | 0.9660 | 0.9932 | 0.9901 | 0.9329 | 0.0099 |
| Random Forest | no | 0.9897 | 0.9789 | 0.9842 | 0.9980 | 0.9983 | 0.9687 | 0.0143 |
| NLI-only | no | 0.6479 | 0.6967 | 0.6714 | 0.7177 | 0.7531 | 0.3189 | 0.0678 |
| TF-IDF (all) | no | 0.6397 | 0.5647 | 0.5999 | 0.6754 | 0.6785 | 0.2484 | 0.0679 |
| TF-IDF (answer) | no | 0.9550 | 0.8920 | 0.9224 | 0.9696 | 0.9665 | 0.8519 | 0.0831 |
| TF-IDF (context) | no | 0.5000 | 1.0000 | 0.6667 | — | — | 0.0000 | 0.0000 |
| XGBoost (ours) | no | 0.9932 | 0.9762 | 0.9846 | 0.9979 | 0.9983 | 0.9697 | 0.0045 |
McNemar (XGBoost vs best baseline rf_full): p = 0.0442 · best baseline F1 0.9842 · Bootstrap 95% CI F1 [0.981, 0.990]
Pipeline provenance (manifest.json)
Frozen run metadata — commit, fingerprint, hardware.
- Model version
- b6-ec-xgb-v1.0
- Git commit
- 5710f90fcdc8
- Source fingerprint
- —
- Seeds
- 42, 123, 456
- GPU
- NVIDIA GeForce RTX 3060 Laptop GPU (6 GB)
- RAM
- —
- Split
- leakage-free (10000 groups)
Leakage-removal impact (B2)
Historical leaked split vs the corrected grouped pipeline.
- Historical (leaky row-level)
- F1 0.9886 · AUROC 0.9980
- Version A (corrected grouped split)
- F1 0.9840 · AUROC 0.9981
- B2 (this work, grouped 5-fold CV)
- F1 0.9846 · AUROC 0.9979
Historical leaked numbers come from the pre-repair README table (row-level split). B2 uses the corrected grouped split AND grouped 5-fold CV for tuning; Version A used the corrected split with row-level stratified CV.
Tuning (per seed, 5-fold group CV)
- seed 42cv_auc 0.9972 · lr 0.05 · depth 5 · n 300
- seed 123cv_auc 0.9972 · lr 0.05 · depth 4 · n 200
- seed 456cv_auc 0.9972 · lr 0.1 · depth 5 · n 100
Every value above is read from artifacts/results/b2/* at request time. ECE/calibration details live in the Calibration tab; transfer evidence in the Robustness tab.