Version B evidence
Experiment dashboard
Every value is read from artifacts/results at request time, never hardcoded.
How to read these metrics
- F1
- Balance of precision and recall. The headline detection metric on the test split.
- AUROC
- Probability that a random positive case scores above a random negative one.
- PR-AUC
- Precision-recall area. More informative than AUROC when classes are imbalanced.
- MCC
- Matthews correlation coefficient. A balanced single-number quality score.
- ECE
- Expected calibration error. Gap between predicted probabilities and observed frequencies.
- Brier
- Mean squared error of the predicted probabilities. Lower is better.
- Platt / isotonic
- Post-hoc calibration methods. Platt is the deployable default here.
- SHAP
- Per-feature attribution explaining how the model reached a score.
- McNemar
- Paired significance test between two classifiers on the same test cases.
- Bootstrap CI
- Confidence interval from resampling with replacement.
Importance triangulation (B5)
Mean |SHAP| vs permutation importance (test split, B2 seed-42 model); group ablation via neutralization proxy.
Kendall τ (SHAP ranking vs permutation importance): 0.6625
Top features (mean |SHAP|)
- 1overlap_answer_context3.0367
- 2n_words1.1909
- 3n_chars1.1741
- 4nli_ctx_contradicts_ans0.4125
- 5cosine_q_ans0.3399
- 6nli_ctx_neutral_ans0.2835
- 7cosine_ctx_ans0.2293
- 8jaccard_ans_q0.2280
- 9entity_overlap_ratio0.1866
- 10overlap_answer_question0.1810
- 11nli_ans_neutral_ctx0.1747
- 12avg_word_len0.1693
Group ablation ΔF1 (neutralization proxy)
- lexical+0.14703
- length+0.03769
- nli+0.00443
- entity+0.00071
- semantic+0.00069
- numeric+0.00033
- hedging+0.00000
Retrain-based ablation lives in ablation_results.csv (Version A); B5 reports the neutralization proxy (roadmap B5.1).
Neutralization curve (B5.2)
Top-k features set to their test median — prediction change.
| k | Top features | Mean Δscore | ΔF1 |
|---|---|---|---|
| 1 | overlap_answer_context | -0.0793 | +0.0713 |
| 3 | overlap_answer_context, n_words, n_chars | -0.2215 | +0.3029 |
| 5 | overlap_answer_context, n_words, n_chars, nli_ctx_contradicts_ans, cosine_q_ans | -0.2831 | +0.5930 |
| 10 | overlap_answer_context, n_words, n_chars, nli_ctx_contradicts_ans, cosine_q_ans, nli_ctx_neutral_ans, cosine_ctx_ans, jaccard_ans_q, entity_overlap_ratio, overlap_answer_question | -0.4120 | +0.9710 |
Perturbation stability (B5.3)
Full feature re-extraction after text edits; no fixed thresholds.
| Perturbation | n | Mean |Δscore| | |Δ| > 0.3 rate | Top-1 flip | SHAP ρ |
|---|---|---|---|---|---|
| clause_shuffle | 14 | 0.0012 | 0.0% | 0.0% | 0.9212 |
| date | 13 | 0.2786 | 30.8% | 0.0% | 0.8704 |
| entity | 78 | 0.3602 | 37.2% | 0.0% | 0.8397 |
| irrelevant_insert | 120 | 0.0059 | 0.8% | 0.0% | 0.9483 |
| numeric | 22 | 0.3449 | 36.4% | 0.0% | 0.8525 |
| support_removal | 66 | 0.2721 | 28.8% | 0.0% | 0.8742 |
Bootstrap stability (B5.4)
1,000 resamples; top-k set Jaccard + mean-|SHAP| envelope.
Top-5 feature set Jaccard: 1.000 ± 0.000
Reviewer export (B5.5)
40-case manual plausibility sheet — two reviewers, disagreements recorded.
26 cases exported (10 FP / 10 FN / 20 borderline) to b5_review_cases.csv with reviewer_1 / reviewer_2 / agreement columns to fill.
| ID | Label | Raw | Calibrated | Top-5 SHAP features |
|---|---|---|---|---|
| q_8518_correct | 0 | 0.8016 | 0.9354 | overlap_answer_context, n_words, nli_ctx_contradicts_ans, n_chars, jaccard_ans_ctx |
| q_3061_correct | 0 | 0.9135 | 0.9751 | overlap_answer_context, n_words, n_chars, nli_ctx_contradicts_ans, nli_ctx_entails_ans |
| q_7742_correct | 0 | 0.8612 | 0.9609 | overlap_answer_context, n_words, n_chars, nli_ctx_neutral_ans, nli_ans_neutral_ctx |
| q_1402_correct | 0 | 0.7594 | 0.9087 | overlap_answer_context, n_words, nli_ctx_contradicts_ans, n_chars, nli_ans_neutral_ctx |
| q_8650_correct | 0 | 0.9790 | 0.9860 | overlap_answer_context, n_words, n_chars, cosine_q_ans, jaccard_ans_q |
| q_4156_correct | 0 | 0.5273 | 0.5578 | overlap_answer_context, n_words, n_chars, overlap_answer_question, nli_ctx_contradicts_ans |
| q_5922_correct | 0 | 0.6374 | 0.7706 | overlap_answer_context, n_words, n_chars, overlap_answer_question, nli_ctx_neutral_ans |
| q_4329_correct | 0 | 0.7779 | 0.9214 | overlap_answer_context, n_words, n_chars, nli_ctx_contradicts_ans, entity_overlap_ratio |
| q_7876_correct | 0 | 0.7563 | 0.9064 | overlap_answer_context, n_words, n_chars, nli_ctx_contradicts_ans, nli_ctx_neutral_ans |
| q_4539_hallucinated | 1 | 0.2038 | 0.0663 | overlap_answer_context, n_words, n_chars, nli_ctx_neutral_ans, nli_ctx_contradicts_ans |
| q_2910_hallucinated | 1 | 0.2300 | 0.0822 | overlap_answer_context, n_chars, n_words, cosine_q_ans, nli_ctx_contradicts_ans |
| q_1921_hallucinated | 1 | 0.0713 | 0.0214 | overlap_answer_context, n_chars, n_words, overlap_answer_question, jaccard_ans_q |