Version B evidence
Experiment dashboard
Every value is read from artifacts/results at request time, never hardcoded.
How to read these metrics
- F1
- Balance of precision and recall. The headline detection metric on the test split.
- AUROC
- Probability that a random positive case scores above a random negative one.
- PR-AUC
- Precision-recall area. More informative than AUROC when classes are imbalanced.
- MCC
- Matthews correlation coefficient. A balanced single-number quality score.
- ECE
- Expected calibration error. Gap between predicted probabilities and observed frequencies.
- Brier
- Mean squared error of the predicted probabilities. Lower is better.
- Platt / isotonic
- Post-hoc calibration methods. Platt is the deployable default here.
- SHAP
- Per-feature attribution explaining how the model reached a score.
- McNemar
- Paired significance test between two classifiers on the same test cases.
- Bootstrap CI
- Confidence interval from resampling with replacement.
Per-module latency (measured)
latency_analysis.json — p50 / p95 / mean milliseconds per sample.
| Module | p50 (ms) | p95 (ms) | Mean (ms) |
|---|---|---|---|
| Lexical | 0.1 | 0.1 | 0.1 |
| Entity | 21.1 | 47.2 | 25.3 |
| NLI | 27.1 | 68.6 | 40.1 |
| Semantic | 9.3 | 15.6 | 12.6 |
| XGBoost | 1.3 | 8.8 | 2.6 |
| SHAP | 1.9 | 15.5 | 3.6 |
Total per sample: p50 62 ms · model artifacts ≈ 1.64 MB
Cost per 1,000 predictions (measured)
From latency/cost artifacts — never hardcoded.
- HaluRISC (local CPU/GPU)
- $0.0010
- LLM-as-Judge estimate
- $0.1050
- Ratio
- 105× cheaper
LLM-as-Judge vs XGBoost (200 samples, gpt-5.6-luna)
llm_judge_results.json — agreement and cost measured on the same subset.
| Model | Accuracy | F1 | Latency p50 (ms) |
|---|---|---|---|
| gpt-5.6-luna judge | 0.860 | 0.841 | 1310 |
| XGBoost (ours) | 0.985 | 0.985 | 62 |
Agreement with XGBoost: 0.845 · McNemar p = 0.0000 · judge run cost $0.0210