Offline walkthrough
Presenter demo
Every score, table, and figure below is rendered from the frozen artifacts/, with no API calls, no OpenAI key, and no network. It works on a clean clone after unzipping the Colab artifact bundle.
01 · The problem in one line
Black-box LLMs hallucinate; you cannot inspect their weights. HaluRISC scores candidate answers against the provided context with 26 evidence-consistency features plus 8 claim-level NLI aggregates (35 features in the deployed EC-XGB model) — and measures honestly where that works and where it breaks.
02 · Three real cases (from the B5 reviewer export)
Sampled from b5_review_cases.json — real scores, no staging data.
Low risk
calibrated 0.066 · raw 0.204Q: The Cincinnati Reds record holder for most stolen bases by a rookie in a season is a native of which city ?
A: Jason Campbell holds the record.
Top-5 SHAP features: overlap_answer_context, n_words, n_chars, nli_ctx_neutral_ans, nli_ctx_contradicts_ans · label 1
High risk
calibrated 0.935 · raw 0.802Q: "Touch Too Much" is the fourth track on an album by a band that hails from which country?
A: Australia
Top-5 SHAP features: overlap_answer_context, n_words, nli_ctx_contradicts_ans, n_chars, jaccard_ans_ctx · label 0
Medium risk
calibrated 0.558 · raw 0.527Q: The 2010 Insight Bowl was the 22nd edition, of the college football bowl game, played at Sun Devil Stadium in Tempe, Arizona on Tuesday, December 28, 2010, it f…
A: 2010 Missouri Tigers football team
Top-5 SHAP features: overlap_answer_context, n_words, n_chars, overlap_answer_question, nli_ctx_contradicts_ans · label 0
03 · A real failure (B5 perturbation protocol)
Controlled edit that flipped or moved the model's score.
Sample q_8327_correct under entity perturbation: raw score 0.0014 → 0.9976 (Δ +0.9962), SHAP top-1 stable (overlap_answer_context → overlap_answer_context).
Zero-shot error case (ragtruth): score 0.9996, true label 0.
04 · Calibration under shift, the key evidence (B4)


Source calibration does not transfer: ECE ≈ 0.8185 on RAGTruth QA test. Fitting on target data restores calibration: ECE ≈ 0.1335 (Platt, disjoint source groups).
05 · EC-XGB under shift (B9)
Standard XGBoost (HaluEval only) vs EC-XGB (multi-source). Flagged = share above the 0.5 threshold.
| Corpus | Model | F1 | AUROC | Flagged | ECE |
|---|---|---|---|---|---|
| RAGTruth (all test tasks) | Standard XGBoost | 0.5181 | 0.4753 | 99.9% | 0.6346 |
| RAGTruth (all test tasks) | EC-XGB (deployed) | 0.5608 | 0.5815 | 61.3% | 0.2782 |
| RAGTruth QA test (held out) | Standard XGBoost | 0.3022 | 0.5465 | 99.9% | 0.8149 |
| RAGTruth QA test (held out) | EC-XGB (deployed) | 0.3081 | 0.5491 | 97.4% | 0.7433 |
| FaithBench | Standard XGBoost | 0.8130 | 0.5089 | 99.5% | 0.2917 |
| FaithBench | EC-XGB (deployed) | 0.1201 | 0.5941 | 7.3% | 0.4500 |
Display calibrator: platt on 5,034 RAGTruth QA rows, held-out ECE 0.7257 → 0.1255.
06 · Transfer robustness (B3)
Zero-shot on external data — figures from artifacts/figures/b3.


| Subset | Rows | F1 | AUROC | ΔF1 vs in-domain |
|---|---|---|---|---|
| ragtruth_qa_test | 900 | 0.3022 | 0.5398 | -0.6825 |
| ragtruth_all | 17790 | 0.6031 | 0.4965 | -0.3815 |
| ragtruth_train | 15090 | 0.6173 | 0.5012 | -0.3673 |
| faithbench | 750 | 0.8130 | 0.5314 | -0.1716 |
| ragtruth_task_qa | 5934 | 0.4515 | 0.5868 | -0.5331 |
| ragtruth_task_summarization | 5658 | 0.4603 | 0.6317 | -0.5244 |
| ragtruth_task_data_to_text | 6198 | 0.8140 | 0.5777 | -0.1706 |
Live mode needs the local API (uvicorn src.api.main:app) — or the chat needs OPENAI_API_KEY. This page needs neither.