A: Australia
top signal: overlap_answer_context, n_words, nli_ctx_contradicts_ans, n_chars, jaccard_ans_ctx
Loading view…
Calibrated · Explainable · Lightweight
HaluRISC scores candidate answers against the evidence you provide. It runs on CPU in milliseconds, needs no model internals, and shows where its own limits are.
A: Australia
top signal: overlap_answer_context, n_words, nli_ctx_contradicts_ans, n_chars, jaccard_ans_ctx
A: Jason Campbell holds the record.
top signal: overlap_answer_context, n_words, n_chars, nli_ctx_neutral_ans, nli_ctx_contradicts_ans
Real cases from the B5 reviewer export · scores are the evidence-calibrated risk
Preflight for the booth. Offline pages (/demo, /dashboard) never depend on these checks.
A question, the evidence you trust (pasted, uploaded, or searched), and the candidate answer. No model weights or activations needed.
26 features across seven groups: lexical overlap, entity coverage, numeric consistency, NLI contradiction, hedging, semantic drift, and length. A calibrated XGBoost turns them into a risk score.
Claim-level supported / contradicted / unsupported verdicts, SHAP attribution for every score, and the failure cases kept in the open.
Type or paste a case and inspect the score, thresholds, and SHAP values side by side.
Converse with a grounded assistant; every answer is checked after it streams.
All B2 to B5 tables, charts, and statistics, rendered straight from the frozen artifacts.
A fully offline walkthrough: no API key, no network, works on a clean clone.
Zero-shot transfer to RAGTruth QA drops well below in-domain performance, and source calibration does not transfer to new data without refitting. Both effects ship in the dashboard with the exact numbers.