Version B evidence · B1 to B5
About HaluRISC
Calibrated, explainable hallucination risk prediction for black-box LLM outputs.
Pipeline Architecture
HaluRISC evaluates candidate LLM answers against provided context without inspecting model weights or activations. Features are computed across seven feature groups: length/style, lexical overlap, entity coverage, numeric consistency, hedging density, NLI contradiction, and semantic embeddings (26 features total — read from feature_names.json).
1. Feature Extraction
26 engineered features across 7 groups measuring grounding, entity coverage, NLI consistency, numeric novelty, and semantic drift (B1/B2).
2. Calibrated XGBoost
XGBoost with group-aware splits (no leakage) and Platt scaling; calibration-under-shift is measured in B4 — source calibration does not transfer, target calibration fixes it.
3. SHAP Explanations
TreeExplainer attributions identifying features raising or lowering risk. B5 validates SHAP stability (top-1 never flips under controlled perturbations).
4. Conversational AI UI
Next.js + assistant-ui for natural-language explanations with Generative UI; offline /demo walkthrough needs no API key.
Evidence (all numbers from artifacts)
In-domain test F1 ≈ 0.98 with leakage-free grouped splits (B2) · zero-shot transfer drops on RAGTruth QA (AUROC ≈ 0.54) while FaithBench stays usable (F1 ≈ 0.81) (B3) · ECE 0.81 → 0.13 after target calibration (B4) · SHAP top-5 set Jaccard = 1.0 and top-1 feature never flips under perturbation (B5). Open the Dashboard for the full tables.
Reading the interface
Short version for a first demo. Every screen also carries its own "How to read" notes next to the numbers, and the repository guide docs/04-interface-guide.md explains each element with examples.
Chat mode
Ask a question and every answer gets a risk card: headline verdict, claim-level supported or contradicted checks, citations, and the features that moved the score.
Analyze mode
Score one answer, or two side by side. The gauge shows the calibrated probability with its low, medium, and high cutoffs. The SHAP chart shows what pushed the raw score up or down.
Dashboard
One tab per research question, from baselines to calibration and failure cases. A glossary at the top of the page defines every metric name.
Presenter demo
The same evidence in one offline page. No API key and no network, so it works on a clean clone.
The thumbs buttons on a chat risk card record your agreement into a local feedback file on this machine. They feed the error-analysis appendix, not the model.
Core Tech Stack
Versions above are read from manifest.json (the frozen Colab run environment); trained on NVIDIA GeForce RTX 3060 Laptop GPU. Seeds 42/123/456.