Evaluation plan
Create a dataset that reflects your domains, languages, answer lengths, structured outputs, calculations, source quality, and retrieval failure modes. Measure F1 and AUROC, but also inspect false positives, unsupported-span usefulness, latency after warm-up, memory, model downloads, judge cost, and the action your system takes after detection. Run separate tests for wrong evidence, no evidence, tool failure, stale cache, and legitimate conversational memory.
For production, pair answer grounding with input security and source integrity. PrismGuard handles prompt injection; PrismShine handles cause/effect evidence; neither converts untrusted source content into world truth.