LLM Evaluation

3 min read

What 'Hallucination' Actually Means in Clinical LLMs (And How to Measure It)

Why standard NLP benchmark metrics fail to quantify clinical hallucination risk, and how a domain-specific error taxonomy bridges model evaluation and bedside safety.

Clinical LLMs AI Safety LLM Evaluation Evidence Grounding
4 min read

EviTrace: Evidence-Grounded PDF Extraction for Clinical Research

An auditable, evidence-grounded research pipeline for extracting structured clinical attributes from scientific PDFs with W3C JSON-LD provenance and a 4-stage quality control loop.

Evidence Grounding LLM Evaluation Research Tooling Clinical AI Safety