Why standard NLP benchmark metrics fail to quantify clinical hallucination risk, and how a domain-specific error taxonomy bridges model evaluation and bedside safety.
An auditable, evidence-grounded research pipeline for extracting structured clinical attributes from scientific PDFs with W3C JSON-LD provenance and a 4-stage quality control loop.