Clinical LLMs

3 min read

What 'Hallucination' Actually Means in Clinical LLMs (And How to Measure It)

Why standard NLP benchmark metrics fail to quantify clinical hallucination risk, and how a domain-specific error taxonomy bridges model evaluation and bedside safety.

Clinical LLMs AI Safety LLM Evaluation Evidence Grounding
Evaluating Clinical LLMs: Beyond Standard NLP Benchmarks
3 min read

Evaluating Clinical LLMs: Beyond Standard NLP Benchmarks

Why general LLM benchmarks like MMLU or GSM8K fall short in medicine, and how evidence grounding, hallucination bounds, and FHIR interoperability redefine clinical AI safety.

Clinical LLMs AI Safety Evidence Grounding Health Informatics

Early Evidence for Context-Aware Large Language Models (LLMs) in Sensitive Health Data Classification

Presented early evidence that context-aware LLMs can classify sensitive health data for consent-driven record sharing.

Clinical LLMs Sensitive Health Data