Blog & Research Notes
Reflections, technical deep-dives, and methodology notes on clinical AI evaluation, evidence grounding, and health data infrastructure by Dr. Soroush Dianaty.
Reflections, technical deep-dives, and methodology notes on clinical AI evaluation, evidence grounding, and health data infrastructure by Dr. Soroush Dianaty.
Following an unprecedented 19-day export control freeze, Anthropic’s Claude Fable 5 and Mythos 5 are back online under strict restrictions. Here is a deep analysis of Project Glasswing, real-time KYC, and how health-tech teams can build resilient AI architectures.
Why standard NLP benchmark metrics fail to quantify clinical hallucination risk, and how a domain-specific error taxonomy bridges model evaluation and bedside safety.
Personal reflections on transitioning from practicing family medicine across rural and urban clinics to developing rigorous clinical AI evaluation frameworks at ASU.
A practical guide to HL7 FHIR Security Labels, 42 CFR Part 2 compliance, and context-aware LLM classifiers for sensitive health data exchange.
Why general LLM benchmarks like MMLU or GSM8K fall short in medicine, and how evidence grounding, hallucination bounds, and FHIR interoperability redefine clinical AI safety.
Anthropic’s Fable 5 suspension shows why health-tech AI needs due process, lifecycle governance, validated fallbacks, and risk-based regulation—not opaque shutdowns over imperfect models.