Cortico Launches MedSafe-Dx, an Open Benchmark Exposing Clinical Safety Gaps in Frontier AI Models

Across the 11 frontier large language models (LLMs) evaluated in the launch paper, MedSafe-Dx found the safest model still unnecessarily escalated 71% of routine cases, while the worst missed 17% of medical emergencies.