FDA weighs clinician-style oversight for generative AI medical devices
Generative AI-enabled medical devices challenge traditional FDA pathways such as 510(k) premarket clearance, De Novo classification and premarket approval because their inputs and outputs can be open ended, while models can change frequently. Tools built on foundation models from OpenAI, Anthropic or Google also raise attribution questions when harmful outputs occur and visibility into training data, architecture or evaluation is limited.
The FDA is considering a framework that sorts devices by what they do and the severity of harm if an output is wrong. It would treat directiveness as a spectrum, distinguish informational outputs from recommendations or autonomous actions, and assess whether clinicians can independently verify results. Premarket review could use clinician-style competency benchmarks, staged clinical confirmation and independent adjudication.
Postmarket oversight would become central, with re-benchmarking after updates, clinician review of real-world outputs and drift monitoring. The UK’s MHRA has drawn a related line for ambient voice tools: summarizing encounters or formatting transcripts is not a device, while generating new clinical insights, diagnoses, follow-up or treatment options is. Open issues include accountability, who funds ongoing monitoring and what comparator should define acceptable performance.