The Food and Drug Administration on August 18 opened a public discussion about a different way to evaluate medical devices powered by generative artificial intelligence: assess whether the finished device can demonstrate defined clinical competencies, then confirm its performance under realistic conditions and continue monitoring it after deployment.

The Center for Devices and Radiological Health presented the concept in a 31-page discussion paper and requested feedback by October 19 under docket FDA-2026-N-7874. The agency repeatedly emphasizes that the document is neither draft nor final guidance, establishes no new evidentiary requirement, and does not determine whether every approach discussed fits within its existing legal authority.

That distinction matters. FDA is asking how a regulatory framework might work; it has not adopted one or authorized any product through the proposed framework.

From static testing to demonstrated competency

Traditional software can often be tested against a representative set of bounded inputs and expected outputs. Generative systems complicate that model because they can accept open-ended prompts, produce different responses to similar inputs, perform multiple tasks, and change when underlying models, retrieval systems, prompts, guardrails, or interfaces are updated.

FDA’s possible competency-based approach has two main components. First, sponsors would benchmark the final user-facing device—not merely its underlying foundation model—for safety behavior, clinical proficiency, communication, reproducibility, subgroup performance, and, where relevant, agentic capabilities. Tests and acceptance criteria would be specified in advance and tailored to the device’s intended use and risk.

Second, clinical confirmation would examine whether the device performs as intended in real or clinically representative use. FDA says this would not necessarily require a prospective clinical study in every case. The appropriate design and evidentiary rigor would depend on the device’s function and risk profile.

This is a regulatory framework, not a clinical trial. It has no enrolled patient population, comparator assignment, or clinical endpoint, and it supplies no evidence that generative-AI devices improve health outcomes.

Risk would depend on what the system does

The paper proposes a two-axis risk heuristic combining the independence of a device’s activity with the consequences of an incorrect output. A system providing general information would occupy a different position from one directing a patient’s action, assigning a diagnosis, prescribing medication, or initiating a clinical order.

FDA also highlights risks that can emerge during extended conversations. A chatbot may begin with general information but become increasingly directive, while a failure in emergency-care escalation can cause harm through either under-referral or unnecessary over-referral. Patient-facing tools may deserve different scrutiny when users cannot independently recognize an incorrect output.

For performance comparisons, the agency asks whether some devices should be measured against qualified clinicians’ consensus or a median clinician performing the same task. Other products could still be assessed against conventional clinical reference standards, such as biopsy-confirmed diagnoses. The paper leaves unresolved whether the correct comparator is a generalist, specialist, clinician-AI team, or autonomous system.

Postmarket monitoring becomes central

Because premarket tests cannot cover every future prompt, clinical setting, or software change, FDA is considering periodic re-benchmarking, clinician review of sampled real-world outputs, and monitoring for performance degradation. It also asks whether greater postmarket surveillance could justify accepting more uncertainty before authorization—a consequential question that remains open.

Public benchmarks have important limitations. Training-data contamination can allow models to recognize test material, while saturated or unrepresentative datasets may overstate performance in actual clinics. Independent test sets, expert adjudication, subgroup analysis, and clear triggers for reassessment could therefore determine whether competency testing is meaningful.

For clinicians and patients, nothing changes immediately. The paper does not certify generative AI as equivalent to a physician, expand FDA’s jurisdiction over general-purpose AI, or replace existing medical-device requirements. Its practical significance is that the regulator has now put a detailed evaluation model into public discussion, including the question of how much evidence should be required before and after increasingly autonomous clinical software reaches patients.

Primary sourceFDA discussion paper: Considerations for the Regulation of Generative AI-Enabled Medical Devices

The source ledger and revision history are retained with the newsroom record.

AI-assisted reporting disclosure

AI assisted with source organization and drafting. Vitalspan Wire is accountable for the published text and maintains a revision record.

Medical note

This article provides general information, not diagnosis or treatment advice. Consult a qualified clinician before making medical decisions.