“For generative AI-enabled medical devices, the regulatory question is no longer only whether AI can generate an answer—it is whether the complete device can generate clinically appropriate outputs, operate within defined boundaries, and maintain safe performance throughout its lifecycle.”
Generative artificial intelligence (GenAI) is increasingly being incorporated into healthcare software, creating new opportunities for clinical decision support, information generation, patient interaction, and workflow automation. Unlike conventional software or many traditional AI-enabled device functions, GenAI systems can produce variable outputs, process open-ended inputs, perform multiple tasks, and potentially operate with increasing levels of autonomy.
These characteristics create new considerations for demonstrating the safety and effectiveness of medical devices that incorporate GenAI.
Importantly, the publication is a discussion paper, not draft or final guidance. It does not establish new regulatory requirements or represent FDA’s final regulatory expectations. Rather, it provides insight into the scientific and regulatory questions FDA is considering as GenAI capabilities become increasingly integrated into medical devices.
AI is already used in medical devices for functions such as image analysis, detection, prediction, monitoring, and clinical decision support. GenAI introduces additional complexity because its outputs may be variable and context-dependent.
Depending on the intended use, a GenAI-enabled medical device may:
Therefore, evaluation based solely on predefined input-output pairs may not fully characterize the safety and performance of a GenAI-enabled device.
The discussion paper explores a potential framework for considering GenAI device risk using two primary dimensions:
| Risk dimension | Regulatory question |
| Degree of device activity | How independently does the device generate, recommend, or execute an action? |
| Consequence of an incorrect output | What is the potential harm if the output is incorrect and relied upon? |
For example, a GenAI system providing general educational information may present a different risk profile from a system that provides patient-specific treatment recommendations or autonomously initiates an action.
A central concept discussed by FDA is a potential competency-based approach to evaluating GenAI-enabled medical devices.
Rather than relying on a single accuracy metric, FDA discusses evaluating multiple dimensions of device performance through:
Device Benchmarking + Clinical Confirmation
Device Benchmarking
Benchmarking could assess whether the device demonstrates appropriate:
For systems with agentic capabilities, additional evaluation may address planning, tool use, human-oversight checkpoints, and other autonomous behaviors.
Regulatory expectations for AI enabled devices can involve device characterization, verification and validation, software documentation, risk management, and performance evidence. Maven can help you assess your regulatory pathway and prepare the necessary documentation for the US market.
Talk to our FDA regulatory consultants
Safety evaluation may need to determine whether the device can:
A fluent response should not be interpreted as evidence of clinical correctness.
Depending on the intended use, evaluation may consider:
This is particularly relevant where outputs involve medication dosing, laboratory values, physiological measurements, or other clinical calculations.
Performance should be considered across relevant variations in:
Competencies applicable only to agentic GenAI-enabled devices
For conventional software, a validated version may remain relatively stable. GenAI systems can introduce additional sources of change.
Performance may be affected by:
Consequently, FDA discusses risk-proportionate postmarket monitoring as an important element of lifecycle oversight.
Potential activities could include:
This supports a shift from a “test once and release” mindset toward continuous performance assurance throughout the product lifecycle.
Many GenAI-enabled medical devices may rely on foundation models developed by third parties.
This creates an additional regulatory challenge:
The medical device manufacturer may control the application, while another organization controls the underlying model.
A change to the foundation model could potentially affect the performance of the finished medical device.
Manufacturers may therefore need appropriate mechanisms to understand and manage:
FDA’s discussion paper explores potential mechanisms including contractual controls, technical controls, and Predetermined Change Control Plans (PCCPs).
The fundamental regulatory issue remains continued manufacturer control over the safety and effectiveness of the finished medical device.
FDA also discusses the potential use of Foundation Model Master Files (MAFs).
The concept could provide a mechanism for foundation-model developers to make relevant information available to FDA for potential reference by medical device manufacturers.
However, a Foundation Model MAF would not constitute FDA authorization of the foundation model for medical use.
The manufacturer of the finished medical device would remain responsible for demonstrating that its device is safe and effective for its intended use.
The regulatory considerations become more complex when GenAI systems are capable of autonomous, multi-step activities.
An agentic system may:
Plan → Select tools → Execute tasks → Evaluate results → Continue acting
This introduces considerations beyond conventional text or content generation, including:
The potential consequences become particularly significant where an AI system can independently execute actions that affect patient care or control another medical device.
Before moving toward a US market submission, manufacturers should consider how the AI model, intended use, software architecture, validation strategy, and ongoing changes affect the device’s regulatory pathway.
Discuss your regulatory requirements with Maven
Although the discussion paper does not establish new requirements, it highlights several areas that manufacturers developing GenAI-enabled medical devices should consider as part of a lifecycle-based development strategy.
| Lifecycle area | Core Consideration |
| Intended use | Define clinical purpose, users, population, environment and autonomy |
| Risk management | Assess consequences of incorrect outputs and device activity |
| System characterization | Understand models, prompts, retrieval, guardrails and system architecture |
| Verification & validation | Evaluate safety, clinical performance and generalizability |
| Human factors | Assess comprehension, reliance and automation bias |
| Cybersecurity | Consider prompt injection, data integrity and system vulnerabilities |
| Change control | Define controls for model and system modifications |
| Postmarket surveillance | Monitor real-world performance and emerging risks |
| Third-party governance | Establish oversight of foundation-model dependencies |
No
The FDA publication is a discussion paper and request for feedback. It is not:
The concepts discussed by FDA should therefore not be presented as mandatory regulatory requirements.
Instead, the paper provides an indication of the scientific and regulatory issues FDA is evaluating as it considers future approaches to GenAI-enabled medical devices.
The FDA’s discussion paper marks an important development in the regulatory conversation surrounding GenAI-enabled medical devices. It recognizes that the characteristics that make GenAI powerful—flexible inputs, variable outputs, broad capabilities, third-party foundation models, and increasing autonomy—also create challenges for conventional approaches to medical device evaluation.
FDA’s discussion points toward a potential risk-based and lifecycle-oriented framework in which device benchmarking, clinical confirmation, postmarket monitoring, model-change management, and assessment of autonomous capabilities may collectively contribute to demonstrating safety and effectiveness.
For manufacturers, the central consideration is not simply whether a GenAI model performs well in isolation. The more important question is whether the complete medical device consistently performs its intended function, within defined boundaries and under clinically relevant conditions, while risks remain controlled throughout the product lifecycle.
The FDA discussion paper explores scientific and regulatory considerations for medical devices that use generative artificial intelligence. It discusses potential approaches to evaluating safety, clinical performance, generalizability, agentic capabilities, foundation model dependencies, and postmarket monitoring. The publication is a discussion paper and request for feedback, not draft or final guidance.
The FDA discussion paper does not establish new regulations or mandatory requirements specifically for generative AI enabled medical devices. It presents regulatory and scientific considerations that may inform future FDA approaches to GenAI enabled medical devices.
The FDA discussion paper explores a potential risk based approach that considers factors such as the degree of device activity and the potential consequences of incorrect outputs. It also discusses competency based evaluation involving safety, clinical proficiency, generalizability, and, where applicable, agentic AI capabilities.
Potential risks can arise from variable outputs, incorrect or misleading information, inappropriate clinical recommendations, limited generalizability, model changes, third party foundation models, cybersecurity vulnerabilities, and autonomous actions. The level of risk depends significantly on the device’s intended use and the consequences of an incorrect output.
Manufacturers should consider the device’s safety, clinical proficiency, generalizability, intended use, human oversight, system architecture, foundation model dependencies, cybersecurity, change control, and postmarket performance. For agentic systems, evaluation may also need to consider planning, tool use, autonomous actions, and human oversight checkpoints.
Agentic AI refers to GenAI systems capable of performing multi step activities with varying degrees of autonomy. An agentic medical device may plan actions, select or use tools, execute tasks, evaluate results, and continue acting. These capabilities can introduce additional considerations related to autonomy, human oversight, authorization, failure handling, and interaction with other systems or devices.
Manufacturers may need controls to understand and manage foundation model versions, changes, limitations, performance characteristics, update processes, and validation requirements. The FDA discussion paper also explores potential mechanisms such as contractual controls, technical controls, and Predetermined Change Control Plans.
A Foundation Model Master File is a concept discussed by FDA that could allow foundation model developers to provide relevant information to FDA for potential reference by medical device manufacturers. It would not represent FDA authorization of the foundation model for medical use. The manufacturer of the finished medical device would remain responsible for demonstrating that its device is safe and effective for its intended use.
GenAI device performance can potentially be affected by foundation model updates, system changes, data source changes, deployment environments, user populations, and emerging real world use patterns. Risk proportionate postmarket monitoring can therefore help manufacturers identify performance degradation, emerging failure modes, subgroup performance issues, and changes requiring further evaluation.
No. Accuracy alone may not fully characterize the performance of a GenAI enabled medical device. Depending on the intended use, evaluation may also need to consider safety, clinical proficiency, generalizability, communication of uncertainty, human oversight, and agentic capabilities.
Recent Post
Clinical Investigation Plan (CIP) Under EU MDR 2017/745
Are You Looking For Medical Devices Certifications?
Contact UsHave questions? We're here to help.
We'll respond shortly