...
Maven
Generative AI-Enabled Medical Devices

Generative AI-Enabled Medical Devices: FDA’s Emerging Regulatory Considerations.

“For generative AI-enabled medical devices, the regulatory question is no longer only whether AI can generate an answer—it is whether the complete device can generate clinically appropriate outputs, operate within defined boundaries, and maintain safe performance throughout its lifecycle.”

Generative artificial intelligence (GenAI) is increasingly being incorporated into healthcare software, creating new opportunities for clinical decision support, information generation, patient interaction, and workflow automation. Unlike conventional software or many traditional AI-enabled device functions, GenAI systems can produce variable outputs, process open-ended inputs, perform multiple tasks, and potentially operate with increasing levels of autonomy.

These characteristics create new considerations for demonstrating the safety and effectiveness of medical devices that incorporate GenAI.

Importantly, the publication is a discussion paper, not draft or final guidance. It does not establish new regulatory requirements or represent FDA’s final regulatory expectations. Rather, it provides insight into the scientific and regulatory questions FDA is considering as GenAI capabilities become increasingly integrated into medical devices.

Why Does Generative AI Require Additional Regulatory Consideration?

AI is already used in medical devices for functions such as image analysis, detection, prediction, monitoring, and clinical decision support. GenAI introduces additional complexity because its outputs may be variable and context-dependent.

Depending on the intended use, a GenAI-enabled medical device may:

  • Generate text, images, or other content;
  • Process open-ended or conversational inputs;
  • Produce different outputs for similar inputs;
  • Perform multiple subtasks;
  • Retrieve information from external sources;
  • Depend on a third-party foundation model; or
  • Perform multi-step actions with varying degrees of autonomy.

Therefore, evaluation based solely on predefined input-output pairs may not fully characterize the safety and performance of a GenAI-enabled device.

FDA’s Potential Risk-Based Approach

The discussion paper explores a potential framework for considering GenAI device risk using two primary dimensions:

Risk dimension Regulatory question
Degree of device activity How independently does the device generate, recommend, or execute an action?
Consequence of an incorrect output What is the potential harm if the output is incorrect and relied upon?

For example, a GenAI system providing general educational information may present a different risk profile from a system that provides patient-specific treatment recommendations or autonomously initiates an action.

Premarket Evaluation: Beyond Conventional Accuracy

A central concept discussed by FDA is a potential competency-based approach to evaluating GenAI-enabled medical devices.

Rather than relying on a single accuracy metric, FDA discusses evaluating multiple dimensions of device performance through:

Device Benchmarking + Clinical Confirmation

Device Benchmarking

Benchmarking could assess whether the device demonstrates appropriate:

  • Safety;
  • Clinical proficiency;
  • Generalizability.
  • Agentic AI Capabilities

For systems with agentic capabilities, additional evaluation may address planning, tool use, human-oversight checkpoints, and other autonomous behaviors.

Planning an FDA Submission for an AI Enabled Medical Device?

Regulatory expectations for AI enabled devices can involve device characterization, verification and validation, software documentation, risk management, and performance evidence. Maven can help you assess your regulatory pathway and prepare the necessary documentation for the US market.

Talk to our FDA regulatory consultants

What Should Be Evaluated?

Safety

Safety evaluation may need to determine whether the device can:

  • Recognize safety-critical situations;
  • Appropriately escalate concerning cases;
  • Remain within its intended scope; and
  • Communicate uncertainty appropriately.

A fluent response should not be interpreted as evidence of clinical correctness.

Clinical Proficiency

Depending on the intended use, evaluation may consider:

  • Clinical knowledge;
  • Clinical reasoning;
  • Information gathering;
  • Interpretation of incomplete or conflicting information; and
  • Quantitative reasoning.

This is particularly relevant where outputs involve medication dosing, laboratory values, physiological measurements, or other clinical calculations.

Generalizability

Performance should be considered across relevant variations in:

  • Users;
  • Inputs;
  • Clinical scenarios;
  • Conversational context;
  • Patient populations; and
  • Repeated interactions.

Agentic AI Capabilities

Competencies applicable only to agentic GenAI-enabled devices

Postmarket Monitoring: Maintaining Performance After Authorization

For conventional software, a validated version may remain relatively stable. GenAI systems can introduce additional sources of change.

Performance may be affected by:

  • Foundation-model updates;
  • Changes in system architecture;
  • Changes to data sources;
  • Changes in deployment environments;
  • Changes in user populations; or
  • Emerging real-world use patterns.

Consequently, FDA discusses risk-proportionate postmarket monitoring as an important element of lifecycle oversight.

Potential activities could include:

  • Periodic benchmarking;
  • Monitoring for performance degradation;
  • Review of real-world outputs;
  • Monitoring for emerging failure modes;
  • Subgroup performance assessment; and
  • Re-evaluation following significant system or model changes.

This supports a shift from a “test once and release” mindset toward continuous performance assurance throughout the product lifecycle.

Foundation Models: A New Layer of Regulatory Dependency

Many GenAI-enabled medical devices may rely on foundation models developed by third parties.

This creates an additional regulatory challenge:

The medical device manufacturer may control the application, while another organization controls the underlying model.

A change to the foundation model could potentially affect the performance of the finished medical device.

Manufacturers may therefore need appropriate mechanisms to understand and manage:

  • Model versions;
  • Model changes;
  • Known limitations;
  • Performance characteristics;
  • Update processes;
  • Change notifications; and
  • Validation requirements.

FDA’s discussion paper explores potential mechanisms including contractual controls, technical controls, and Predetermined Change Control Plans (PCCPs).

The fundamental regulatory issue remains continued manufacturer control over the safety and effectiveness of the finished medical device.

Foundation Model Master Files

FDA also discusses the potential use of Foundation Model Master Files (MAFs).

The concept could provide a mechanism for foundation-model developers to make relevant information available to FDA for potential reference by medical device manufacturers.

However, a Foundation Model MAF would not constitute FDA authorization of the foundation model for medical use.

The manufacturer of the finished medical device would remain responsible for demonstrating that its device is safe and effective for its intended use.

Agentic AI: When Generative AI Begins to Act

The regulatory considerations become more complex when GenAI systems are capable of autonomous, multi-step activities.

An agentic system may:

Plan → Select tools → Execute tasks → Evaluate results → Continue acting

This introduces considerations beyond conventional text or content generation, including:

  • Degree of autonomy;
  • Human oversight;
  • Tool access;
  • Multi-step execution;
  • Failure handling;
  • Action authorization;
  • Prompt-injection resistance; and
  • Interaction with other medical devices.

The potential consequences become particularly significant where an AI system can independently execute actions that affect patient care or control another medical device.

Developing a Medical Device With Generative or Agentic AI?

Before moving toward a US market submission, manufacturers should consider how the AI model, intended use, software architecture, validation strategy, and ongoing changes affect the device’s regulatory pathway.

Discuss your regulatory requirements with Maven

What Should Manufacturers Consider?

Although the discussion paper does not establish new requirements, it highlights several areas that manufacturers developing GenAI-enabled medical devices should consider as part of a lifecycle-based development strategy.

Lifecycle area Core Consideration
Intended use Define clinical purpose, users, population, environment and autonomy
Risk management Assess consequences of incorrect outputs and device activity
System characterization Understand models, prompts, retrieval, guardrails and system architecture
Verification & validation Evaluate safety, clinical performance and generalizability
Human factors Assess comprehension, reliance and automation bias
Cybersecurity Consider prompt injection, data integrity and system vulnerabilities
Change control Define controls for model and system modifications
Postmarket surveillance Monitor real-world performance and emerging risks
Third-party governance Establish oversight of foundation-model dependencies

Is the FDA Discussion Paper a Regulatory Requirement?

No

The FDA publication is a discussion paper and request for feedback. It is not:

  • Final guidance;
  • Draft guidance;
  • A new regulation;
  • A new device classification system; or
  • A binding set of submission requirements.

The concepts discussed by FDA should therefore not be presented as mandatory regulatory requirements.

Instead, the paper provides an indication of the scientific and regulatory issues FDA is evaluating as it considers future approaches to GenAI-enabled medical devices.

Regulatory Takeaways

  • The regulatory significance of GenAI depends on intended use, device activity, human oversight, and consequences of error.
  • Accuracy is only one component of performance.
  • Safety, clinical proficiency, communication, generalizability, and—in applicable systems—agentic capabilities may also be relevant.
  • Premarket evaluation may need to be complemented by lifecycle monitoring.
  • Real-world performance and changes to the underlying system may require continued assessment.
  • Third-party foundation models introduce additional controls.
  • Autonomy changes the risk profile.
  • The transition from generating information to recommending or executing actions introduces additional regulatory considerations.

Conclusion

The FDA’s discussion paper marks an important development in the regulatory conversation surrounding GenAI-enabled medical devices. It recognizes that the characteristics that make GenAI powerful—flexible inputs, variable outputs, broad capabilities, third-party foundation models, and increasing autonomy—also create challenges for conventional approaches to medical device evaluation.

FDA’s discussion points toward a potential risk-based and lifecycle-oriented framework in which device benchmarking, clinical confirmation, postmarket monitoring, model-change management, and assessment of autonomous capabilities may collectively contribute to demonstrating safety and effectiveness.

For manufacturers, the central consideration is not simply whether a GenAI model performs well in isolation. The more important question is whether the complete medical device consistently performs its intended function, within defined boundaries and under clinically relevant conditions, while risks remain controlled throughout the product lifecycle.

Reference

1. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback

Frequently Asked Questions

The FDA discussion paper explores scientific and regulatory considerations for medical devices that use generative artificial intelligence. It discusses potential approaches to evaluating safety, clinical performance, generalizability, agentic capabilities, foundation model dependencies, and postmarket monitoring. The publication is a discussion paper and request for feedback, not draft or final guidance.

The FDA discussion paper does not establish new regulations or mandatory requirements specifically for generative AI enabled medical devices. It presents regulatory and scientific considerations that may inform future FDA approaches to GenAI enabled medical devices.

The FDA discussion paper explores a potential risk based approach that considers factors such as the degree of device activity and the potential consequences of incorrect outputs. It also discusses competency based evaluation involving safety, clinical proficiency, generalizability, and, where applicable, agentic AI capabilities.

Potential risks can arise from variable outputs, incorrect or misleading information, inappropriate clinical recommendations, limited generalizability, model changes, third party foundation models, cybersecurity vulnerabilities, and autonomous actions. The level of risk depends significantly on the device’s intended use and the consequences of an incorrect output.

Manufacturers should consider the device’s safety, clinical proficiency, generalizability, intended use, human oversight, system architecture, foundation model dependencies, cybersecurity, change control, and postmarket performance. For agentic systems, evaluation may also need to consider planning, tool use, autonomous actions, and human oversight checkpoints.

Agentic AI refers to GenAI systems capable of performing multi step activities with varying degrees of autonomy. An agentic medical device may plan actions, select or use tools, execute tasks, evaluate results, and continue acting. These capabilities can introduce additional considerations related to autonomy, human oversight, authorization, failure handling, and interaction with other systems or devices.

Manufacturers may need controls to understand and manage foundation model versions, changes, limitations, performance characteristics, update processes, and validation requirements. The FDA discussion paper also explores potential mechanisms such as contractual controls, technical controls, and Predetermined Change Control Plans.

A Foundation Model Master File is a concept discussed by FDA that could allow foundation model developers to provide relevant information to FDA for potential reference by medical device manufacturers. It would not represent FDA authorization of the foundation model for medical use. The manufacturer of the finished medical device would remain responsible for demonstrating that its device is safe and effective for its intended use.

GenAI device performance can potentially be affected by foundation model updates, system changes, data source changes, deployment environments, user populations, and emerging real world use patterns. Risk proportionate postmarket monitoring can therefore help manufacturers identify performance degradation, emerging failure modes, subgroup performance issues, and changes requiring further evaluation.

No. Accuracy alone may not fully characterize the performance of a GenAI enabled medical device. Depending on the intended use, evaluation may also need to consider safety, clinical proficiency, generalizability, communication of uncertainty, human oversight, and agentic capabilities.

Get In Touch With Us

Have questions? We're here to help.

Business Enquiries

We'll respond shortly