The FDA Asked How to Regulate Generative AI Devices. Physicians Have Until October 19.
On August 18, 2026, the FDA's Center for Devices and Radiological Health published a discussion paper on generative AI medical devices. It is not draft guidance. It is not final policy. It is a set of questions, and the comment window closes on October 19, 2026.
That distinction matters. A discussion paper is the moment when a physician's comment can still change the questions the agency treats as serious. After a draft guidance, the frame is already set.
The paper is Considerations for the Regulation of Generative AI-Enabled Medical Devices. Comments go to docket FDA-2026-N-7874.
The FDA is not asking whether generative AI belongs in medicine. It is asking how to tell a competent device from a fluent one.What the paper is actually proposing
CDRH starts from a fact clinicians already live with. A generative model does not behave like a calculator. The same question can produce different answers. The input can be an open-ended conversation. The model underneath the product may belong to a third party, and that third party can change it after the hospital has gone live.
The paper organizes the response in three moves.
Risk, on two axes. One axis is what the software does: offer information, direct an action, or take the action. The other axis is the harm if someone relies on a wrong output. A note that lisinopril is sometimes increased when blood pressure stays high is not the same object as an instruction to raise the dose from 10 mg to 20 mg. A suggestion to use over-the-counter hydrocortisone for poison ivy is not the same object as an instruction to change a basal insulin dose. Those examples are the FDA's, and they are the right clinical instinct.
A competency exam, not an endless test of every possible input. CDRH is considering a premarket approach modeled, loosely, on how physicians are credentialed: structured tests of knowledge and judgment, then supervised performance in realistic conditions. In the paper this is device benchmarking, followed by clinical confirmation. The thing being tested is the product a user will actually see, not the foundation model sitting alone in a lab.
Watching the device after it ships. Because these products keep changing, CDRH is considering whether some uncertainty at approval can be traded for harder postmarket monitoring: repeat the benchmark, have independent clinicians sample real outputs, and watch for drift when the patients, the data, or the underlying model shift.
None of this is a rule yet. The paper says so in the first paragraph.
The risk that a disclaimer does not cancel
The most useful section for practicing physicians is the one on "action-directing" language.
CDRH is considering a continuum, not a switch. General information. A statement tied to this patient's situation. An endorsement of a specific action. A direct order. The agency is also considering that a patient-facing answer does not become safer because it ends with "talk to your doctor" or "I am not a medical professional," if the sentence before that already told the patient what to do.
That matches what happens in clinic. The dangerous output is rarely a wild invention. It is a specific, personalized, confident next step, attached to a chart that looks like the patient in front of you.
Two further points are easy to skip and should not be.
A conversation can start as information and end as an order. Risk has to be judged across the whole exchange, not on the first reply.
For triage, both errors count. Missing an emergency delays care. Sending everyone with chest discomfort to the emergency department creates its own harm: anxiety, unnecessary testing, crowded departments, and patients who stop believing the tool the next time.
Benchmarks are not the same thing as practice
The proposed exam has four families of tasks: safety, clinical proficiency, generalizability, and, for agentic systems, extra competencies. Safety includes recognizing when to escalate, staying inside the intended use, and telling the user when the device is unsure. Proficiency includes knowledge, information gathering, numbers, and whether a human can actually understand the answer. Generalizability includes whether the result holds up across runs and across subgroups.
Then comes clinical confirmation, because a benchmark can be thorough and still miss the workflow. CDRH lists a ladder, and a randomized trial is the top rung, not the only rung: retrospective review of real records, shadow mode in a live clinic where the output is recorded and hidden from the care team, standardized patients, independent clinician review, and, when the risk requires it, a prospective study.
Shadow mode is the idea I hope physicians defend. A device should have to be right in the workflow before its words are allowed to touch the chart.The paper is also honest about a trap. Public benchmarks get memorized, saturated, and detached from the patients a hospital actually sees. A sponsor-built benchmark can be tuned until the product passes. If an LLM is the "independent" grader, that grader still has to be independent of the company that built the device and of the company that built the model. Independence is a structure, not a slogan.
Where this meets the evidence problem
In July I wrote on LinkedIn about a related failure. Clinical AI had learned to cite papers and still treat them as interchangeable. A large randomized trial and a 47-patient chart review could sit under the same confident paragraph. Physicians do not need a faster footnote. They need to see how much the evidence can be trusted for the question in front of them. That is the GRADE distinction: certainty belongs to a specific question and a specific outcome, not to a paper, and not to the tone of the model.
The FDA paper is the device-regulation version of that argument. Question 1 even asks whether traceability of an output back to primary sources belongs in the risk framework. It should.
A device can score well on a knowledge benchmark and still do the thing I described in July. It can retrieve real citations, arrange them fluently, and hide the fact that the body of evidence for this patient is low-certainty. Accuracy of the sentence and certainty of the evidence are different properties. A regulatory exam that measures only the first will license tools that sound finished.
I have written separately about how to catch that class of error before it reaches a patient, in AI hallucinations in medicine. The regulatory question is what the manufacturer has to prove before the physician is left to catch it alone.
What is worth saying before October 19
If you submit a comment, or if your hospital's digital health group submits one, these are the points I would not leave to device lawyers alone.
-
Put traceability on the risk axis. An output the clinician cannot check against a primary source is a higher-consequence output, even when the words are only "informational." Measurement and signal tools already get this treatment in the paper, because the user cannot see the basis of the number. Cited prose deserves the same suspicion when the citations are decorative.
-
Treat calibration as a safety behavior, not a style choice. The paper's own benchmark list includes uncertainty communication and clinical deferral. A product that is equally sure on a settled question and on a thin evidence base has failed a safety item. Saying so in public comments keeps that item from being negotiated down to a tooltip.
-
Require shadow deployment before action-directing uses. Retrospective vignettes are a start. They are not a substitute for watching the device on real patients while a human is still the only one allowed to act.
-
Name who the comparator is. "As good as a clinician" can mean a specialist panel, a typical generalist, the human-AI pair, or the device working alone. Those are four different products. The intended user should determine the comparator. A tool sold to generalists should not be graded only against subspecialists who would never have missed the trap.
-
Make third-party model changes the manufacturer's problem on a clock. If the foundation model updates on a Tuesday, the device that sits on top of it is a different device on Wednesday. A comment that asks for detection, re-benchmarking, and a hold on action-directing functions until that re-check is done is a patient-safety comment, not an anti-innovation comment.
-
Keep manufacturer accountability intact. The paper invites hospitals, societies, and payers into postmarket monitoring. That help is useful. It becomes dangerous if "the health system will watch it" replaces a duty the sponsor can be held to.
Comments from people who have had to decide whether to trust an AI sentence, with a patient waiting, are the comments this docket is thin on. The manufacturers will write. The physicians should too.
Frequently Asked Questions
Q: Is this new FDA guidance I have to follow?
A: No. The August 18, 2026 paper is a discussion document. CDRH says it is not draft guidance, not final guidance, and not a statement of what evidence a future submission must contain. The actionable date is October 19, 2026, when comments on docket FDA-2026-N-7874 are due.
Q: Does the FDA regulate ChatGPT itself?
A: The paper draws the line the agency already uses for software. FDA does not regulate generative AI as a technology. It regulates products that meet the definition of a medical device. A general chatbot is not automatically a device. A product that uses generative AI to direct or take a clinical action may be.
Q: What is the "competency" approach in plain language?
A: Test the finished product the way we test a trainee. First, a structured exam of knowledge, safety behavior, communication, and whether the result holds up across patients. Then confirmation in realistic clinical use, which might be a chart review, a silent shadow period, standardized patients, independent clinicians, or a prospective study. The rigor is supposed to rise with risk.
Q: How does this relate to evidence grading and CertaintyGrade?
A: A device can be factually fluent and still rest on weak evidence. GRADE rates certainty for a defined question and outcome. CertaintyGrade applies that standard to a question you specify, instead of treating every cited paper as equally trustworthy. The FDA is asking how to test device competence. Evidence grading asks a second question the exam can miss: even if the words are competent, how strong is the evidence a clinician is being asked to act on?
Q: What should a physician do with a generative AI tool already in the workflow?
A: Use it as a draft. Check one primary source before you act. Notice when the tool stops informing you and starts instructing you. If your hospital is deploying a generative feature that writes orders, patient advice, or triage, ask whether it has been watched in shadow mode on your patients, and what happens when the underlying model changes.
Dr. Adarsh K. Gupta, DO, MS is a clinical AI strategist and the author of The Science of Habits. He is a Cochrane Associate Editor and the founder of CertaintyGrade, a GRADE-based platform that rates the certainty of clinical evidence for a specific question. Follow him at adarshgupta.com.
