The U.S. Food and Drug Administration (FDA) has awarded a $1.29 million research contract to Cognita Imaging, a move intended to revolutionize the way generative artificial intelligence is validated and regulated in the medical imaging sector. This 18-month initiative, which officially commenced on June 22, focuses on developing a scalable methodology for evaluating AI-generated radiology reports. At the heart of the project is a novel "jury" system, where multiple large language models (LLMs) are utilized to assess the accuracy and safety of reports generated by other AI systems, potentially reducing the heavy burden currently placed on human radiologists during the software validation process.
As generative AI continues to permeate the healthcare landscape, the FDA is seeking more robust, automated ways to ensure that these complex tools are safe for clinical use. The contract signifies a shift from the evaluation of narrow, task-specific AI to broad, generative models capable of interpreting multifaceted medical data. Cognita Imaging, an emerging leader in the field of vision-language models, will lead the research in collaboration with the FDA’s Center for Devices and Radiological Health (CDRH) regulatory science team.
The Evolution of AI in Medical Imaging
To understand the significance of the Cognita contract, one must examine the current state of AI in radiology. Historically, the FDA has authorized hundreds of AI-enabled medical devices, the vast majority of which are "discriminative" or "narrow" AI. These tools are designed to perform a single, specific task: identifying a pulmonary embolism in a CT scan, detecting a fracture in an X-ray, or flagging a potential intracranial hemorrhage.
The evaluation of these narrow tools is relatively straightforward. Performance is measured against a "ground truth" established by human experts, typically using metrics like sensitivity, specificity, and the area under the receiver operating characteristic (ROC) curve. Because the output is binary or limited—either the pathology is there or it isn’t—human verification is manageable.
However, the industry is now moving toward generative AI and Vision-Language Models (VLMs). These tools do not simply flag a finding; they draft comprehensive narrative reports that describe everything from the size of a heart to the presence of subtle nodules in the lungs. Akshay Chaudhari, co-founder of Cognita Imaging and an associate professor of radiology and biomedical data science at Stanford University, notes that these models are "open-ended in nature." A single chest X-ray report might contain dozens of distinct clinical findings, observations, and recommendations. Evaluating such complex output across thousands of cases becomes an insurmountable task for human experts alone, creating a bottleneck in regulatory approval and clinical deployment.
The "LLM Jury" Methodology
The core of the Cognita-FDA project is the development of a framework that uses a "jury" of LLMs to act as an automated oversight committee. This approach addresses the scalability problem by tasking high-performing language models with the objective of reviewing the textual outputs of radiology AI.
In this framework, several different LLMs will analyze the same AI-generated report. By comparing the interpretations of multiple models, the system can identify discrepancies or uncertainties. If the "jury" reaches a consensus that a report is accurate, it may require less human intervention. Conversely, if the models disagree or flag a high-risk statement, the case is escalated to a human radiologist for final adjudication.
This methodology aims to solve two of the most persistent issues in generative AI: hallucinations and omissions. Hallucinations occur when an AI model generates information that sounds clinically plausible but is factually incorrect—for instance, describing a "normal gallbladder" in a patient who has previously undergone a cholecystectomy. Omissions, perhaps more dangerous in a clinical setting, occur when a model fails to mention a critical finding that was present in the imaging data.
The project will apply this "jury" framework to a massive dataset of 1 million patient exams sourced from a large U.S. cohort. This scale is unprecedented for AI validation studies and will allow researchers to examine model performance across a vast array of variables, including different patient demographics, various types of imaging equipment (from different manufacturers), diverse care settings (from rural clinics to urban teaching hospitals), and rare diseases that are often underrepresented in smaller datasets.
Timeline and Project Deliverables
The 18-month contract is structured to provide the FDA with actionable data and tools to inform future regulatory guidance. Following the initial development and testing of the framework, Cognita will provide several key deliverables:
- A Comprehensive Findings Report: A detailed analysis of the project’s results, including the success rate of the LLM jury in identifying errors compared to human radiologists.
- Software Code: The underlying code used to build and implement the LLM jury system, allowing the FDA to audit the methodology.
- Regulatory Guidance Templates: Documentation that helps define how other developers might use similar "AI-on-AI" evaluation techniques to validate their own generative models.
- Comparative Analysis: A study comparing the effectiveness of large-scale validation cohorts (like the 1-million-exam set) against smaller, traditional validation sets.
- Discrepancy Studies: A deep dive into cases where the AI jury and human radiologists disagreed, providing insights into the "reasoning" of generative models.
Strategic Context: Radiology Partners and the Workforce Crisis
The timing of this contract is particularly relevant given the current state of the radiology workforce. Cognita Imaging, founded in 2024, was acquired only a year later by Radiology Partners, the largest radiology practice in the United States. This acquisition provides Cognita with a unique "living laboratory." Radiology Partners employs thousands of radiologists who provide services across a wide spectrum of clinical environments.
The feedback loop between Cognita’s developers and Radiology Partners’ practicing clinicians is a critical component of the project. By understanding the real-world challenges faced by radiologists—such as high volume, burnout, and the need for rapid turnaround times—Cognita can tailor its AI tools to be truly assistive rather than disruptive.
Furthermore, the U.S. is currently facing a significant shortage of radiologists. According to data from the Association of American Medical Colleges (AAMC) and the American College of Radiology (ACR), the demand for imaging services is outstripping the supply of trained specialists. This gap leads to longer wait times for patients and increased pressure on physicians. Generative AI is seen as a potential "force multiplier" that could automate the more tedious aspects of reporting, allowing radiologists to focus on complex diagnostic decision-making. However, for these tools to be adopted, trust in their accuracy is paramount—a trust that the FDA-Cognita project aims to build through rigorous validation.
Breakthrough Device Designation and Future Prospects
While the research contract focuses on the evaluation of AI, Cognita is also moving forward with its own generative products. The company recently received the FDA’s "Breakthrough Device Designation" for a vision-language model designed specifically to interpret chest X-rays and generate preliminary reports.
The Breakthrough Device program is intended to speed up the development and review of medical devices that provide more effective treatment or diagnosis of life-threatening or irreversibly debilitating diseases. While this designation does not mean the device is authorized for sale yet, it grants Cognita prioritized interactive communication with FDA experts during the premarket review phase.
Chaudhari has clarified that the $1.29 million research contract and the Breakthrough Device project are separate but complementary. The lessons learned from the research contract—specifically how to quantify hallucination rates and clinical significance—will likely inform the clinical study plan for Cognita’s own commercial products.
"There’s still work to be done in developing the exact study plan and actually implementing the study," Chaudhari said in a statement. "But we’re confident that through these very nice partnerships that the FDA fosters that there is a path forward, because there is a real clinical need for these tools."
Broader Impact on the Regulatory Landscape
The FDA’s investment in Cognita reflects a broader agency-wide focus on the governance of artificial intelligence. In early 2024, the FDA released a discussion paper and held public workshops regarding the use of generative AI in medical devices. The agency is grappling with how to regulate "adaptive" models—those that continue to learn and change after they are deployed in the field.
Traditional regulatory frameworks are designed for static devices. If a scalpel or a heart valve is approved, its design does not change overnight. Generative AI, however, is dynamic. The "LLM jury" approach could provide a mechanism for continuous monitoring, where AI systems are constantly being checked against a standard of truth as they process new data.
If successful, the Cognita project could serve as a blueprint for other medical specialties beyond radiology, such as pathology, dermatology, and cardiology, where visual data is translated into narrative reports. By establishing a standardized, automated way to evaluate these models, the FDA can ensure that the rapid pace of AI innovation does not outstrip the agency’s ability to protect public health.
As the 18-month project progresses, the medical and technology communities will be watching closely. The results will likely dictate whether "AI-on-AI" oversight becomes a standard requirement for the next generation of medical software, potentially ushering in an era where the collaboration between human expertise and machine intelligence is validated by the machines themselves.

