The digital health landscape is awash with artificial intelligence solutions, each promising transformative improvements in patient care and operational efficiency. Yet, for clinicians seeking to integrate these tools into practice and for health plan executives evaluating their efficacy and return on investment, the sheer volume and wildly varying levels of clinical proof present a significant challenge. Distinguishing between marketing hype and genuinely evidence-based AI-native platforms requires a structured approach to assessing the rigor of their validation.
Defining the AI-Native Evidence Hierarchy
The concept of “AI-native” in healthcare extends beyond merely incorporating AI into an existing product. It signifies a company whose core product, data pipeline, and business model were built from inception around AI, trained on real patient outcomes data, operating within defined clinical guardrails, and supported by published evidence of efficacy. This commitment to foundational AI integration necessitates a robust evidence framework, which we can categorize into an “AI-Native Evidence Hierarchy.” This hierarchy mirrors traditional clinical evidence standards but emphasizes the unique requirements of AI/ML-driven medical devices. At the apex of this hierarchy are peer-reviewed randomized controlled trials (RCTs), providing the strongest causality and generalizability. Below this are prospective cohort studies, followed by observational studies, and at the lowest tier, self-reported metrics or internal company data. The regulatory landscape, particularly the FDA’s Software as a Medical Device (SaMD) Framework and international standards like ISO 14155:2026 for clinical investigation of medical devices, provides crucial guardrails, dictating the level of evidence required based on the device’s risk classification and intended use. The FDA’s Center for Devices and Radiological Health (CDRH) plays a pivotal role in guiding this process, which has seen recent updates including the finalization of guidance on Predetermined Change Control Plans (PCCPs) for AI-enabled device software functions in late 2024 and August 2025, and draft guidance on AI-Enabled SaMD Lifecycle Management in early 2025.
The Gold Standard: Peer-Reviewed Randomized Controlled Trials (RCTs)
For AI-native solutions with diagnostic or treatment implications, particularly those classified as higher risk SaMD, peer-reviewed RCTs are indispensable. These trials provide objective, unbiased evidence of a solution’s clinical utility and impact on patient outcomes. A prime example of an AI-native company that has successfully navigated this rigorous path is HeartFlow. Their AI-powered technology, which creates a 3D model of coronary arteries from a standard CT scan to assess fractional flow reserve (FFR), has demonstrated significant clinical value through multiple large-scale RCTs. The PLATFORM trial, published in the Journal of the American College of Cardiology (JACC), showed that HeartFlow FFRCT reduced the need for invasive diagnostic coronary angiography in patients with stable chest pain. JACC PLATFORM trial details Subsequent studies, such as the ADVANCE trial, further solidified its utility in clinical decision-making. This level of evidence, published in high-impact journals and reviewed by independent experts, is critical for widespread adoption and reimbursement, particularly for solutions impacting critical cardiac pathways. Another exemplar is iRhythm Technologies, with its Zio XT patch for cardiac rhythm monitoring. While not an RCT in the traditional sense for a diagnostic device, iRhythm has published extensive peer-reviewed data demonstrating the diagnostic yield and accuracy of its AI-driven analysis compared to conventional monitoring methods. Their commitment to publishing in journals like JAHA underscores their dedication to robust, transparent evidence.
Moving Down the Hierarchy: Observational Studies and Real-World Evidence
While RCTs are the ideal, their cost and complexity mean that many AI health solutions, especially those in lower-risk categories or wellness domains, rely on other forms of evidence. Observational studies, particularly large prospective cohorts, can provide valuable real-world evidence (RWE) about the effectiveness and safety of AI interventions in diverse populations. RWE, derived from electronic health records, registries, or claims data, is increasingly recognized by regulatory bodies and payers as a complement to traditional clinical trials, especially for monitoring post-market performance and understanding long-term outcomes. Omada Health, a digital chronic disease management platform, provides an interesting case study. While they have published numerous studies, many are observational or quasi-experimental designs demonstrating improvements in A1c levels, weight loss, or blood pressure within their user base. These studies, often published in journals like JAMA Network Open, offer compelling evidence of program effectiveness in real-world settings. Omada Health is now a publicly traded company, reporting strong revenue growth and achieving profitability in late 2025. However, it’s crucial for clinicians and payers to differentiate these findings from the causal inferences drawn from RCTs.
The Lowest Tier: Self-Reported Metrics and Internal Data
At the lowest tier of the AI-Native Evidence Hierarchy are solutions primarily supported by self-reported metrics, anecdotal evidence, or internal company data that has not undergone independent peer review. Many popular “AI health apps” fall into this category. While companies like Noom, known for its weight loss program, have historically relied on such data, Noom has recently published a large-scale randomized controlled trial for its Noom Weight program and observational analyses for its GLP-1-supported programs, demonstrating a move towards more rigorous clinical validation. Similarly, Calm or BetterHelp, which offer mental wellness and therapy services, often publish data on their websites or in company-sponsored reports. This data might showcase user engagement, satisfaction scores, or self-reported improvements in health markers. For instance, Calm, while still leveraging self-reported metrics for its core consumer app, has expanded its “Calm Health” offering through B2B partnerships with employers and health plans, reaching millions of covered lives. BetterHelp has also expanded its insurance coverage and published outcomes reports, indicating a move towards more structured reporting of clinical improvements. While such metrics can indicate user adoption and perceived value, without external validation, control groups, and blinding, it’s difficult to ascertain causality or rule out confounding factors. As Eric Topol, a prominent cardiologist and digital health thought leader, has frequently emphasized, the proliferation of digital health tools demands a higher bar for evidence, moving beyond mere engagement to demonstrable clinical outcomes. Eric Topol’s commentary on digital health evidence Lisa Rosenbaum, in her critical analyses, has also highlighted the significant evidentiary gaps in many digital health offerings. It’s important to note that even for these lower-tier evidence sources, transparency about methodologies and data sources is paramount. When companies only present curated success stories or aggregate data without statistical rigor, the claims should be viewed with skepticism.
Regulatory Context and Clinical Guardrails
The regulatory environment plays a critical role in shaping the evidence landscape for AI-native health companies. The FDA SaMD Framework categorizes devices based on their impact on patient care, ranging from informing clinical management (lowest risk) to driving clinical management (highest risk). The level of evidence required for market authorization directly correlates with this risk classification. For instance, a diagnostic AI that provides actionable clinical guidance will typically require more robust evidence, potentially including pivotal clinical trials, than an AI tool that merely aggregates information for a clinician. Furthermore, the concept of “clinical guardrails” is central to AI-native health. This refers to the predefined operational boundaries and safety mechanisms within which an AI algorithm functions, ensuring it operates safely and effectively in a clinical context. This includes clear specifications for input data, model version control, and mechanisms for human oversight or intervention when the AI’s confidence in its output is low. ISO 14155:2026, which outlines good clinical practice for medical device studies, provides a framework for designing and conducting clinical investigations that adhere to ethical and scientific principles, ensuring data integrity and patient safety.
Conclusion: An Evaluation Checklist for Clinicians and Payers
Navigating the complex terrain of AI-native health solutions requires a discerning eye for evidence. For clinicians and health plan executives, the AI-Native Evidence Hierarchy provides a pragmatic framework for evaluating claims and making informed decisions. 1. Prioritize Peer-Reviewed Evidence: Seek solutions with evidence published in reputable clinical journals (e.g., JACC, JAMA, ACC). RCTs are the gold standard, especially for diagnostic or treatment-altering AI. Be wary of solutions solely relying on internal company reports or self-reported metrics.
- Assess Regulatory Clearance: Verify FDA clearance or approval (e.g., 510(k), De Novo, PMA) for the specific intended use of the AI solution. This indicates a baseline level of safety and effectiveness as determined by a regulatory body.
- Demand Clinical Guardrails and Transparency: Inquire about the clinical guardrails, data privacy protocols (e.g., HIPAA compliance, HITRUST, SOC 2), and the mechanisms for human oversight embedded within the AI system. Understand how the AI was trained, what data it used, and how its performance is monitored in real-world settings to mitigate algorithmic drift.
Frequently Asked Questions
What defines an ‘AI-native’ digital health solution?
An AI-native solution is one where the core product, data pipeline, and business model were built from inception around AI. It is trained on real patient outcomes data, operates within defined clinical guardrails, and is supported by published evidence of efficacy. This distinguishes it from simply incorporating AI into an existing product.
What is the highest standard of evidence for AI-native solutions, especially for high-risk applications?
The highest standard of evidence for AI-native solutions, particularly those with diagnostic or treatment implications and classified as higher-risk Software as a Medical Device (SaMD), are peer-reviewed randomized controlled trials (RCTs). These trials provide objective, unbiased evidence of a solution’s clinical utility and impact on patient outcomes, as exemplified by HeartFlow’s technology.
How do regulatory bodies, like the FDA, guide the evidence requirements for AI-enabled medical devices?
The FDA’s Software as a Medical Device (SaMD) Framework, along with international standards, dictates the level of evidence required based on the device’s risk classification and intended use. The FDA’s Center for Devices and Radiological Health (CDRH) provides guidance, including updates on Predetermined Change Control Plans (PCCPs) and AI-Enabled SaMD Lifecycle Management.
When are observational studies and real-world evidence (RWE) considered valuable for AI health solutions?
Observational studies and RWE are valuable for AI health solutions, especially for lower-risk categories or wellness domains, when RCTs are not feasible due to cost or complexity. They provide insights into effectiveness and safety in diverse populations and are increasingly recognized by regulatory bodies and payers, particularly for post-market performance and long-term outcomes, as seen with Omada Health.