Truveta brand logo mark in teal on a black background, featuring stacked chevron shapes forming the Truveta symbol.

From fragmented records to living evidence: Health system–governed, AI-driven, continuously updated real-world clinical data from Truveta

by | Aug 6, 2026

Authors: Hannah A Burkhardt, PhD Truveta, Inc, Bellevue, WA, Sarah J Blach, MHS  Truveta, Inc, Bellevue, WA, Srinivasa R Burugapalli, MS  Truveta, Inc, Bellevue, WA, Brianna M Cartwright, MS Truveta, Inc, Bellevue, WA, Ryley Martin, BS Truveta, Inc, Bellevue, WA, Michael Simonov, MD Truveta, Inc, Bellevue, WA, Sarah Stewart, MD Truveta, Inc, Bellevue, WA, Grace Turner, MS Truveta, Inc, Bellevue, WA, Angela L Winegar, PhD Truveta, Inc, Bellevue, WA, Nicholas Stucky, MD Truveta, Inc, Bellevue, WA

from fragmented records to living evidence, a peer-reviewed look at Truveta Data
  • Truveta Data links structured and unstructured EHR data with closed claims, mortality, and social drivers of health.
  • Data are ingested daily, normalized to standard clinical ontologies, de-identified, and monitored through a continuous quality-control framework.
  • AI and natural language processing transform clinical concepts from physician notes, imaging reports, and pathology narratives into standardized variables for observational research.
  • The health system–governed model creates an ongoing feedback loop to improve data quality, provenance, and research utility.

Real-world data (RWD) have historically suffered from fragmentation, delayed availability, variable data quality, and limited analytic utility. Truveta developed an AI-enabled data platform to address these longstanding challenges in using RWD for clinical research. This peer-reviewed paper, published in JAMIA Open, describes Truveta’s partnership model, platform design, data scale, and research applications.

Materials and Methods

Truveta de-identifies, aggregates, and harmonizes electronic health record (EHR) data for 140 million patients—1 in 3 Americans—from US health systems. The platform links structured and unstructured EHR content with closed claims, mortality, and social determinants of health. Data undergo daily ingestion, normalization to standard ontologies, and de-identification. Advanced artificial intelligence (AI), including natural language processing (NLP), extracts key clinical concepts from free-text notes, such as physician notes, imaging reports, and pathology narratives, transforming them into standardized variables suitable for large-scale observational research.

Results

Truveta Data comprise over 130 million de-identified patient records that are updated daily and represent diverse geographic regions, care settings, and patient populations in the US. The data have supported over 100 scientific publications to date; additionally, they support health system participants’ own research interests and patient care insights. Published studies have addressed treatment effectiveness, post-market device surveillance, COVID-19 vaccine safety, and health equity.

Discussion

Truveta addresses critical barriers that have hindered the realization of a learning health system. Unlike prior RWD initiatives limited by scope, latency, and vendor dependence, Truveta enables near-real-time, population-scale research grounded in rich clinical data. Its governance model ensures alignment with ethical, privacy, and regulatory standards. By rethinking the role of health systems from passive suppliers into active, incentivized partners, Truveta creates an unprecedented virtuous cycle of data quality and continuous improvement.

Conclusion

Truveta represents a paradigm shift in real-world evidence generation. By aligning incentives and AI driven harmonization, it provides a scalable, sustainable infrastructure for continuously updated clinical data. Providing large-scale, high-fidelity EHR data and daily updates, the platform accelerates clinical discovery, policy and public health decision-making, and improved patient outcomes.