Truveta brand logo mark in teal on a black background, featuring stacked chevron shapes forming the Truveta symbol.

ISPE 2026: A new frontier in case validation for Guillain-Barré syndrome (GBS)

by | Aug 30, 2026

Authors: Joanne Wu ⊕,Pfizer, Inc, New York, NY, Sampada Gandi, MPH ⊕, Pfizer, Inc, New York, NY, Jared Kearn ⊕, Truveta, Inc, Bellevue, WA, Maria Maddalena Lino  ⊕,  Pfizer, Inc, New York, NY,   Joe Engeda ⊕, Truveta, Inc, Bellevue, WA, Michelle R Iannacone Pfizer, Inc, New York, NY, Rebecca Naumann ⊕, Truveta, Inc, Bellevue, WA,  Younus Muhammad ⊕, Pfizer, Inc, New York, NY, Heather Rubino ⊕, Pfizer, Inc, New York, NY, Priyam Shah ⊕, Truveta, Inc, Bellevue, WA, Sumati Ramsinghani ⊕,  Truveta, Inc, Bellevue, WA,  Andy Pangilinan ⊕,Truveta, Inc, Bellevue, WA, Puja Rao ⊕,Truveta, Inc, Bellevue, WA,  Mao Hu ⊕, Truveta, Inc, Bellevue, WA, Jeff Ratto ⊕,  Truveta, Inc, Bellevue, WA, Esther Kim⊕ Truveta, Inc, Bellevue, WA, 

Using Unstructured EHR Data to Validate Guillain-Barré Syndrome Cases
  • Among 64 randomly sampled patients identified with a structured Guillain-Barré syndrome (GBS) diagnosis code, LLM-extracted information from unstructured clinical notes supported classification of 45% as confirmed GBS and 30% as probable GBS. 
  • One in four patients identified using structured GBS diagnosis codes were ultimately classified as unconfirmed GBS or as having confirmed or probable chronic inflammatory demyelinating polyneuropathy (CIDP), highlighting the potential for misclassification when relying on diagnosis codes alone. 
  • LLM-assisted extraction of physician evaluations, laboratory and diagnostic testing, and longitudinal clinical information may offer a scalable approach to improving GBS case validation in large real-world datasets. 

    This report summarizes our poster presented at ISPE 2026, titled A New Frontier in Case Validation for Guillain-Barré Syndrome (GBS): Harnessing Unstructured Data from Truveta Electronic Health Records. 

    Guillain-Barré syndrome (GBS) is a rare neurological disorder and an important safety endpoint in vaccine safety research. Large real-world data sources can help researchers study rare safety outcomes such as GBS across broad populations. However, identifying true cases remains challenging. Structured electronic health record (EHR) diagnosis codes are commonly used to identify potential GBS cases, but diagnosis codes alone can be prone to misclassification. Traditional medical record review can provide greater clinical detail but is often time- and resource-intensive. 

    Unstructured clinical notes provide another source of information for case validation. These notes can capture physician assessments, diagnostic uncertainty, laboratory and testing results, and changes in a patient’s diagnosis over time. Recent advances in artificial intelligence (AI) make it possible to extract these features from large volumes of clinical text at a scale that would be difficult to achieve through manual review alone. The Truveta Language Model (TLM), for example, can extract clinically relevant information from unstructured EHR notes for use in real-world research (1). 

    In this study, we evaluated the feasibility and potential value of an AI-enabled approach to classify GBS cases using information extracted from unstructured EHR notes. Specifically, we assessed whether physician evaluations and results from cerebrospinal fluid (CSF), nerve conduction studies (NCS), and electromyography (EM) could help distinguish confirmed and probable GBS from unconfirmed cases and chronic inflammatory demyelinating polyneuropathy (CIDP). 

    Methods

    Using a subset of Truveta Data, we identified patients with a structured GBS diagnosis code, based on ICD-10-CM or SNOMED codes, between 2017 and 2024. Overall, 39,816 patients had a structured GBS diagnosis code during the study period. 

    We then used TLM to extract features related to GBS from patients’ unstructured clinical notes. Extracted information included physician evaluations of GBS and CIDP as well as results from CSF testing, NCS, and EM. Across the study population, features were extracted from 1,082,371 unstructured notes. 

    For an in-depth feasibility assessment, 64 cases were randomly sampled for human review, representing 4,917 unstructured notes. Reviewers used a prespecified decision tree incorporating the last physician evaluation of GBS during the follow-up period, abnormal diagnostic testing, and the last physician evaluation of CIDP to assign a final case classification. Cases were categorized as confirmed GBS, probable GBS, confirmed CIDP, probable CIDP, or unconfirmed. 

    The classification approach reflects the longitudinal nature of diagnosing GBS. Rather than treating the presence of a diagnosis code as definitive evidence of disease, the approach incorporates clinical information documented over time to determine whether subsequent evidence supports, modifies, or does not confirm the initial diagnosis. 

    Results 

    AI enabled evaluation of more than 1 million clinical notes 

    Among 39,816 patients with a structured GBS diagnosis code between 2017 and 2024, TLM extracted GBS-related features from more than 1 million unstructured notes. The in-depth human-reviewed sample included 64 randomly selected cases and 4,917 corresponding clinical notes. 

    This approach combines the scale of structured EHR data with information contained in clinical narratives. The analysis first identified potential cases using structured diagnosis codes and then evaluated information extracted from unstructured notes to better characterize each patient’s clinical course. 

    Clinical notes captured evolving diagnostic journeys 

    The longitudinal patient examples shown in Figure 2 of the poster illustrate why information beyond a diagnosis code can be important for case validation. Physician assessments and diagnostic test results accumulated over time, providing evidence that could either support GBS or point toward another diagnosis such as CIDP. One example ultimately reflected confirmed CIDP, while another was classified as confirmed GBS. 

    These examples also demonstrate that case status may evolve during the days following the first structured GBS diagnosis. Evaluating a patient’s clinical journey, rather than relying on a single point in time, provides additional context for determining case certainty. 

    75% of sampled cases were classified as confirmed or probable GBS 

    Of the 64 randomly sampled cases, 29 (45%) were classified as confirmed GBS and 19 (30%) as probable GBS. Together, these groups accounted for 75% of cases initially identified through structured GBS diagnosis codes. 

    The remaining quarter were not ultimately classified as confirmed or probable GBS. Six cases (9%) were classified as confirmed CIDP, five (8%) as probable CIDP, and five (8%) as unconfirmed. In total, 25% of patients identified through structured GBS codes were classified as either unconfirmed or as having CIDP. 

    Figure 3 of the poster shows the decision tree used to reach these classifications. Among patients whose last physician evaluation confirmed GBS, abnormal EM, NCS, or CSF findings strengthened case certainty when there was no evidence of CIDP. Patients without abnormal testing could still be classified as probable GBS based on physician evaluation, while evidence supporting CIDP shifted the final assignment toward confirmed or probable CIDP. Patients whose final physician assessment remained provisional or unconfirmed required additional evidence to support a GBS classification. 

    Figure 4 summarizes the resulting classifications and highlights a central finding of the study: a structured GBS diagnosis code alone did not consistently represent the final clinical classification. 

    Discussion 

    This study demonstrates the potential value of combining structured EHR data with AI-extracted information from unstructured clinical notes to improve GBS case classification. This distinction is particularly important for vaccine safety research, where GBS is an important safety endpoint and accurate outcome identification is essential. In our analysis, one in four patients identified through structured GBS diagnosis codes were ultimately classified as unconfirmed or as having confirmed or probable CIDP. Misclassification of outcomes in observational studies can affect estimates and interpretation of safety signals, underscoring the importance of robust case-validation approaches. 

    A key strength of this approach is its ability to use information that is difficult to capture through structured fields alone. Truveta Data provides access to longitudinal laboratory results and physician evaluations, while TLM enables extraction of relevant information from unstructured EHR notes. More than 1 million notes were evaluated in this study, demonstrating the potential scalability of the approach. Following validation against traditional medical record review using an established case definition such as the Brighton Collaboration criteria, this approach has the potential to support more rapid, accurate, and scalable identification of GBS cases in large EHR databases (2). 

    There are important limitations to consider. EHR data may be incomplete when patients receive care outside Truveta member health systems. Missing information may also occur for other systematic reasons, creating the potential for informed-presence or ascertainment bias. Validation against traditional medical record review using established GBS case definitions will be important before applying the method more broadly. 

    Overall, this study demonstrates how AI-assisted extraction of unstructured EHR data can add clinical context to structured diagnosis codes. By combining physician assessments, diagnostic testing, and longitudinal patient journeys, this approach may help researchers distinguish true GBS cases from uncertain or alternative diagnoses. With further validation, AI-enabled case review could provide a scalable method for improving outcome identification in vaccine safety research and other real-world evidence studies where accurate case classification is critical. 

    Data are constantly changing and updating. These findings are consistent with data analyzed for the ISPE 2026 study.

    Citations

    1. Truveta, Whitepaper: Truveta Language Model (2026). 
    1. J. J. Sejvar, K. S. Kohl, J. Gidudu, et al., Guillain–Barré syndrome and Fisher syndrome: Case definitions and guidelines for collection, analysis, and presentation of immunization safety data. Vaccine 29, 599–612 (2011). 

    Disclosures 

    This study was funded by Pfizer, Inc. Pfizer authors are employees of Pfizer, Inc. and may hold shares and/or stock options in the company. Truveta authors are employees of Truveta, Inc. Truveta, Inc. received funding from Pfizer, Inc. to assist with this study.