PEDSnet provides a harmonized, multi-institutional pediatric clinical dataset that enables cross-site research within a unified governance framework. However, traditional EHR analyses rely on diagnosis-based cohort definitions or tabular feature representations that do not fully capture the relational structure of longitudinal pediatric health data. We propose to leverage the Data Distillery Knowledge Graph (DDKG) to model PEDSnet subject-level data as structured networks and generate graph embeddings that enable scalable, interpretable patient similarity analysis. Our central hypothesis is that graph-based patient embeddings derived from harmonized PEDSnet data will enable more meaningful clustering of pediatric phenotypes, treatments, and outcomes than traditional feature-based approaches.
Aim 1: Integrate harmonized PEDSnet subject-level data into the Data Distillery Knowledge Graph framework. We will map structured PEDSnet Common Data Model domains, including conditions, procedures, medications, laboratory results, encounters, and demographics, into the DDKG schema. Each subject’s longitudinal clinical history will be represented as a patient-linked subgraph within a shared biomedical context derived from UMLS-based ontologies. This aim establishes a reproducible pipeline for transforming harmonized pediatric EHR data into a graph-native analytic substrate.
Aim 2: Generate and evaluate patient-level graph embeddings to model phenotype similarity. We will apply graph embedding techniques to derive compact numerical representations of each subject’s longitudinal clinical network. These embeddings will enable rapid subject-to-subject similarity scoring, clustering, and subtype discovery. Using a computational IBD phenotype as a proof-of-concept, we will evaluate whether embedding-based clustering identifies clinically coherent subgroups independent of predefined diagnostic categories.
Aim 3: Simulate cross-institutional embedding generation and assess robustness across site partitions. Although PEDSnet data are harmonized centrally, we will simulate cross-institutional participation by partitioning data according to actual and/or simulated site identifiers. We will generate site-specific embeddings and evaluate stability, generalizability, and concordance across partitions. This aim will provide empirical evidence for the feasibility of extending this framework to a fully federated FL-GNN architecture in future phases.
Successful completion of these Aims will demonstrate that biomedical knowledge graph–based patient similarity modeling is feasible, interpretable, and scalable within harmonized pediatric network data. These results will provide the methodological foundation for future deployment of federated graph learning across independent institutional infrastructures.
