Javascript must be enabled to continue!
Phecoder: semantic retrieval for auditing and expanding ICD-based phenotypes in EHR biobanks
View through CrossRef
Abstract
Background
Electronic health record (EHR)–based phenotyping underpins genome-wide association studies, yet current ICD-code phenotypes rely heavily on manually curated lists such as Phecodes. These definitions are labour-intensive to maintain, inherently subjective, and may omit clinically relevant diagnostic codes, reducing study power. Advances in text embedding models offer an opportunity to automate and standardize ICD-based phenotype construction.
Methods
We developed Phecoder, an ensemble of pre-trained text embedding models that rank ICD codes by similarity from free-text phenotype descriptions. Nine embedding models and multiple unsupervised ensemble rank-fusion methods were evaluated against 1,125 PhecodeX phenotypes. Retrieval performance was assessed using recall and average precision at top-100 (R@100, AP@100). Expert clinical review of six neuropsychiatric phenotypes was undertaken to identify relevant ICD codes absent from PhecodeX. Cohort sizes under these new definitions were compared with PhecodeX across sex and ancestry strata in the Million Veteran Program (MVP).
Findings
Among individual models, Qwen3-Embedding-4B achieved the highest median recall (R@100 = 0.86). Ensemble rank-fusion further improved R@100 by 3%, and median AP@100 by 8%. Expert review confirmed that Phecoder retrieved additional clinically relevant ICD codes beyond PhecodeX across all six neuropsychiatric case studies. Median potential case expansion increased by 200%, with 700% increases for bipolar disorder and 2000% increase for eating disorders.
Interpretation
Manually defining ICD phenotypes has been critiqued as subjective, potentially yielding overly restrictive definitions that miss relevant codes. To address this issue, Phecoder algorithmically identifies relevant codes for ICD-based phenotyping. Phecoder extracts relevant ICD codes to expand the potential case pool across different demographic groups. Phecoder is easily applicable to future ICD-code releases and across different ICD coding versions that are used in different countries. Taken together, Phecoder has the potential to improve reproducibility in EHR data research.
Funding
This research was supported by the Department of Veterans Affairs MVP (MVP-000, MVP-076 and MVP-096). The MVP is supported by the Office of Research and Development, Department of Veterans Affairs. The authors thank the MVP staff, researchers, and volunteers, who have contributed to MVP, and especially who previously served their country in the military and now generously agreed to enroll in the study (see
mvp.va.gov
for more information). The contents do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.
This study was supported by the Veterans Affairs Merit grants: BX006500 (to D.B.) and BX004189 (to P.R.). This work was supported by the National Institutes of Health (NIH): R01MH125246 (to P.R.), R01AG078657 (to G.V.), R01AG067025 (to P.R.), and U24AG087563 (to P.R.).
Research in context
Evidence before this study
Phecodes are curated groupings of ICD-9 and ICD-10 diagnostic codes designed to create clinically meaningful phenotypes for large-scale research using electronic health records. They provide a standardized, reproducible way to map tens of thousands of ICD codes into interpretable disease concepts, and they are widely used in genome-wide association studies, phenome-wide association studies, and disease prediction models. PhecodeX, released in 2023, is the most recent update to this framework. It restructures the original catalogue to make full use of ICD-10 granularity and introduces more than 1,700 new phenotypes across a broad range of clinical domains. Phecodes have become foundational tools for EHR-linked biobanks, enabling harmonized phenotyping across institutions and cohorts.
A major barrier to the continued development and updating of Phecodes is their reliance on slow, manual curation, which is inherently subjective. As a result, it is difficult to ensure complete capture of all clinically relevant ICD codes; particularly for conditions with diffuse presentations, heterogeneous coding practices, or evolving diagnostic criteria.
Thus, some studies using predefined ICD-based phenotypes have reported unexpectedly low sensitivity, reinforcing the concern that curated code lists may miss substantial numbers of true cases.
Added value of this study
This study introduces Phecoder, a semantic retrieval framework that streamlines ICD-based phenotyping by ranking codes according to their similarity to any free-text phenotype description. This approach replaces static, manually curated lists with a flexible workflow in which phenotype definitions can be rapidly generated, audited, and refined simply by adjusting the input text. Phecoder therefore supports the continuous development of more comprehensive and responsive phenotype mappings.
We provide the first systematic benchmark of text embedding models for ICD-level phenotyping, evaluating nine encoders and several unsupervised ensemble methods against 1,125 PhecodeX phenotypes. Leveraging existing curated mappings as a reference standard, we show that a score-level ensemble improves retrieval performance over individual models and achieves perfect median recall for mental health phenotypes. Phecoder also identifies clinically relevant ICD codes beyond PhecodeX.
Expert review of six neuropsychiatric case studies confirms the clinical relevance of these additional codes, and their incorporation exposes sizable untapped patient cohorts in the Million Veteran Program.
Implications of all the available evidence
Together with existing use of Phecodes in major biobanks, these findings show that curated phenotype systems remain indispensable but benefit from scalable, transparent tools that support their maintenance and evolution. Phecoder enables continuous auditing and expansion of ICD-based phenotype definitions, and improves cohort completeness across demographic groups. Consequently, we anticipate that Phecoder will facilitate improved reproducibility across demographic groups in scientific research.
Title: Phecoder: semantic retrieval for auditing and expanding ICD-based phenotypes in EHR biobanks
Description:
Abstract
Background
Electronic health record (EHR)–based phenotyping underpins genome-wide association studies, yet current ICD-code phenotypes rely heavily on manually curated lists such as Phecodes.
These definitions are labour-intensive to maintain, inherently subjective, and may omit clinically relevant diagnostic codes, reducing study power.
Advances in text embedding models offer an opportunity to automate and standardize ICD-based phenotype construction.
Methods
We developed Phecoder, an ensemble of pre-trained text embedding models that rank ICD codes by similarity from free-text phenotype descriptions.
Nine embedding models and multiple unsupervised ensemble rank-fusion methods were evaluated against 1,125 PhecodeX phenotypes.
Retrieval performance was assessed using recall and average precision at top-100 (R@100, AP@100).
Expert clinical review of six neuropsychiatric phenotypes was undertaken to identify relevant ICD codes absent from PhecodeX.
Cohort sizes under these new definitions were compared with PhecodeX across sex and ancestry strata in the Million Veteran Program (MVP).
Findings
Among individual models, Qwen3-Embedding-4B achieved the highest median recall (R@100 = 0.
86).
Ensemble rank-fusion further improved R@100 by 3%, and median AP@100 by 8%.
Expert review confirmed that Phecoder retrieved additional clinically relevant ICD codes beyond PhecodeX across all six neuropsychiatric case studies.
Median potential case expansion increased by 200%, with 700% increases for bipolar disorder and 2000% increase for eating disorders.
Interpretation
Manually defining ICD phenotypes has been critiqued as subjective, potentially yielding overly restrictive definitions that miss relevant codes.
To address this issue, Phecoder algorithmically identifies relevant codes for ICD-based phenotyping.
Phecoder extracts relevant ICD codes to expand the potential case pool across different demographic groups.
Phecoder is easily applicable to future ICD-code releases and across different ICD coding versions that are used in different countries.
Taken together, Phecoder has the potential to improve reproducibility in EHR data research.
Funding
This research was supported by the Department of Veterans Affairs MVP (MVP-000, MVP-076 and MVP-096).
The MVP is supported by the Office of Research and Development, Department of Veterans Affairs.
The authors thank the MVP staff, researchers, and volunteers, who have contributed to MVP, and especially who previously served their country in the military and now generously agreed to enroll in the study (see
mvp.
va.
gov
for more information).
The contents do not represent the views of the U.
S.
Department of Veterans Affairs or the United States Government.
This study was supported by the Veterans Affairs Merit grants: BX006500 (to D.
B.
) and BX004189 (to P.
R.
).
This work was supported by the National Institutes of Health (NIH): R01MH125246 (to P.
R.
), R01AG078657 (to G.
V.
), R01AG067025 (to P.
R.
), and U24AG087563 (to P.
R.
).
Research in context
Evidence before this study
Phecodes are curated groupings of ICD-9 and ICD-10 diagnostic codes designed to create clinically meaningful phenotypes for large-scale research using electronic health records.
They provide a standardized, reproducible way to map tens of thousands of ICD codes into interpretable disease concepts, and they are widely used in genome-wide association studies, phenome-wide association studies, and disease prediction models.
PhecodeX, released in 2023, is the most recent update to this framework.
It restructures the original catalogue to make full use of ICD-10 granularity and introduces more than 1,700 new phenotypes across a broad range of clinical domains.
Phecodes have become foundational tools for EHR-linked biobanks, enabling harmonized phenotyping across institutions and cohorts.
A major barrier to the continued development and updating of Phecodes is their reliance on slow, manual curation, which is inherently subjective.
As a result, it is difficult to ensure complete capture of all clinically relevant ICD codes; particularly for conditions with diffuse presentations, heterogeneous coding practices, or evolving diagnostic criteria.
Thus, some studies using predefined ICD-based phenotypes have reported unexpectedly low sensitivity, reinforcing the concern that curated code lists may miss substantial numbers of true cases.
Added value of this study
This study introduces Phecoder, a semantic retrieval framework that streamlines ICD-based phenotyping by ranking codes according to their similarity to any free-text phenotype description.
This approach replaces static, manually curated lists with a flexible workflow in which phenotype definitions can be rapidly generated, audited, and refined simply by adjusting the input text.
Phecoder therefore supports the continuous development of more comprehensive and responsive phenotype mappings.
We provide the first systematic benchmark of text embedding models for ICD-level phenotyping, evaluating nine encoders and several unsupervised ensemble methods against 1,125 PhecodeX phenotypes.
Leveraging existing curated mappings as a reference standard, we show that a score-level ensemble improves retrieval performance over individual models and achieves perfect median recall for mental health phenotypes.
Phecoder also identifies clinically relevant ICD codes beyond PhecodeX.
Expert review of six neuropsychiatric case studies confirms the clinical relevance of these additional codes, and their incorporation exposes sizable untapped patient cohorts in the Million Veteran Program.
Implications of all the available evidence
Together with existing use of Phecodes in major biobanks, these findings show that curated phenotype systems remain indispensable but benefit from scalable, transparent tools that support their maintenance and evolution.
Phecoder enables continuous auditing and expansion of ICD-based phenotype definitions, and improves cohort completeness across demographic groups.
Consequently, we anticipate that Phecoder will facilitate improved reproducibility across demographic groups in scientific research.
Related Results
The Effect of Clinical Knee Measurement in Children with Genu Varus
The Effect of Clinical Knee Measurement in Children with Genu Varus
Abstract
Introduction
Children with genu varus needs frequent assessment and follow up that may need several radiographies. This study investigates the effectiveness of the clinica...
Benefit of Implantable Cardioverter Defibrillator Use in Japanese Patients Based on Modified MADIT-ICD Benefit Score
Benefit of Implantable Cardioverter Defibrillator Use in Japanese Patients Based on Modified MADIT-ICD Benefit Score
Abstract
Aims
The MADIT-ICD benefit score is used to stratify the risk of life-threatening arrhythmia and non-arrhythmic ...
Arrythmic storm in patients with and without an implantable cardioverter defibrillator
Arrythmic storm in patients with and without an implantable cardioverter defibrillator
Abstract
Introduction
Available data on arrhythmic storm (AS) is frequently obtained from retrospective observational series of ...
Impact of Cardiac Resynchronization Therapy on Hospitalizations in the Resynchronization-Defibrillation for Ambulatory Heart Failure Trial
Impact of Cardiac Resynchronization Therapy on Hospitalizations in the Resynchronization-Defibrillation for Ambulatory Heart Failure Trial
Background—
This study reports the impact of cardiac resynchronization therapy (CRT) on hospitalizations in patients randomized to implantable cardioverter-defi...
Clinical outcomes of subcutaneous vs. transvenous implantable defibrillator therapy in a polymorbid patient cohort
Clinical outcomes of subcutaneous vs. transvenous implantable defibrillator therapy in a polymorbid patient cohort
BackgroundThe subcutaneous implantable cardioverter-defibrillator (S-ICD) has been designed to overcome lead-related complications and device endocarditis. Lacking the ability for ...
Electronic Health Record Acceptance by Physicians: A Single Hospital Experience in Daily Practice
Electronic Health Record Acceptance by Physicians: A Single Hospital Experience in Daily Practice
Introduction: Potential benefits of implementing an electronic health record (EHR) to increase the efficiency of health services and improve the quality of health care are often ob...
Biobanking for Genetic Diseases
Biobanking for Genetic Diseases
Abstract
Biobanks are bioresources of human samples linked to relevant personal and health data of the participants, which are collected, proces...
Registro Italiano Pacemaker e Defibrillatori - Bollettino Periodico 2019. Associazione Italiana di Aritmologia e Cardiostimolazione
Registro Italiano Pacemaker e Defibrillatori - Bollettino Periodico 2019. Associazione Italiana di Aritmologia e Cardiostimolazione
Razionale. Il Registro Italiano Pacemaker (PM) e Defibrillatori (ICD) dell’Associazione Italiana di Aritmologia e Cardiostimolazione (AIAC) raccoglie annualmente i principali dati ...

