Defining Disease

This programme supported researchers to identify and define heart and circulatory conditions in health records, making research more reliable and reproducible

Scientific leadership: Professor Spiros Denaxas

About this work

Cardiovascular research using large amounts of health data relies on robust and reusable, computational definitions that faithfully represent clinical concepts, like a diagnosis of a specific health condition – these are known as phenotype definitions. 

Working with clinicians, data scientists and researchers, we developed and shared reusable cardiovascular disease definitions, helping researchers use and interpret health data. We also developed tools and guidance to make sure phenotype definitions are developed and used as effectively as possible in cardiovascular research. 

What this programme delivered

Established best practices and standards 

We worked with researchers and clinicians to define best practices for creating, storing, and sharing phenotyping algorithms using electronic health record (EHR) data. Our white paper outlined key recommendations to ensure that phenotyping algorithms are findable, accessible, interoperable, and reusable, meeting the needs of the cardiovascular research community.  

We developed guidance for researchers on phenotyping. 

Made disease definitions available for researchers 

We worked with experts to develop validated phenotyping algorithms, many of which emerged from research projects within CVD-COVID-UK. These definitions are openly available in the BHF Data Science Centre Phenotype Library, which includes over 90 cardiovascular phenotypes, covering:  

  • Core cardiovascular diseases  
  • Medications  
  • Comorbidities (e.g., COVID-19)  
  • Risk factors  

Access the BHF Data Science Centre collection in the Library here.

Developed frameworks and tools for phenotype creation 

We established a framework for building and applying phenotyping algorithms, ensuring consistency across research studies. This approach was successfully used in collaboration with our clinical trials work in the SCORE-CVD project, which defined clinical trial outcomes in cardiovascular research. 

We also developed a code list comparison tool to help researchers:   

  • Compare different phenotyping algorithms   
  • Assess code variations between datasets  
  • Select the most suitable algorithm for their study

Assessed phenotyping algorithms and datasets

We systematically evaluated the accuracy and completeness of cardiovascular event recording and outcomes in different datasets. Our initial analysis focused on the Sentinel Stroke National Audit Programme (SSNAP) dataset, comparing stroke event data with electronic health records.  

This research (manuscript in preparation) highlighted important variations in stroke event recording, demonstrating how linked data can improve stroke measurement and healthcare quality assessment at a lower cost. 

Generated population-wide insights on cardiovascular diseases 

We worked to define and provide information on cardiovascular diseases within the Disease Atlas, an ambitious project leveraging health data from 56 million people to generate new insights into diseases.  

This work aimed to:  

  • Provide comparative insights into cardiovascular health across different populations  
  • Inform policy and healthcare practices  
  • Unlock new opportunities to improve patient outcomes  

Enhanced international research collaboration 

We made it easier for researchers to compare cardiovascular studies across countries by:  

  • Collaborating with global leaders, such as Vanderbilt/eMERGE/All of Us  
  • Standardising cardiovascular phenotype definitions between UK Biobank and Germany’s NAKO study