Defining Disease

This programme supported researchers to identify and define heart and circulatory conditions in health records, making research more reliable and reproducible. 

Scientific leadership: Professor Spiros Denaxas

About this work

Cardiovascular research using health data relies on robust and reusable, computational definitions that faithfully represent clinical concepts, like a diagnosis of a specific health condition – these are known as phenotype definitions.

Working with clinicians, data scientists and researchers, we developed and shared reusable cardiovascular disease definitions, helping researchers use and interpret health data. We also developed tools and guidance to make sure phenotype definitions are developed and used as effectively as possible in cardiovascular research.

What this programme delivered

Established best practices and standards

We worked with researchers and clinicians to define best practices for creating, storing, and sharing phenotyping algorithms using electronic health record (EHR) data. Our white paper outlines key recommendations to ensure that phenotyping algorithms are findable, accessible, interoperable, and reusable (FAIR), meeting the needs of the cardiovascular research community.  

Read our full report and FAIR recommendations and guidance for researchers on sharing phenotype definitions.

Made disease definitions available for researchers

We worked with experts to develop validated phenotyping algorithms, many of which have emerged from research projects within CVD-COVID-UK. These definitions are openly available in the BHF Data Science Centre Phenotype Library, which includes over 300 phenotype definitions, covering:  

  • Core cardiovascular diseases  
  • Medications  
  • Comorbidities (e.g., COVID-19)  
  • Risk factors  

Access the BHF Data Science Centre collection in the Library here.

Developed frameworks and tools for phenotype creation

We established a framework for building and applying phenotyping definitions, ensuring consistency across research studies. This approach was successfully used in collaboration with our clinical trials work in the SCORE-CVD project, which defined clinical trial outcomes in cardiovascular research. 

We also developed a code list comparison tool to help researchers:   

  • Compare different phenotyping definitions 
  • Assess code variations between datasets  
  • Select the most suitable definition for their study

Assessed phenotyping definitions and datasets

We systematically evaluated the accuracy and completeness of cardiovascular event recording and outcomes in different datasets. Our initial analysis focused on the Sentinel Stroke National Audit Programme (SSNAP) dataset, comparing stroke event data with electronic health records.  

This research highlighted important variations in stroke event recording, demonstrating how linked data can improve stroke measurement and healthcare quality assessment at a lower cost.

Read the news story about this work.

You can also read the full paper on BMJ Open.

Generated population-wide insights on cardiovascular disease

We worked on defining and providing information on cardiovascular diseases within the Disease Atlas, an ambitious project leveraging health data from 56 million people to generate new insights into diseases.  

This work aimed to change the way we think about and research diseases by:  

  • Providing comparative insights into cardiovascular health across different populations  
  • Informing policy and healthcare practices  
  • Unlocking new opportunities to improve patient outcomes  

Enhanced international research collaboration

We made it easier for researchers to compare cardiovascular studies across countries by:  

  • Collaborating with global leaders, such as Vanderbilt/eMERGE/All of Us  
  • Standardising cardiovascular phenotype definitions between UK Biobank and Germany’s NAKO study