A report from DATAMIND’s Roadbuilder 3 (Linking Physical and Mental Health in SMI) highlights the importance for coordinated efforts to enhance electronic health record resources in the UK, particularly to advance research on the physical health of individuals with severe mental illness (SMI). This means improving access to comprehensive primary care data nationwide and providing thorough documentation of available electronic health record data across the four nations, facilitating analyses across multiple datasets and data pooling. Additionally, there is a need for further exploration into the availability of data for individuals with SMI within longitudinal cohorts and the potential for linking these with electronic health records. An overarching strategy for the UK to expedite the investigation of mental health, along with the combined impact of mental and physical health, in routinely collected data, is advocated, potentially through initiatives like the Health Data Research UK Research Driver Programmes.
Challenges to Mental Health Research
The above complexities create a barrier for the field of mental health. Beyond the extraction of the required information itself, it fuels the proliferation of heterogeneous phenotypes with mixed quality and validity, hindering replication of results and slowing down progress across all areas of mental health research, from understanding of conditions to development of treatments and public health interventions. The lack of evidence-based, established phenotypes also limits the use of mental health variables by researchers from other specialties, preventing exploitation of existing synergies across medical specialties and slowing our understanding of mental-physical comorbidities.
Multi-disciplinary Approach to Solution
Viable solutions require multi-disciplinary teams including mental health clinicians, epidemiologists, patients and public involvement, database administrators, and data scientists with expertise in the processes generating and recording mental health data, and in data wrangling and analysis techniques. DATAMIND brings together these professionals and is uniquely situated to address the described problem, by generating novel solutions and making existing solutions discoverable and accessible. Our unique strategy is to create a one-stop repository of trusted, high-quality mental health phenotypes to halt the proliferation of mixed quality heterogeneous definitions and to concentrate all the knowledge in one place, facilitating discoverability and accessibility. We will achieve this by providing a set of mental health phenotypes co-created with expert clinicians and externally validated.
Expansion of Phenotypes
One arm of these phenotypes pertains to coded clinical data. So far, we have 39 mental health code-based phenotypes for mental health disorders and behaviours, some new and some existing and updated. These include definitions based on routinely collected EHRs from primary and secondary healthcare, as well as lists of Read and ICD-10 codes, and relevant algorithms. We aim to continue growing this list and expanding existing phenotypes by adding SNOMED definitions (also in link with UK’s Mental Health Mission) and run external validations using linked cohort and routine data. There are also plans to develop tools that can make direct use of these phenotypes for the extraction of mental health variables from EHRs, facilitating the job of researchers even further. DATAMIND’s phenotypes have been used in 30+publications to date.
Leveraging Natural Language Processing (NLP)
The second arm of phenotypes pertains to the automatic extraction of information directly from clinical notes through natural language processing (NLP) – a modality of machine learning that extracts information from natural, human-readable text (i.e., free text). NLP experts within DATAMIND have been instrumental in the setting up and development of the Clinical Record Interactive Search (CRIS) NLP Service at the South London and Maudsley since 2008, which includes the extracting and processing of anonymised information from free text in mental healthcare records. CRIS currently deploys over 100 NLP applications for symptoms, physical health conditions, contextual factors, intervention, outcomes and clinical status, and more, and has begun translating these into higher-level mental health phenotypes akin to those described above (for example, symptom profiles in psychosis). NLP meta-data from one or more of these apps has been used in nearly all 300+ publications from the Maudsley CRIS data resource.
Ensuring Discoverability and Accessibility
To achieve our vision, the above tools and services need to be discoverable and accessible. We have worked closely with the team developing HDR UK’s phenotype library, providing feedback to shape existing and new tools as well as the user interface. We have published our phenotypes in this library as a DATAMIND collection. In the period between 1st to 27th February 2024 alone, DATAMIND’s phenotypes have been visited 2,579 times and downloaded 450 times (note that these are crude numbers and likely an overestimation of true human activity). Starting with the CRIS catalogue, we are seeking to render NLP resources more findable through HDR UK’s Gateway, in active discussions with the Gateway team. We are also rendering them more accessible via our MH Text Analytics Cloud (MH-TAC) platform, which allows Trusts to upload text clinical data and receive NLP-derived variables, and which has been developed to established prototype stage. Finally, via an HDRUK Improving Transparency Award to the CRIS team (£6.6k; 2023-24) and with extensive patient consultation and input, the CRIS interface for researchers has been improved, including for NLP resources.
Impact on Mental Health Research
Overall, DATAMIND’s phenotypes and work with CRIS and MH-TAC have created new tools for the extraction of MH variables, and expanded existing ones. These tools have facilitated sharing of information across research teams, and increased efficiency in extracting mental health variables from EHRs. We have improved the discoverability and accessibility of such tools, providing a centralised resource of information on how to accurately extract mental health related information from the free text and clinical codes in EHRs. As a result, we expect to see even more researchers and papers making use of these resource.
