Tag: UK Biobank

  • Where Do UK Medical and Epidemiology PhD Students Get Their Data? UK Biobank, CPRD, HES, OpenSAFELY and the Cohorts in Numbers (2026)

    Where Do UK Medical and Epidemiology PhD Students Get Their Data? UK Biobank, CPRD, HES, OpenSAFELY and the Cohorts in Numbers (2026)

    A UK medical or epidemiology doctorate can be built entirely on data that already exists: UK Biobank holds around 500,000 participants and had 30,000 registered researchers by 2023, CPRD reaches a population of up to 64 million patients, and OpenSAFELY covers roughly 58 million patient records — but every one of these resources is gated by an application, and the application timeline, not the analysis, is the constraint that shapes a three-year candidature.

    What follows is the sourced inventory: what each resource holds, how large it is, how a doctoral researcher gets in, and what the figures imply for planning. Every number carries its source and year in the table; where a figure is a stated reach rather than an active count, the text says so.

    The datasets at a glance

    Resource Scale What it holds Access route for a doctoral researcher Source and year
    UK Biobank About 500,000 participants aged 40–69 at recruitment (2006–2010) Genotyping and sequencing, imaging, biomarkers, lifestyle questionnaires, linked hospital, cancer and death records; over 15 million biological samples Registration and approval of a health-related, public-interest project; results returned to the resource; analysis on the Research Analysis Platform UK Biobank, 2023
    CPRD (Clinical Practice Research Datalink) Reach of up to 64 million patients across the four UK nations Anonymised primary care records: consultations, diagnoses, prescriptions, laboratory results; linkages to HES, ONS deaths, the national cancer registry and MINAP Protocol approval and a data licence, usually through the supervisor’s institution CPRD, MHRA and NIHR, 2012 onwards
    Hospital Episode Statistics (HES) About 16 million admitted episodes a year Every episode of admitted patient care in England, including NHS-funded care by private providers; outpatient and emergency department datasets alongside Application to NHS England, which absorbed NHS Digital on 1 February 2023 NHS England, 2023
    OpenSAFELY About 58 million patient records Pseudonymised primary care records held in place; researchers submit code and receive only aggregated results Approved project with an accredited team; code run against the data inside the platform Bennett Institute, University of Oxford, 2020–2025
    SAIL Databank Population-scale linked data for Wales Anonymised health and administrative datasets inside a trusted research environment at Swansea University Project approval and analysis within the secure environment SAIL Databank, 2026
    ALSPAC (Children of the 90s) More than 14,000 pregnant women recruited April 1991 to December 1992 Birth cohort followed for three decades: the mothers, their children and now the grandchildren, with biological samples and linked records Proposal to the ALSPAC executive at the University of Bristol; data access fees apply University of Bristol, 2026
    Understanding Society 40,000 households, about 100,000 individuals, since January 2009 Annual household panel with health modules, nurse-collected biomarkers and genetic data Free through the UK Data Service; special licence for sensitive files ISER, University of Essex, and ESRC, 2026
    ONS Longitudinal Study A 1% sample of England and Wales: over 500,000 people at each census point, about one million members across 40 years Linked census records since 1971 with births, deaths and cancer registrations Application through CeLSIUS; analysis only in ONS secure settings in London, Newport and Titchfield Office for National Statistics, 2025
    UK Data Service Funded by UKRI through ESRC with £37.5 million to 2030 The national repository for social and economic survey data, including the health-relevant cohorts and panels Registration; safeguarded and controlled data by application UK Data Service, 2025

    UK Biobank: the deep-phenotype cohort

    UK Biobank recruited approximately 500,000 volunteers aged 40 to 69 between 2006 and 2010 and intends to follow them for at least 30 years. It holds more than 10,000 variables per participant, over 15 million biological samples and more than 30 petabytes of data; its genome sequencing, proteomic and imaging datasets are described by the resource as the largest in the world. Access opened to researchers in March 2012 on a single condition set: the project must be health-related and in the public interest, results must be published openly, and findings must be returned to the resource. By 2023, 30,000 researchers from more than 90 countries had registered, more than 9,000 peer-reviewed papers had used the data by November 2023, and the Research Analysis Platform launched in 2021 had over 5,000 users.

    For a doctoral project the implications are two. First, a UK Biobank analysis is almost never a solo application: the approved project is usually the supervisor’s, and the student is added as a named collaborator, which is the fastest route in. Second, the 9,000 papers are the competition. With that many analyses already published on the same 500,000 people, the doctoral contribution has to come from the question, the linkage or the method, not from the dataset itself. Check the published-papers register before committing a chapter to an outcome someone has already reported.

    UK Biobank’s own induction module for new researchers, covering what the resource holds and how access works.

    CPRD and HES: primary care and hospital records at national scale

    The Clinical Practice Research Datalink was launched on 29 March 2012, consolidating the older General Practice Research Database with the Health Research Support Service, and it is jointly funded by the Medicines and Healthcare products Regulatory Agency and the National Institute for Health and Care Research. Its stated reach is a population of potentially up to 64 million patients across England, Wales, Scotland and Northern Ireland — a reach figure, not a count of currently contributing patients, and the distinction matters when you write the methods chapter. The records cover consultations, coded diagnoses, prescribed drugs and laboratory data, and CPRD links them to Hospital Episode Statistics, ONS death registrations, the national cancer registry and the MINAP cardiovascular registry. The historic series shows how the resource grew: 543,100 patients across 57 practices in 1988, 1.2 million in 1990, 4.4 million active patients from 650 practices by 1994.

    Hospital Episode Statistics is the other half of the routinely collected picture: a record of every episode of admitted patient care delivered by the NHS in England, including care done under contract by private providers, with outpatient and emergency department datasets alongside. It runs to about 16 million episodes of care a year. NHS Digital, which collected and published HES, merged into NHS England on 1 February 2023, so applications now go to NHS England’s data access service. For a doctoral researcher the practical point is the same for both: the application is made by or through an institution, it takes months rather than weeks, and it needs a protocol that already states the cohort definition, the codes and the outcomes. Our guide to writing the governance section of a health methodology chapter covers the approvals layer that sits underneath these applications.

    A data access application and data sharing agreement for a doctoral epidemiology study laid out on a desk
    Every national dataset is reached through an application; the protocol it demands is most of the methods chapter written early.

    OpenSAFELY and SAIL: the trusted research environments

    OpenSAFELY inverts the usual model. Rather than releasing copies of raw data, it leaves the pseudonymised primary care records where they are already stored and takes the researcher’s code to the data, returning only aggregated results. The platform covers approximately 58 million patient records and is developed and operated by the Bennett Institute for Applied Data Science in the University of Oxford’s Nuffield Department of Primary Care Health Sciences, having begun as a collaboration between Oxford’s DataLab, the electronic health records group at the London School of Hygiene and Tropical Medicine and the record-system suppliers. Its first analysis, of COVID-19 mortality risk factors, was published in Nature in July 2020; the NHS extended its use to other major diseases by 2023; and in 2025 the Wellcome Trust awarded £17 million, £7 million for talking-therapy outcomes and £10 million for the infrastructure. For a doctoral researcher the model has a cost and a benefit: the analysis must be written as reproducible code before it touches data, and in exchange the whole pipeline is publishable and reviewable, which examiners in the field increasingly expect.

    Wales runs its own trusted research environment, the SAIL Databank at Swansea University, holding anonymised health and administrative datasets for the Welsh population that are analysed inside the secure environment after project approval. Scottish and Northern Irish routes exist on the same pattern. The common feature of every environment in this section is that data never leaves; the doctoral researcher travels to it, by code or by secure desktop, and plans the thesis timetable around approval.

    The cohorts and surveys: ALSPAC, Understanding Society and the ONS Longitudinal Study

    Three long-running studies give a medical doctorate a life-course or population dimension that routine records cannot. The Avon Longitudinal Study of Parents and Children, based at the University of Bristol and known as Children of the 90s, recruited more than 14,000 pregnant women between April 1991 and December 1992 and has followed the mothers, the children and now the next generation since, with biological samples throughout. Understanding Society, directed by the Institute for Social and Economic Research at the University of Essex and funded principally by the Economic and Social Research Council, has surveyed 40,000 households — about 100,000 individuals — annually since January 2009, and includes nurse-collected biomarkers and genetic data alongside the socio-economic panel. The ONS Longitudinal Study links census records for a 1% sample of England and Wales at every census since 1971 with births, deaths and cancer registrations: over 500,000 people at each point and about one million members across its 40-year history, accessible through CeLSIUS and analysed only in the secure settings at ONS offices in London, Newport and Titchfield.

    The gateway for the survey data is the UK Data Service, whose continuation to 2030 was secured by £37.5 million of UKRI funding through ESRC. Understanding Society is downloadable there after registration; the ALSPAC and ONS resources sit behind their own application processes, and ALSPAC charges data access fees that a studentship budget needs to anticipate. The question of how many participants a doctoral analysis needs from any of these is a design question, set out with its conventions in our guide to sample size in postgraduate health research.

    A doctoral researcher planning a data application timeline against the years of a UK PhD
    Application lead times decide the order of chapters: the data request goes in before the literature review is finished, not after.

    What the figures mean for a three-year candidature

    Read together, the numbers say three things. First, the resources are already crowded: 30,000 registered UK Biobank researchers and 9,000 papers mean that novelty lives in the question and the linkage. Second, every route in is institutional. None of the nine resources in the table is opened by a student acting alone; the supervisor’s existing approvals, the university’s data sharing agreements and the sponsor’s governance are the real credentials, so choose a supervisory team by the data it can already reach. Third, the lead time is the cost. Applications that need a protocol, an ethics opinion and a signed agreement routinely consume a year of a registration period whose limits are set out in our piece on how long a UK PhD actually takes, which is why experienced supervisors have the data application drafted in the first term and the literature review written while it is being processed.

    If the doctorate combines routine data with a systematic review, run the review while the application is pending: our guide to the PRISMA systematic review chapter is built for that sequencing. And once data arrives it arrives under conditions: most of the agreements above forbid taking record-level data outside the approved environment, which includes pasting it into any external service — see our note on unpublished thesis data and AI tools before drafting a results chapter anywhere but the approved machine. Tesify is used by doctoral researchers in this position to hold the protocol, the chapter plan and the writing, with only aggregated, publishable results ever leaving the safe haven; start with the free plan and keep the record-level data where the agreement says it must stay.

    Frequently asked questions

    Can a PhD student apply for UK Biobank data directly?

    Registration is open to researchers whose project is health-related and in the public interest, but in practice doctoral researchers are added as collaborators on a supervisor’s approved application. By 2023 the resource had 30,000 registered researchers from more than 90 countries.

    How many people are in UK Biobank?

    Approximately 500,000 volunteers aged 40 to 69 at recruitment, enrolled between 2006 and 2010, with over 15 million biological samples and more than 30 petabytes of data held, and follow-up planned for at least 30 years.

    How many patients does CPRD cover?

    CPRD states a reach of potentially up to 64 million patients across England, Wales, Scotland and Northern Ireland. That is a reach figure rather than a count of currently contributing patients, and the methods chapter should say which.

    What is the difference between CPRD and HES?

    CPRD is anonymised primary care data from general practice, jointly funded by the MHRA and NIHR and launched in 2012. HES records every episode of admitted patient care in NHS England, around 16 million a year, plus outpatient and emergency department datasets. CPRD can be linked to HES for approved projects.

    How does OpenSAFELY differ from downloading a dataset?

    Researchers never receive record-level data. The code is taken to the roughly 58 million pseudonymised records where they are stored, and only aggregated results come back. The platform is run by the Bennett Institute at the University of Oxford.

    Is ALSPAC data free for doctoral research?

    No. Access is by proposal to the study executive at the University of Bristol and data access fees apply, so the cost belongs in the studentship budget. The cohort began with more than 14,000 pregnant women recruited in 1991 and 1992.

    Where can I analyse the ONS Longitudinal Study?

    Only in ONS secure settings in London, Newport and Titchfield, after application through CeLSIUS. The study links census records since 1971 for a 1% sample of England and Wales, over 500,000 people at each census point.

    Which health datasets can I download from the UK Data Service?

    Understanding Society, with its 40,000 households and nurse-collected biomarkers, is available after registration, with some sensitive files under special licence. The service’s funding through ESRC runs to 2030.