,

How Big Should Your Sample Be? Sample Size Rules for Postgraduate Nursing and Health Research (2026)

There is no sector-wide dataset recording what sample sizes UK postgraduate nursing dissertations actually achieve. Anyone who quotes you a national average has invented it. What genuinely exists — and what an examiner will expect you to cite — is a body of published methodological conventions, each with a named source, a year, and a clearly stated boundary beyond which it should not be applied.

This article sets out those conventions in citable form. It is deliberately not a power calculator. The point is that “how many participants do I need?” has no answer until you have specified a design, an outcome and an effect you would consider clinically meaningful — and the honest work is in that specification, not in the arithmetic that follows it.

What is and is not evidenced

Three distinctions worth holding before you read any figure below.

First, methodological conventions are not empirical distributions. Cohen’s effect size benchmarks tell you what counts as a medium effect by convention; they do not tell you what effect sizes occur in nursing research. Cohen himself warned against applying them mechanically where field-specific evidence exists.

Second, rules of thumb encode assumptions. The widely-cited “ten events per variable” rule for logistic regression was derived under specific simulation conditions and has been challenged in later work. Citing it is fine; citing it as a law is not.

Third, a target sample size and an achieved sample size are different numbers, and dissertations are examined on how honestly the gap between them is handled. Recruitment through NHS services routinely runs below projection.

Conventions for quantitative designs

Convention What it states Source Year
Effect size benchmarks Cohen’s d: 0.2 small, 0.5 medium, 0.8 large; correlation r: 0.1, 0.3, 0.5 Cohen, Statistical Power Analysis for the Behavioral Sciences 1988
Conventional power target 80% power at a two-sided alpha of 0.05 as the default minimum Cohen 1988
Events per variable, logistic regression Approximately 10 events per candidate predictor Peduzzi et al., Journal of Clinical Epidemiology 1996
Multiple regression rule of thumb N ≥ 50 + 8m for testing overall model fit, where m is the number of predictors Green, Multivariate Behavioral Research 1991
Prediction model sample size Criteria-based calculation replacing events-per-variable rules Riley et al., Statistics in Medicine 2019–2020
Measurement property studies Samples of at least 100 rated as adequate for most psychometric properties COSMIN methodology 2018

The practical route through this table: identify your primary outcome and its analysis, find the convention that governs that analysis, and state the assumption you are making about the effect you expect. A power calculation reported without its assumed effect size is uninterpretable, and examiners ask about it.

Conventions for pilot and feasibility studies

Many postgraduate nursing dissertations are pilot or feasibility studies, either by design or because recruitment made them so. This is a legitimate design with its own literature, and it is not a consolation prize.

Convention What it states Source Year
Rule of 12 Around 12 participants per group as a general pilot rule of thumb Julious, Pharmaceutical Statistics 2005
Minimum for variance estimation At least 30 participants to estimate a parameter with reasonable precision Browne, Statistics in Medicine 1995
Effect-size-dependent pilot sizing Pilot size should scale to the effect size anticipated in the main trial Whitehead et al., Statistical Methods in Medical Research 2016
Pilot trial sizing review Recommendations vary widely across the literature; no single accepted figure Teare et al., Trials 2014

The critical point, and the one most often missed: a pilot study is not powered to detect an effect, so it should not report a significance test of effectiveness as a primary finding. Its outcomes are recruitment rate, retention, protocol adherence, acceptability and variance estimates for the main study. Reporting a non-significant p-value from a pilot as evidence of no effect is a substantive error, not a stylistic one.

Conventions for qualitative designs

Qualitative sample size is where the weakest justifications appear, usually as a single unreferenced sentence claiming saturation. The literature here is better than most students realise.

Convention What it states Source Year
Saturation in homogeneous samples Most themes emerged within 12 interviews; basic elements within 6 Guest, Bunce & Johnson, Field Methods 2006
Information power Required N depends on study aim, sample specificity, theory use, dialogue quality and analysis strategy Malterud, Siersma & Guassora, Qualitative Health Research 2016
Saturation is design-dependent Saturation means different things in grounded theory, content analysis and thematic analysis Braun & Clarke, Qualitative Research in Sport, Exercise and Health 2021

Two cautions. Guest and colleagues’ figure of 12 is regularly cited as though it were universal; their own conditions were a homogeneous sample and a relatively narrow objective, and heterogeneous samples need more. And Braun and Clarke have argued directly that saturation is a poor fit for reflexive thematic analysis, since meaning generation does not simply stop. If you are using that method, the information power framework is the safer justification.

If your design involves interviews, budget for transcription as a resource question with its own ethics dimension — the constraints are set out in the comparison of transcription tools for research interviews.

How to write the justification paragraph

Whatever your design, the section your examiner reads should contain four elements. State the primary outcome and the analysis it will receive. State the effect you would consider meaningful, and where that figure came from — a prior study, a published minimal clinically important difference, or a stated convention. State the parameters used, including power, alpha and any assumed attrition. Then state the resulting target, and the software or method used to derive it.

A defensible quantitative example reads roughly as follows: the primary outcome was change in a validated fatigue score; a difference of 4 points was taken as the minimal clinically important difference, based on a named prior study; assuming a standard deviation of 9 points from that study, 80 per cent power and a two-sided alpha of 0.05, 82 participants per group were required; inflating for 15 per cent attrition gave a recruitment target of 190. Every number in that sentence is traceable to a source or a stated assumption.

A defensible qualitative example names the framework rather than asserting saturation: a sample of 15 to 20 was planned on information power grounds, given a narrow aim, a specific sample of advanced practice nurses in one specialty, and an analysis using established theory, with recruitment to continue until the analytic team judged the material sufficient for the stated aim.

The calculation itself is usually trivial in G*Power or in R; the reasoning is what carries marks. Which environment you use for it, and whether the output reaches your thesis cleanly, is covered in the comparison of statistical software for postgraduate dissertations.

When the sample you achieve is smaller than the sample you planned

This is common enough that examiners expect to see it handled well rather than concealed. The strong response has three parts.

Report the recruitment pathway honestly, with numbers: how many were approached, how many were eligible, how many consented, how many completed. Reframe the study to match what you have, most often as a feasibility study or an exploratory analysis. Then state the consequence for interpretation in plain terms — an underpowered study that finds no significant difference has not shown that no difference exists, and saying so explicitly is a mark of methodological maturity rather than a weakness.

Discuss the reframing with your supervisor early, and check whether an ethics amendment is required if the design changes materially. The general principle of writing limitations as analytic judgements rather than apologies is set out in the guide to writing a discussion chapter.

Sample size in a systematic review

If your dissertation is a review rather than a primary study, sample size still matters — but it is a property of the studies you include, not of your own recruitment. A synthesis of seventeen small underpowered trials supports weaker conclusions than one of three adequately powered ones, and saying so is part of the quality appraisal rather than an aside. The procedure is set out in writing a PRISMA systematic review chapter.

Where a portfolio thesis combines a review with an empirical study, the two sample size discussions should speak to each other: the review is often where your effect size assumption legitimately comes from. That structural relationship is one of the things examiners probe, as described in the anatomy of a professional doctorate portfolio.

Get the justification written while the reasoning is fresh

Sample size justifications are almost always written twice — once badly in the proposal, and once properly months later when nobody remembers where the assumed standard deviation came from. Tesify keeps the source, the assumption and the draft paragraph attached to each other in one workspace, so the methods section you submit still cites the study your calculation was actually built on.

Start your dissertation with Tesify

Frequently asked questions

What is the average sample size for a masters nursing dissertation?

No national dataset records this, so any stated average is an estimate without a source. What examiners assess is not the size but the justification: whether the target followed from a stated design, effect and set of assumptions, and whether shortfalls were reported honestly.

Is 12 interviews enough for a qualitative dissertation?

It can be. The figure comes from Guest, Bunce and Johnson (2006), whose conditions were a homogeneous sample and a narrow objective. Heterogeneous samples or broader aims require more, and for reflexive thematic analysis an information power justification is generally stronger than a saturation claim.

Do you need a power calculation for a pilot study?

Not in the conventional sense, because a pilot is not powered to detect effectiveness. You justify the size against the pilot’s own objectives — estimating recruitment rate, retention and variance — citing conventions such as Julious (2005) or Whitehead et al. (2016).

What power and alpha should you use?

Eighty per cent power at a two-sided alpha of 0.05 is the conventional minimum, following Cohen (1988). Ninety per cent power is increasingly expected for definitive trials. State whichever you use and why, rather than leaving it implicit.

Where do you find an effect size to assume?

In order of preference: a published minimal clinically important difference for your outcome measure, a prior study in a comparable population, a meta-analysis, and only then a conventional benchmark. Assuming a large effect purely because it yields a feasible sample size is the failure mode examiners look for.

Is the ten events per variable rule still accepted?

It remains widely cited from Peduzzi et al. (1996) but has been challenged as too simple. For prediction models, the criteria-based approach of Riley and colleagues is now preferred. Citing the older rule is acceptable if you acknowledge the critique.

Can you change your sample size after ethics approval?

Usually yes, but a material change typically requires an amendment to your approval rather than a note in the thesis. Check with your sponsor and research office before recruiting beyond or substantially below an approved target.

Does a small sample mean the dissertation will fail?

No. Dissertations fail on unjustified claims, not on small numbers. A small study reported accurately, reframed appropriately and interpreted within its limits is examined more favourably than a small study whose conclusions outrun its data.

Should you report a post hoc power analysis?

Generally no. Observed power calculated from your own non-significant result is a direct function of the p-value and adds no information, a point made repeatedly in the statistical literature. Report confidence intervals instead, which show the range of effects your data remain compatible with.