Free tool · Sample size · Continuous outcomeUpdated October 9, 2026

Two-means sample size calculator for continuous outcomes

Blood pressure, HbA1c, a pain score, a walk distance: when the primary endpoint is a mean in two groups, enter the difference and SD and get n per group, with the small-sample t-correction, unequal allocation and dropout built in.

  • t-correction, not just the z formula
  • Allocation ratio and dropout
  • What-if table for a wrong SD

Free sandbox · No credit card · 21 CFR Part 11 aligned

Demo: systolic BP, difference 5 mmHg, two-sided alpha 0.05, 80% power

Evaluable per group

92

Enrol per group (10% dropout)

103

Total enrolment

206

SD 10 mmHg64/150

64 per group

SD 12 mmHg92/150

92 per group

SD 15 mmHg143/150

143 per group

The SD is squared in the formula: a 25% higher SD needs about 56% more participants.

What this calculator does

  • For a superiority comparison of two group means, the sample size per group is n = (1 + 1/k) σ² (z₁₋α/2 + z₁₋β)² / δ², where k is the allocation ratio (n₂ / n₁), σ the common SD and δ the difference you want to detect.
  • The z formula ignores that σ is estimated. The t-distribution option finds the smallest n whose non-central t power reaches your target, which adds 1 or 2 per group in a typical trial and more in a small one.
  • Worked example: a 5 mmHg difference with an SD of 12 mmHg, two-sided alpha 0.05 and 80% power needs 92 per group, or 103 per group to enrol if 10% will drop out.
  • The inputs that move the answer most are the SD (squared) and the difference (squared). Alpha and power matter less, and they are conventions, not findings.
  • Adjusting for the baseline value in an ANCOVA can cut the sample size by a factor of (1 − ρ²); the calculator gives the plain t-test number, which is the safe starting point.

Free tool

Calculate a two-means sample size

Enter the difference you want to detect (the smallest clinically important one, not the one you hope for) and the SD. Choose a one- or two-sided test, the allocation ratio and the method. It runs in your browser; nothing is saved or sent.

%
Test
Method
Group 1
92
Group 2
92
Total
184
evaluable

Enrol 103 + 103 = 206 with 10% dropout. Standardised effect d = 0.42. Power at these sizes: 80.3%.

n₁ = (1 + 1/k) σ² (z + z)² / δ² = 2.00 × 12² × 2.802² / 5² = 90.4 (normal formula); the t-distribution search gives 92

If the assumptions are off (total evaluable)
SD 20% higher than assumed264
True difference 20% smaller286
Both410

Planning aid and reference only, not validated software. Assumes two independent groups, a common SD, a continuous outcome that is roughly normal, and the analysis planned as a two-sample t-test. The normal formula ignores that σ is estimated, so it is slightly low in small trials; the t-distribution option uses the non-central t distribution (as G*Power does). Dropout inflation divides by (1 − dropout). Confirm the final number with a statistician.

The calculation

The formula, the t-correction and a worked example

The planned analysis is a two-sample comparison of means: a t-test, or a regression that tests the same contrast. The null hypothesis is that the true means are equal, the alternative is that they differ by δ, and you want probability 1 − β (power) of rejecting the null at level α if the alternative is true. With equal groups, the normal-theory answer is n = 2σ²(z₁₋α/2 + z₁₋β)² / δ² per group (Chow, Shao and Wang). For a two-sided alpha of 0.05 and 80% power, the two z values are 1.960 and 0.842, and their sum squared is 7.85, so n = 15.7 / d² where d = δ/σ is the standardised effect size.

Take systolic blood pressure. You want to detect a 5 mmHg difference between a lifestyle programme and usual care, and earlier studies suggest an SD of 12 mmHg. Then d = 0.417 and the z formula gives 2 x 144 x 7.85 / 25 = 90.4, so 91 per group. The t-distribution method returns 92, because the t-test critical value is a little larger than 1.96 when the SD must be estimated from the data. With d = 0.5 the same correction moves the answer from 63 to 64 per group, which is the figure G*Power reports. The correction shrinks as n grows and matters most below about 50 per group.

Power, alpha and sidedness are choices you state in the protocol, not outputs. At 90% power the same example needs 123 per group. A one-sided test at 0.05 needs 72, but regulators and journals expect a two-sided test unless a one-sided hypothesis is justified and the alpha is halved to 0.025, which brings you back to the two-sided number. Tightening alpha to 0.01 raises the requirement to 137. Dropout is a separate step: enrolment is the evaluable number divided by (1 minus the dropout fraction), so 92 evaluable at 10% dropout is 103 to enrol, and the dropout-adjusted sample size calculator explains why dividing is correct and adding 10% is not.

Unequal allocation

Set the allocation ratio to 2 for a 2:1 design and the calculator returns n₁ and n₂ = 2 n₁. In the example, 69 and 138 give 80% power, 207 in total instead of 184. Unequal allocation costs total participants (the loss is small for 3:2 and larger beyond 2:1) and is justified by something else: more safety exposure to the new treatment, or lower cost of the cheaper arm. If you use it, say why in the protocol.

Sensitivity

Evaluable participants per group by difference and SD

Two-sided alpha 0.05, 80% power, 1:1, t-distribution method, rounded up.

Difference to detectSD 10SD 12SD 15
3 units176253394
4 units100143222
5 units6492143
6 units4564100
8 units263757

Illustrative values from the calculator. Halving the difference roughly quadruples n; a 50% larger SD multiplies it by about 2.2.

Assumptions

What you are assuming when you press calculate

Every sample size is a statement about a world you have not seen yet. Writing the assumptions down, with a source for each number, is what lets an ethics committee, a statistician and a regulator judge whether the study is realistic. The tool states them under the result; this is the longer version.

The assumptions behind the number

The calculator assumes all of the following:

  • Two independent groups, one measurement per participant, analysed with a two-sample t-test or equivalent.
  • The outcome is continuous and roughly symmetric. A heavy-tailed or skewed variable (many biomarkers, length of stay) is usually log-transformed first, and then δ and σ must be on the log scale.
  • A common SD in both groups. If the groups plausibly differ a lot, use a Welch-type calculation or simulation.
  • The SD is known without error. It never is, which is why you stress it with the what-if table.
  • Complete data on every evaluable participant, with dropout handled by inflating enrolment, not by modelling.
Sample size inputs to agree before the protocol is final

Primary endpoint

Change in SBP, wk 12

Result

92 per group

  • Difference: smallest clinically important5 mmHg, cited MCID
  • SD from a comparable population12 mmHg, SD of the change
  • Alpha two-sided, power0.05, 80%
  • Baseline adjustment (ANCOVA) planned?Needs the baseline-outcome correlation
  • SD sensitivity rangeAt SD 15 the study needs 143 per group
Each line becomes a sentence in the sample size section of the protocol and SAP.

Common mistakes

Where two-means sample sizes go wrong

Most errors are not arithmetic. They are inputs that sound reasonable and are not, and the standard deviation is the usual culprit.

Using the wrong SD

If the primary endpoint is the change from baseline, you need the SD of the change, not the SD of the final value or of the baseline. These can differ by a factor of two. If the primary endpoint is the final value adjusted for baseline, the relevant variance is the residual one. A published abstract rarely tells you which SD it reports, so look for the paper that does. Where the SD comes from a small pilot, plan on a conservative limit rather than the observed value; the pilot study sample size calculator shows how much that costs.

Choosing the difference to fit the budget

The target difference should be the smallest effect that would change practice, often framed as a minimal clinically important difference. Working backwards from the 80 participants you can afford to whatever difference they happen to detect, and then calling that difference clinically important, is a well-known way to produce an underpowered trial that is later described as inconclusive. If you can only afford a smaller study, say so and consider a different design, or use the power analysis calculator to state honestly what the study can detect.

Forgetting the analysis model

Adjusting for the baseline measurement with ANCOVA multiplies the t-test sample size by (1 − ρ²), where ρ is the correlation between baseline and outcome (Borm et al., 2007). With ρ = 0.5 that turns 92 into about 69 per group. The saving is real, but only if the SAP names the adjusted analysis as primary and ρ comes from data, not hope. Conversely, if the design is clustered, or the same participants contribute repeated measurements, a two-sample formula is wrong and the answer needs a design effect.

Turn the sample size into a study you can test

Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.

Build your study free

From the number to the protocol

How this feeds the protocol, the SAP and the study build

ICH E9 asks that the protocol state the sample size, how it was determined, the primary variable, the assumed values and the method, with the clinical and statistical reasoning behind each. A sentence that satisfies this reads: "A total of 184 evaluable participants (92 per group) provides 80% power to detect a between-group difference of 5 mmHg in change in systolic blood pressure at week 12, assuming a common SD of 12 mmHg, using a two-sided t-test at the 5% level. Allowing for 10% dropout, 206 participants will be enrolled." Put the same numbers in the statistical analysis plan, and put the dropout assumption in the monitoring plan so someone is watching it. The how to write a clinical trial protocol guide places this section in context, and the sample size calculation glossary entry defines the terms.

In Capture the number turns into build decisions. The primary endpoint becomes a numeric eCRF field with a unit and a range edit check, so an implausible blood pressure raises a query while the site is still with the participant. If the endpoint is a calculated value, such as a mean of three readings, a calculated field shows it read-only during entry. Visit windows around the week-12 anchor define which measurements count, randomization allocates participants in the ratio you planned, and the monitoring dashboard shows enrolment against the target so a shortfall is seen early. At the end, CSV or Excel exports come with a data dictionary, or SDTM datasets as SAS XPT files with Define-XML, and the field-level audit trail records old value, new value, user, time and reason for every change. For superiority comparisons of proportions use the two-proportions sample size calculator, and for a margin-based question use the non-inferiority sample size calculator. The broader clinical trial sample size calculator covers the quick binary case.

Statistical analysis plan (demo)
SAP_v1.0_DEMO.docx3/5 complete
  1. 4

    Analysis sets

    Complete
  2. 6

    Primary endpoint: change in SBP at week 12

    t-test, two-sided 5%

    Complete
  3. 9

    Sample size justification

    92 per group, SD 12, difference 5

    Complete
  4. 9.2

    Dropout allowance and monitoring

    10% assumed; review at 50% enrolled

    Draft
  5. 10

    Sensitivity: SD 15 scenario

    To do
The sample size section is written once and referenced by the protocol, the SAP and the data monitoring plan.

Before you lock it

Two-means sample size checklist

Endpoint and time point defined

Final value or change from baseline, and the analysis model that goes with it.

Difference justified

Smallest clinically important difference, with a source, not the number the budget allows.

SD matches the endpoint

SD of the change if the endpoint is a change; from a similar population and measurement method.

Sensitivity checked

Re-run with an SD 20% higher and a difference 20% smaller and decide if the trial survives.

Sidedness and alpha stated

Two-sided 0.05 unless a one-sided test is justified in advance.

Dropout converted to enrolment

Enrolment = evaluable / (1 - dropout), rounded up.

FAQ

Questions teams ask before they switch

Something not covered here? Ask us directly.

What is the sample size formula for comparing two means?

For equal groups, n per group = 2σ²(z₁₋α/2 + z₁₋β)² / δ², where σ is the common SD, δ the difference to detect, and the z values are normal quantiles for alpha and power. For unequal allocation n₁ = (1 + 1/k)σ²(z₁₋α/2 + z₁₋β)² / δ² and n₂ = k n₁. The t-distribution version finds the smallest n whose non-central t power reaches the target and is slightly larger.

What is the difference between the normal formula and the t-distribution option?

The normal formula treats the SD as known. The t option accounts for estimating it from the trial data. With a standardised effect of 0.5, alpha 0.05 and 80% power the answers are 63 and 64 per group; the gap shrinks as n rises.

How do I choose the standard deviation?

Use the SD of the same endpoint (final value or change) from a comparable population and measurement method, ideally pooled from several sources. Then test a higher value. A small pilot gives a noisy SD, so plan on an upper confidence limit instead of the point estimate.

Should I use a one-sided or a two-sided test?

Two-sided at 0.05 is the default for confirmatory trials. A one-sided test is acceptable only when a difference in one direction has been ruled out in advance and the one-sided alpha is set at 0.025, which gives the same sample size as the two-sided 0.05 test.

How does adjusting for baseline change the sample size?

An ANCOVA that adjusts for the baseline value needs about (1 - ρ²) times the t-test sample size, where ρ is the correlation between baseline and outcome. With ρ = 0.5 that is a 25% saving, but the adjusted analysis must be the pre-specified primary analysis.

Does this work for a binary or time-to-event endpoint?

No. Use the two-proportions calculator for binary endpoints and the survival sample size calculator for time-to-event endpoints. This tool is for a difference in means of a continuous outcome.

How do I add dropout?

Divide the evaluable number by (1 - expected dropout) and round up. The calculator does this when you enter a dropout percentage; the dropout-adjusted calculator page explains the reasoning.

Is this calculator validated?

No. It is a free planning aid, checked against published and G*Power reference values. A statistician should confirm the final sample size and the analysis plan before a protocol is submitted.

Build the endpoint your sample size assumes

Numeric eCRF fields with range checks, visit windows and randomization in a free sandbox. No credit card. You pay only when you go live.

Build your study free