Sample size calculators start from power and return n. This one starts from the n you can actually enrol and returns the power for your expected effect, the smallest effect you could detect, and a power curve for means or proportions.
Free sandbox · No credit card · 21 CFR Part 11 aligned
Power at 50 per group
70 / 100
What this calculator does
Free tool
Choose a continuous or binary outcome, enter the participants per group and the effect you expect, and read off the power, the smallest effect detectable at your target power, and power at other sample sizes. It runs in your browser; nothing is saved or sent.
Planning aid and reference only, not validated software. Power is computed for the effect you enter, using the same models as the sample size calculators. Use it before the trial to see what a fixed n can detect. Computing "observed power" from a finished trial's own result adds no information beyond the p-value and is discouraged (Hoenig and Heisey, Am Stat 2001); report the confidence interval instead. Power for the t option uses the non-central t distribution.
The concept
Two errors are possible in a hypothesis test. A type I error is declaring an effect that is not there; alpha, usually 5%, caps its probability. A type II error is missing an effect that is there; its probability is β, and power is 1 − β. Eighty percent power means that if the true effect is exactly the assumed one, four trials in five would reach p < 0.05. It says nothing about what happens if the true effect is smaller, and it is a probability before the trial, not a property of the trial's result.
For a two-group comparison of means, power depends on the signal-to-noise ratio δ / (σ √(2/n)): the difference over its standard error. For the t-test the calculator uses the non-central t distribution, so the answers match G*Power and R's power.t.test. For proportions it uses the normal approximation with pooled or unpooled variance, as the two-proportions sample size calculator does. The relationships to remember are these: power rises with n and with the effect size, falls as alpha is tightened, and rises as the SD falls. Doubling n does not double power; it moves it along an S-shaped curve that flattens near 100%.
Cohen's d is the difference in means divided by the SD. By Cohen's widely quoted benchmarks, 0.2 is small, 0.5 medium and 0.8 large, but those are conventions for behavioural science, not clinical thresholds. A d of 0.2 in a mortality-adjacent endpoint can matter greatly, and a d of 0.8 in a self-rated score can be trivial. For two proportions, the tool reports Cohen's h = 2 asin √p₂ − 2 asin √p₁, an arcsine-transformed effect size with the same use. Anchor the effect in clinical units whenever you can, using something like the minimal clinically important difference.
Reference values
Two-sided alpha 0.05, 1:1, t-test (non-central t). Power in percent.
| Effect size (d) | 20 per group | 30 | 50 | 64 | 100 | 150 | 200 |
|---|---|---|---|---|---|---|---|
| 0.2 (small) | 9% | 12% | 17% | 20% | 29% | 41% | 51% |
| 0.3 | 15% | 21% | 32% | 39% | 56% | 74% | 85% |
| 0.5 (medium) | 34% | 48% | 70% | 80% | 94% | 99% | 100% |
| 0.8 (large) | 69% | 86% | 98% | 99% | 100% | 100% | 100% |
Illustrative values from the calculator. A trial with 50 per group is well powered for large effects and close to useless for small ones.
The inverse question
When n is capped, the more useful way to state a trial's capability is the minimum detectable effect: the smallest true effect for which the trial has your target power. With 50 per group and 80% power the answer is d = 0.57. If the SD of your endpoint is 10 units, that is 5.7 units. With 40 per group it is d = 0.63, or 6.3 units; with 64 per group it is d = 0.50. If the smallest clinically important difference is 5 units, a 40-per-group study cannot reliably detect it. Say that plainly in the protocol or the grant, instead of describing the study as powered "for a medium effect".
For proportions, the equivalent statement names the treatment rate. With a control rate of 30% and 60 per group, the pooled-variance calculation gives 80% power for a treatment rate of about 55%: a 25-point improvement. At 30% vs 50% those 60 per group give 61% power, and 100 per group gives 83%. The calculator reports the detectable rate directly and says when no rate is reachable at the given n.
Planned: 64 per group
d = 0.5, 80% power
Review at 50% of target enrolment
Retention is 78%, not 90%
Re-calculate power for the expected n
50 per group evaluable gives 70%
Decide: extend recruitment or accept
Minimum detectable d = 0.57
Common mistakes
Power calculations are easy to run and easy to misuse. These are the mistakes that matter most.
Calculating power using the effect and SD observed in the finished trial adds nothing: it is a one-to-one function of the p-value, so a non-significant result always shows "low power" and a borderline one always shows about 50%. Hoenig and Heisey (2001) set out why. If a trial was non-significant and you want to know what it could rule out, report the confidence interval. If you want to plan the next study, use a pre-specified, clinically motivated effect, not the observed one.
A small pilot gives a very imprecise estimate of an effect, and the effect it reports is a poor basis for the next study (Kraemer et al., 2006). Use the pilot for feasibility and for the SD, with a conservative limit, as the pilot study sample size calculator shows, and set the effect from clinical reasoning.
Eighty percent is a convention. A confirmatory Phase 3 trial often targets 90%, because the cost of a false negative is a lost development programme. A small rare-disease study may accept less, with the reasoning explained. Alpha is also a convention, but a regulatory one: changing it needs agreement in advance, as does moving from two-sided to one-sided.
If the trial has two co-primary endpoints, or three active arms against one control, the alpha used in each comparison is smaller than 0.05, and the power at the original alpha overstates what you have. Enter the adjusted alpha. For endpoints that must all be significant, the overall power is lower than the lowest individual power.
Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.
From the number to the protocol
The protocol states power as a design property: "With 64 participants per group the study has 80% power to detect a standardised difference of 0.5 using a two-sided t-test at the 5% level." A recruitment-shortfall plan adds a second sentence: the minimum detectable effect if only a stated fraction is evaluable. The SAP carries the same assumptions and, if an interim sample-size re-estimation is planned, specifies who sees what and when. The Bayesian clinical trial design guide covers designs that avoid fixed-n power statements; the how to write a clinical trial protocol guide places the sample size section within the protocol.
In Capture, the monitoring dashboard shows enrolment against target so the recalculation in the diagram above is triggered by data, not by a calendar. Visit windows show which participants are on track to reach the primary assessment, edit checks and auto-queries keep the primary endpoint complete, and randomization keeps allocation on plan. Blinded roles never receive treatment-arm values, so a team can review retention without seeing outcomes by arm. The two-means sample size calculator and survival sample size calculator give the n that this tool then stress-tests, and EDC for Phase 3 clinical trials describes running a trial at that scale.
Randomized
104 / 128
Retained at wk 12
78%
Projected evaluable per group
50
Projected power
70%
31 of 40
27 of 40
22 of 24
24 of 24
Before you quote a power
Smallest clinically important effect, not an observed one from the same trial.
SD from a comparable population; conservative if it comes from a small pilot.
Adjusted for co-primary endpoints, multiple arms or interim looks.
Dropout and missing primary outcomes taken off before computing power.
Stated in clinical units next to the power figure.
t-test for means; pooled or corrected for proportions as planned.
Power is the probability that a trial will produce a statistically significant result if the true effect equals the effect assumed in the calculation. It equals 1 minus the type II error rate. Eighty percent and 90% are the common targets.
Enter the participants per group, the expected difference and SD (or the two rates), and alpha. For means the power is the probability that the t statistic exceeds its critical value under the non-central t distribution with non-centrality δ / (σ √(2/n)).
It is the smallest true effect for which the trial has your target power at its sample size. With 50 per group, alpha 0.05 and 80% power it is d = 0.57. It is a more honest summary of a small trial than a statement of power for one effect.
Eighty percent is the minimum convention; 90% is common for confirmatory trials. The choice is part of the design and should be justified in the protocol, particularly if it is lower than 80%.
Calculating power from the observed effect and SD is uninformative because it is a function of the p-value. Use the confidence interval to describe what the result rules out, and use a clinically motivated effect for planning.
Yes. A one-sided test at alpha 0.05 has more power than a two-sided one at 0.05, but it is the same as a two-sided test at 0.10 in the favoured direction. Regulators generally expect a two-sided test or a one-sided alpha of 0.025.
The normal approximation treats the SD as known. The non-central t distribution accounts for estimating it and matches G*Power. The difference is small above about 50 per group and larger in small studies.
No. It is a free planning aid that reproduces published reference values. A statistician should confirm any power statement that goes into a protocol or grant.
Keep exploring
Two-means sample size calculator
Sample size for a continuous outcome.
Two-proportions sample size calculator
Sample size for a binary outcome.
Clinical trial sample size calculator
Quick two-proportion estimate.
Survival sample size calculator
Events needed for a log-rank test.
Sample size calculation
The concepts behind the formulas.
Dropout-adjusted sample size calculator
How many to enrol to keep the evaluable n.
Monitoring dashboard, visit windows and randomization in a free sandbox. No credit card. You pay only when you go live.