An equivalence trial must rule out a difference in both directions. Enter the limits, the SD, CV or rates, and the expected true difference, and get n per group from exact TOST power, plus enrolment after dropout.
Free sandbox · No credit card · 21 CFR Part 11 aligned
Evaluable per group
70
If true difference is 2
139
Enrol per group (10% dropout)
78
70 per group
82 per group
139 per group
310 per group
What this calculator does
Free tool
Choose a difference in means, a difference in proportions, or a ratio on the log scale with a CV. Enter the limits and the expected true difference. It runs in your browser; nothing is saved or sent.
Power at this size: 80.6%. Equivalence is concluded when the 90% confidence interval lies wholly inside the limits.
n = 2σ² (z₁₋α + z₁₋β/2)² / M² = 2 × 10.0000² × (1.645 + 1.282)² / 5.0000² = 68.5 (normal formula, true difference zero)
Planning aid and reference only, not validated software. Two one-sided tests (TOST, Schuirmann 1987) for two independent groups, 1:1. The null hypothesis is that the true difference lies outside the limits, so an equivalence trial needs a pre-specified margin on both sides. Power is computed exactly for the model (normal, or the usual non-central t approximation); sizes are lower when the true difference is zero and rise sharply as it nears a limit. A 2×2 crossover bioequivalence study uses the within-subject CV and a different variance, which this tool does not model. Confirm the margin and the sample size with a statistician and, where relevant, the applicable guideline.
The design
You cannot prove that two treatments are identical, only that any difference is smaller than one you consider unimportant. TOST (Schuirmann, 1987) turns that into two one-sided hypotheses. H₀₁ says the true difference is at or below −M, H₀₂ says it is at or above +M. Each is tested at level α, and equivalence is concluded only if both are rejected. Equivalent and easier to read: compute the (1 − 2α) confidence interval for the difference and check it lies inside (−M, +M). At α = 0.05 per side that is a 90% interval, which is why bioequivalence studies report a 90% CI, and why a 95% interval requires α = 0.025.
Power is the probability of rejecting both nulls when the true difference is Δ. In the normal case it is Φ((Δ + M)/SE − z₁₋α) + Φ((M − Δ)/SE − z₁₋α) − 1, where SE = σ √(2/n). The calculator evaluates this exactly and searches for the smallest n that reaches your target; the t option uses the standard non-central t approximation. When Δ = 0, the expression simplifies to the familiar closed form with z₁₋β/2, and the example margin ±5, SD 10, 80% power gives 2 x 100 x (1.645 + 1.282)² / 25 = 68.5, so 69 by the normal formula (70 with t).
For regulatory bioequivalence the margin is fixed by guidance: the 90% CI for the ratio of geometric means for AUC and Cmax must lie within 80.00% to 125.00% (EMA guideline CPMP/EWP/QWP/1401/98 Rev. 1 and FDA guidance, with exceptions such as narrow-therapeutic-index drugs). For clinical equivalence trials, including biosimilar comparisons, the margin is a clinical and statistical judgement agreed with the regulator in advance. ICH E9 (section 3.3.2) discusses the design and requires the margin to be specified in the protocol.
Pharmacokinetic measures are log-normal, so the analysis is on the log scale and the limits become ln(0.80) and ln(1.25). The variance of the log values is σ² = ln(1 + CV²), where CV is the coefficient of variation as a fraction. With CV 20% and a true ratio of 1.00, σ = 0.198. A parallel-group study needs 15 per group; if the true ratio is 0.95 it needs 18. The calculator lets you set unequal log limits, and shows how quickly n rises when the true ratio drifts toward a limit. A 2x2 crossover study uses the within-subject CV and needs far fewer participants; this tool covers parallel groups only.
Sensitivity
Difference in means, SD 10, alpha 0.05 per side, 80% power, 1:1, t option, rounded up.
| Equivalence margin (±M) | True difference 0 | True difference 1 | True difference 2 |
|---|---|---|---|
| 3 | 191 | 311 | 1,238 |
| 4 | 108 | 141 | 310 |
| 5 | 70 | 82 | 139 |
| 6 | 49 | 54 | 79 |
| 8 | 28 | 30 | 36 |
Illustrative values from the calculator. A narrower margin is expensive, and a true difference near the margin is more expensive still.
Log-scale ratios
Limits 0.80 to 1.25, alpha 0.05 per side, 1:1, t option, participants per group.
| CV | True ratio 1.00, 80% power | True ratio 0.95, 80% power | True ratio 1.00, 90% power |
|---|---|---|---|
| 15% | 9 | 11 | 11 |
| 20% | 15 | 18 | 18 |
| 25% | 22 | 27 | 28 |
| 30% | 31 | 38 | 39 |
| 40% | 52 | 65 | 66 |
Parallel groups. Crossover designs use the within-subject CV and need fewer participants; use a crossover-specific tool for those.
Choosing the design
The three designs differ in the null hypothesis, the number of limits and the alpha, so a number from one design cannot be used for another. Choose the design from the question, then compute.
Use equivalence when a difference in either direction would matter: a generic or biosimilar compared with a reference product, a new formulation, a change in manufacturing process, or a device that must perform the same as a predicate. Use non-inferiority when only "not worse" matters and a better result would be welcome; the non-inferiority sample size calculator uses one margin and one-sided alpha 0.025. Use superiority when you need to show a difference, as in the two-means and two-proportions calculators.
| Superiority | Non-inferiority | Equivalence | |
|---|---|---|---|
| Limits needed | None | One margin | Two limits |
| Null hypothesis | No difference | Worse by M or more | Outside the limits |
| Conventional alpha | 0.05 two-sided | 0.025 one-sided | 0.05 per side (90% CI) |
| Power budget | z(β) | z(β) | z(β/2) at zero difference |
| Better result is a success |
Common mistakes
Equivalence designs have a few traps beyond those of ordinary sample size calculations.
Two treatments are rarely identical. If you honestly expect a small difference, put it in. At a margin of 5 and SD 10, going from 0 to 2 raises n from 70 to 139. A sample size calculated for a true difference of zero is a best case, and it is the commonest reason equivalence trials fail to show equivalence.
If the SAP says equivalence will be concluded from a 95% confidence interval, each one-sided test has α = 0.025 and the sample size is larger than at α = 0.05. Match the alpha to the interval the analysis will report.
The margin must come from clinical reasoning and prior evidence, and be fixed in the protocol before any data are seen. A margin picked to match the sample size you can afford is the equivalence version of working backwards from the budget, and regulators will ask how it was justified.
Failing to find a difference is not evidence of equivalence. A small or noisy trial has a wide interval that crosses the limits. Only a confidence interval inside the limits shows equivalence, which is what the design guarantees and a superiority trial with a non-significant p-value does not.
Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.
From the number to the protocol
The protocol sentence for an equivalence trial carries more than the usual number: "Equivalence will be concluded if the 90% confidence interval for the difference in mean change in the primary endpoint lies within ±5 units. Assuming a true difference of 1 unit and a common SD of 10, 82 evaluable participants per group provide 80% power. Allowing for 10% dropout, 92 participants per group (184 in total) will be randomized." The SAP specifies the analysis populations: unlike superiority trials, equivalence conclusions should hold in both the full analysis set and the per-protocol set, because poor conduct tends to push the arms together, and regulators look for consistency between them. The clinical trial endpoints entry and the how to write a clinical trial protocol guide cover the surrounding sections.
Because noise and deviations favour a finding of equivalence, the data system is part of the design. In Capture, edit checks raise an auto-query when a value breaks a range or custom rule, source data verification can be set per field for the primary endpoint, and the field-level audit trail records old value, new value, user, time and reason for every change. Protocol deviation tracking records who falls out of the per-protocol set and why. Randomization is blinded by role, with masked views, so outcome entry is not influenced by allocation. Exports come as CSV or Excel with a data dictionary, or as SDTM datasets in SAS XPT files with Define-XML, ready for your statistician. For the dropout step see the dropout-adjusted sample size calculator, and for the quick binary estimate the clinical trial sample size calculator.
Analysis sets: full analysis and per-protocol
Both must support the conclusion
Primary analysis: 90% CI against ±5 units
TOST at 5% per side
Margin justification
Cites prior effect and clinical review
Sample size: 82 per group (true difference 1, SD 10)
80% power, 10% dropout to 92
Sensitivity: true difference 2 needs 139
Before you lock it
Clinical reasoning and prior evidence, fixed before data are seen.
0.05 per side for a 90% CI, 0.025 for a 95% CI.
Not zero unless the treatments are genuinely expected to be identical.
Log scale and CV for ratios; absolute scale for differences.
Parallel groups here; crossover needs the within-subject variance.
Equivalence supported in both full analysis and per-protocol sets.
For a difference in means with a true difference of zero, n per group = 2σ²(z₁₋α + z₁₋β/2)² / M², where M is the equivalence margin. If the true difference is not zero, power has no simple closed form; the calculator finds the smallest n by evaluating TOST power directly.
Equivalence needs two one-sided tests to both succeed. At a true difference of zero, each fails with probability about β/2, so the combined failure probability is β. This is why the closed form uses z₁₋β/2.
Non-inferiority uses one margin and one one-sided test (alpha 0.025), so the denominator involves one distance. Equivalence needs both limits to be cleared, which with the same margin and power takes somewhat more participants. A better result is a success in non-inferiority, not in equivalence.
Two one-sided tests at 5% each are equivalent to checking that the 90% confidence interval lies within the limits. If you use a 95% interval, each test is at 2.5%.
For regulatory bioequivalence it is set by guidance (80.00% to 125.00% for the ratio of geometric means in most cases). For clinical equivalence it comes from the smallest difference that matters clinically and prior evidence, agreed with the regulator, and written in the protocol before the trial.
No. It covers two parallel groups. A 2x2 crossover uses the within-subject variance and needs fewer participants, so use a tool built for that design.
This tool handles differences in means, differences in proportions and ratios of geometric means. Equivalence of time-to-event outcomes needs a different formula or simulation.
No. It is a free planning aid that reproduces closed-form results. A statistician should confirm the margin, the formula and the sample size before the protocol is submitted.
Keep exploring
Non-inferiority sample size calculator
One margin and one-sided alpha.
Two-means sample size calculator
Superiority comparison of means.
Two-proportions sample size calculator
Superiority comparison of rates.
Power analysis calculator
What a fixed n can detect.
Dropout-adjusted sample size calculator
Turn evaluable numbers into enrolment.
Protocol deviation tracking software
Who is in the per-protocol set, and why.
Edit checks, audit trail and blinded randomization in a free sandbox. No credit card. You pay only when you go live.