Free tool · Sample size · EquivalenceUpdated October 9, 2026

Equivalence trial sample size calculator for two one-sided tests (TOST)

An equivalence trial must rule out a difference in both directions. Enter the limits, the SD, CV or rates, and the expected true difference, and get n per group from exact TOST power, plus enrolment after dropout.

  • Means, proportions or log-scale ratios
  • Exact TOST power, t or normal
  • Shows why a true difference of zero is optimistic

Free sandbox · No credit card · 21 CFR Part 11 aligned

Demo: margin ±5 units, SD 10, alpha 0.05 per side, 80% power

Evaluable per group

70

If true difference is 2

139

Enrol per group (10% dropout)

78

True difference 070/320

70 per group

True difference 182/320

82 per group

True difference 2139/320

139 per group

True difference 3310/320

310 per group

Assuming perfect equivalence is the most optimistic input you can enter.

What this calculator does

  • An equivalence trial shows that two treatments differ by less than a margin in either direction. The null hypothesis is that the true difference is outside [−M, +M]; it is rejected only if both one-sided tests (TOST, Schuirmann 1987) reject.
  • At alpha 0.05 per side this is the same as the 90% confidence interval lying wholly inside the limits, the usual regulatory convention for bioequivalence (80.00% to 125.00% for the ratio of geometric means).
  • With a true difference of zero, n per group = 2σ²(z₁₋α + z₁₋β/2)² / M². Note β/2: with two bounds, the power budget is split. Margin ±5, SD 10, 80% power: 69 by the normal formula, 70 with the t correction.
  • The result is far more sensitive to the true difference than a superiority sample size: at a true difference of 2 (still well inside the margin) the same trial needs 139 per group; at 3 it needs 310.
  • Equivalence is not non-inferiority: it needs two limits and a two-sided CI. For a one-sided margin use the non-inferiority calculator.

Free tool

Calculate an equivalence (TOST) sample size

Choose a difference in means, a difference in proportions, or a ratio on the log scale with a CV. Enter the limits and the expected true difference. It runs in your browser; nothing is saved or sent.

Outcome
Method
%
Evaluable per group
70
140 in total, 1:1
Enrol per group
78
156 in total with 10% dropout

Power at this size: 80.6%. Equivalence is concluded when the 90% confidence interval lies wholly inside the limits.

n = 2σ² (z₁₋α + z₁₋β/2)² / M² = 2 × 10.0000² × (1.645 + 1.282)² / 5.0000² = 68.5 (normal formula, true difference zero)

Planning aid and reference only, not validated software. Two one-sided tests (TOST, Schuirmann 1987) for two independent groups, 1:1. The null hypothesis is that the true difference lies outside the limits, so an equivalence trial needs a pre-specified margin on both sides. Power is computed exactly for the model (normal, or the usual non-central t approximation); sizes are lower when the true difference is zero and rise sharply as it nears a limit. A 2×2 crossover bioequivalence study uses the within-subject CV and a different variance, which this tool does not model. Confirm the margin and the sample size with a statistician and, where relevant, the applicable guideline.

The design

How TOST turns "the same" into two hypotheses

You cannot prove that two treatments are identical, only that any difference is smaller than one you consider unimportant. TOST (Schuirmann, 1987) turns that into two one-sided hypotheses. H₀₁ says the true difference is at or below −M, H₀₂ says it is at or above +M. Each is tested at level α, and equivalence is concluded only if both are rejected. Equivalent and easier to read: compute the (1 − 2α) confidence interval for the difference and check it lies inside (−M, +M). At α = 0.05 per side that is a 90% interval, which is why bioequivalence studies report a 90% CI, and why a 95% interval requires α = 0.025.

Power is the probability of rejecting both nulls when the true difference is Δ. In the normal case it is Φ((Δ + M)/SE − z₁₋α) + Φ((M − Δ)/SE − z₁₋α) − 1, where SE = σ √(2/n). The calculator evaluates this exactly and searches for the smallest n that reaches your target; the t option uses the standard non-central t approximation. When Δ = 0, the expression simplifies to the familiar closed form with z₁₋β/2, and the example margin ±5, SD 10, 80% power gives 2 x 100 x (1.645 + 1.282)² / 25 = 68.5, so 69 by the normal formula (70 with t).

For regulatory bioequivalence the margin is fixed by guidance: the 90% CI for the ratio of geometric means for AUC and Cmax must lie within 80.00% to 125.00% (EMA guideline CPMP/EWP/QWP/1401/98 Rev. 1 and FDA guidance, with exceptions such as narrow-therapeutic-index drugs). For clinical equivalence trials, including biosimilar comparisons, the margin is a clinical and statistical judgement agreed with the regulator in advance. ICH E9 (section 3.3.2) discusses the design and requires the margin to be specified in the protocol.

Ratios on the log scale

Pharmacokinetic measures are log-normal, so the analysis is on the log scale and the limits become ln(0.80) and ln(1.25). The variance of the log values is σ² = ln(1 + CV²), where CV is the coefficient of variation as a fraction. With CV 20% and a true ratio of 1.00, σ = 0.198. A parallel-group study needs 15 per group; if the true ratio is 0.95 it needs 18. The calculator lets you set unequal log limits, and shows how quickly n rises when the true ratio drifts toward a limit. A 2x2 crossover study uses the within-subject CV and needs far fewer participants; this tool covers parallel groups only.

Sensitivity

Participants per group by margin and true difference

Difference in means, SD 10, alpha 0.05 per side, 80% power, 1:1, t option, rounded up.

Equivalence margin (±M)True difference 0True difference 1True difference 2
31913111,238
4108141310
57082139
6495479
8283036

Illustrative values from the calculator. A narrower margin is expensive, and a true difference near the margin is more expensive still.

Log-scale ratios

Parallel-group bioequivalence-style sample sizes by CV

Limits 0.80 to 1.25, alpha 0.05 per side, 1:1, t option, participants per group.

CVTrue ratio 1.00, 80% powerTrue ratio 0.95, 80% powerTrue ratio 1.00, 90% power
15%91111
20%151818
25%222728
30%313839
40%526566

Parallel groups. Crossover designs use the within-subject CV and need fewer participants; use a crossover-specific tool for those.

Choosing the design

Equivalence, non-inferiority and superiority are different questions

The three designs differ in the null hypothesis, the number of limits and the alpha, so a number from one design cannot be used for another. Choose the design from the question, then compute.

When equivalence is the right question

Use equivalence when a difference in either direction would matter: a generic or biosimilar compared with a reference product, a new formulation, a change in manufacturing process, or a device that must perform the same as a predicate. Use non-inferiority when only "not worse" matters and a better result would be welcome; the non-inferiority sample size calculator uses one margin and one-sided alpha 0.025. Use superiority when you need to show a difference, as in the two-means and two-proportions calculators.

Three designs, three sample size questions
SuperiorityNon-inferiorityEquivalence
Limits neededNoneOne marginTwo limits
Null hypothesisNo differenceWorse by M or moreOutside the limits
Conventional alpha0.05 two-sided0.025 one-sided0.05 per side (90% CI)
Power budgetz(β)z(β)z(β/2) at zero difference
Better result is a success
With the same margin and power, equivalence needs somewhat more participants than non-inferiority.

Common mistakes

Where equivalence sample sizes go wrong

Equivalence designs have a few traps beyond those of ordinary sample size calculations.

Assuming the true difference is exactly zero

Two treatments are rarely identical. If you honestly expect a small difference, put it in. At a margin of 5 and SD 10, going from 0 to 2 raises n from 70 to 139. A sample size calculated for a true difference of zero is a best case, and it is the commonest reason equivalence trials fail to show equivalence.

Using a 95% CI and the 90% sample size

If the SAP says equivalence will be concluded from a 95% confidence interval, each one-sided test has α = 0.025 and the sample size is larger than at α = 0.05. Match the alpha to the interval the analysis will report.

Margin chosen after the fact

The margin must come from clinical reasoning and prior evidence, and be fixed in the protocol before any data are seen. A margin picked to match the sample size you can afford is the equivalence version of working backwards from the budget, and regulators will ask how it was justified.

Reading a wide interval as equivalence

Failing to find a difference is not evidence of equivalence. A small or noisy trial has a wide interval that crosses the limits. Only a confidence interval inside the limits shows equivalence, which is what the design guarantees and a superiority trial with a non-significant p-value does not.

Keep both arms' data clean enough to show equivalence

Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.

Build your study free

From the number to the protocol

How this feeds the protocol, the SAP and the study build

The protocol sentence for an equivalence trial carries more than the usual number: "Equivalence will be concluded if the 90% confidence interval for the difference in mean change in the primary endpoint lies within ±5 units. Assuming a true difference of 1 unit and a common SD of 10, 82 evaluable participants per group provide 80% power. Allowing for 10% dropout, 92 participants per group (184 in total) will be randomized." The SAP specifies the analysis populations: unlike superiority trials, equivalence conclusions should hold in both the full analysis set and the per-protocol set, because poor conduct tends to push the arms together, and regulators look for consistency between them. The clinical trial endpoints entry and the how to write a clinical trial protocol guide cover the surrounding sections.

Because noise and deviations favour a finding of equivalence, the data system is part of the design. In Capture, edit checks raise an auto-query when a value breaks a range or custom rule, source data verification can be set per field for the primary endpoint, and the field-level audit trail records old value, new value, user, time and reason for every change. Protocol deviation tracking records who falls out of the per-protocol set and why. Randomization is blinded by role, with masked views, so outcome entry is not influenced by allocation. Exports come as CSV or Excel with a data dictionary, or as SDTM datasets in SAS XPT files with Define-XML, ready for your statistician. For the dropout step see the dropout-adjusted sample size calculator, and for the quick binary estimate the clinical trial sample size calculator.

Statistical analysis plan, equivalence trial (demo)
SAP_equivalence_DEMO.docx3/5 complete
  1. 5

    Analysis sets: full analysis and per-protocol

    Both must support the conclusion

    Complete
  2. 7

    Primary analysis: 90% CI against ±5 units

    TOST at 5% per side

    Complete
  3. 7.3

    Margin justification

    Cites prior effect and clinical review

    Complete
  4. 9

    Sample size: 82 per group (true difference 1, SD 10)

    80% power, 10% dropout to 92

    Draft
  5. 9.1

    Sensitivity: true difference 2 needs 139

    To do
The margin justification and the sensitivity table are what a reviewer reads first.

Before you lock it

Equivalence sample size checklist

Margins pre-specified and justified

Clinical reasoning and prior evidence, fixed before data are seen.

Alpha matches the interval

0.05 per side for a 90% CI, 0.025 for a 95% CI.

Realistic true difference

Not zero unless the treatments are genuinely expected to be identical.

Right scale

Log scale and CV for ratios; absolute scale for differences.

Design matches the formula

Parallel groups here; crossover needs the within-subject variance.

Analysis sets planned

Equivalence supported in both full analysis and per-protocol sets.

FAQ

Questions teams ask before they switch

Something not covered here? Ask us directly.

What is the sample size formula for an equivalence trial?

For a difference in means with a true difference of zero, n per group = 2σ²(z₁₋α + z₁₋β/2)² / M², where M is the equivalence margin. If the true difference is not zero, power has no simple closed form; the calculator finds the smallest n by evaluating TOST power directly.

Why is power split as β/2 in equivalence trials?

Equivalence needs two one-sided tests to both succeed. At a true difference of zero, each fails with probability about β/2, so the combined failure probability is β. This is why the closed form uses z₁₋β/2.

What is the difference between equivalence and non-inferiority sample sizes?

Non-inferiority uses one margin and one one-sided test (alpha 0.025), so the denominator involves one distance. Equivalence needs both limits to be cleared, which with the same margin and power takes somewhat more participants. A better result is a success in non-inferiority, not in equivalence.

Why does a 90% confidence interval appear in equivalence studies?

Two one-sided tests at 5% each are equivalent to checking that the 90% confidence interval lies within the limits. If you use a 95% interval, each test is at 2.5%.

How do I choose the equivalence margin?

For regulatory bioequivalence it is set by guidance (80.00% to 125.00% for the ratio of geometric means in most cases). For clinical equivalence it comes from the smallest difference that matters clinically and prior evidence, agreed with the regulator, and written in the protocol before the trial.

Does the calculator handle crossover bioequivalence studies?

No. It covers two parallel groups. A 2x2 crossover uses the within-subject variance and needs fewer participants, so use a tool built for that design.

What if my primary outcome is a ratio such as a hazard ratio?

This tool handles differences in means, differences in proportions and ratios of geometric means. Equivalence of time-to-event outcomes needs a different formula or simulation.

Is this calculator validated?

No. It is a free planning aid that reproduces closed-form results. A statistician should confirm the margin, the formula and the sample size before the protocol is submitted.

Keep the data clean enough to show equivalence

Edit checks, audit trail and blinded randomization in a free sandbox. No credit card. You pay only when you go live.

Build your study free