A pilot does not need 80% power, because it is not testing the treatment. Pick what it must establish: a usable SD for the main trial, a recruitment or retention rate, or a chance to spot a design problem, and get a justified n with the formula shown.
Free sandbox · No credit card · 21 CFR Part 11 aligned
Pilot: 24 participants (12 per group)
Feasibility and SD, not efficacy
Progression criteria checked
Recruitment, retention, adherence thresholds
Main-trial SD set from the 80% upper limit
Observed 10 becomes 11.6
Main trial sized
86 per group instead of 64
What this calculator does
Free tool
Choose what the pilot should establish. Each mode shows its formula and what it assumes. It runs in your browser; nothing is saved or sent.
| SD used for the main trial | n per group |
|---|---|
| Observed (10) | 64 |
| Upper 80% limit (11.6) | 86 |
| Upper end of the 95% CI (14.2) | 127 |
Two-sided alpha 0.05, 1:1, t-distribution. The pilot SD is itself an estimate: plan the main trial on a pessimistic limit, not the point value.
Planning aid and reference only, not validated software. A pilot is for learning about feasibility and about design parameters such as the SD; it is not a small efficacy trial, and it should not be used to estimate the treatment effect for the main trial's sample size (Kraemer et al., Arch Gen Psychiatry 2006). The SD limits use the chi-square distribution with 2(n − 1) degrees of freedom (two groups, pooled SD). The rate formula is the Wald interval, which is rough for rates near 0% or 100%; the problem-detection formula is from Viechtbauer et al. (2015). Your statistician should confirm the choice.
The principle
A pilot has a job, and the sample size follows from the job. The CONSORT extension for pilot and feasibility trials (Eldridge et al., BMJ 2016) says a pilot should state its feasibility objectives and progression criteria, and should not present effect estimates as the main result. Three purposes cover most pilots, and each has a different formula.
The first is estimating a parameter for the main trial, usually the SD of a continuous outcome. The SD is estimated with degrees of freedom 2(n − 1) for two groups of n, and its precision is read from the chi-square distribution: the 95% limits are the observed SD multiplied by √(df / χ²) at the 97.5th and 2.5th percentiles. With 12 per group (22 df) that gives 0.77 and 1.42; with 6 per group the range is 0.70 to 1.75; with 35 per group, 0.86 to 1.20. The second is estimating a feasibility rate: the proportion who consent, stay in the study or take the intervention. The third is detecting design problems such as an unclear eligibility criterion or a questionnaire item people misread. Each purpose gives a defensible n, and none of them needs an effect size.
It is tempting to take the treatment effect from the pilot and size the main trial on it. Kraemer and colleagues (2006) showed this is unreliable: a small trial estimates the effect so imprecisely that the resulting sample size is likely to be too small, or, if the pilot got lucky, far too large, and selecting on a promising result makes the bias worse. Take the effect from clinical reasoning (the minimal clinically important difference is the usual anchor) and use the pilot for the SD and for feasibility.
Precision of the SD
Pooled SD from two equal groups, df = 2(n - 1). Multipliers apply to the observed SD.
| Pilot size per group | Degrees of freedom | 95% CI for the SD | Upper 80% limit |
|---|---|---|---|
| 6 | 10 | x0.70 to x1.75 | x1.27 |
| 10 | 18 | x0.76 to x1.48 | x1.18 |
| 12 | 22 | x0.77 to x1.42 | x1.16 |
| 15 | 28 | x0.79 to x1.35 | x1.14 |
| 20 | 38 | x0.82 to x1.29 | x1.12 |
| 35 | 68 | x0.86 to x1.20 | x1.08 |
Values from the chi-square distribution, as in the calculator. The upper 80% limit is the inflation used by Browne (1995).
Worked example
A pilot with 12 per group observes an SD of 10 on the primary outcome. The main trial wants to detect a difference of 5 with 80% power at a two-sided alpha of 0.05. Taking 10 at face value, the two-means sample size calculator returns 64 per group. But the true SD could be as high as 14.2 (the top of the 95% interval), where the same trial needs 127 per group. Browne (1995) proposed using an upper confidence limit instead of the point estimate, and an 80% limit is the level usually cited; the upper 80% limit here is 10 x 1.16 = 11.6, which gives 86 per group. That is 34% more participants than the naive plan, in exchange for much lower odds of an underpowered main trial.
For a feasibility rate, suppose the pilot must show that at least 70% of those enrolled finish the primary assessment. Estimating a 70% retention rate to within ±10 points at 95% confidence needs n = 1.96² x 0.70 x 0.30 / 0.10² = 81 participants. If you have no prior view of the rate, the worst case (50%) needs 97. Narrowing the interval to ±5 points when you expect 80% takes 246. This is why pilots that try to estimate many rates precisely become larger than the main trial's first cohort: choose the one or two rates that decide whether to proceed, and set progression thresholds ahead of time.
For problem detection, the chance of seeing at least one participant with a problem of probability π among n is 1 − (1 − π)ⁿ, so n = ln(1 − confidence) / ln(1 − π). A problem that affects 1 in 20 participants needs 59 for 95% confidence; 1 in 10 needs 29; 1 in 5 needs 14; 1 in 100 needs 299. Seeing a problem needs an observer, so the pilot should record deviations, queries and participant comments systematically.
Rules of thumb
Different recommendations answer different questions. Always state which one you are using.
| Source | Recommendation | Aim |
|---|---|---|
| Julious, Pharm Stat 2005 | 12 per group (24 in total) | Minimum with usable precision for the mean and the SD, when no prior information exists. |
| Browne, Stat Med 1995 | Use an upper confidence limit of the pilot SD, rather than the point estimate (80% is the level commonly cited) | Protect the main trial against an underestimated SD. |
| Teare et al., Trials 2014 | About 35 per group (70 in total) for a continuous outcome | Estimate the SD with little further gain from more participants. More are needed to estimate a binary outcome rate. |
| Whitehead et al., SMMR 2016 | Stepped by planned standardised effect size; for a 90% powered main trial, roughly 75, 25, 15 and 10 per arm for effects up to about 0.1, 0.1 to 0.3, 0.3 to 0.7 and above 0.7 | Minimise the combined sample of pilot and main trial. Check the paper's table before citing in a protocol. |
| Viechtbauer et al., J Clin Epidemiol 2015 | n = ln(1 - confidence) / ln(1 - π), for example 59 for π = 5% and 95% confidence | Detect design or procedural problems. |
The Whitehead figures are summarised from secondary sources; confirm them in the original paper.
Common mistakes
These are the mistakes that most often undermine a pilot or the main trial that follows it.
A pilot still needs a justification: a stated purpose, a formula or a precision target, and a plan for what happens next. "Pilot, n = 30" with no reason invites the question from funders and ethics committees. Use the modes above and write the reasoning down, and see how to run a pilot study for the wider process.
Progression criteria should be about feasibility (recruitment, retention, adherence, data completeness), with thresholds set in advance, often as a traffic light with a go, amend or stop zone. Statistical significance of the treatment effect in a pilot is not one of them.
The table above shows how wide the uncertainty is. If the main trial is expensive, spend more on the pilot or build in a sample-size re-estimation point, which is cheaper than an underpowered main trial. The power analysis calculator shows what a given n can detect if the SD turns out higher.
An internal pilot (the first part of the main trial) can be rolled into the final analysis; an external pilot generally cannot. Decide before the pilot starts, because it changes the protocol, the consent forms and the data system.
Consented / approached
41%
Retained at week 8
79%
Diary completion
92%
Decision
Amend
Go
Amend
Go
Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.
From the number to the protocol
Write the pilot sample size as a justification of purpose: "Twenty-four participants (12 per group) will be randomized. This provides a 95% confidence interval for the SD of the primary outcome of 0.77 to 1.42 times the observed value, and the upper 80% limit will be used to size the main trial. The proportion of eligible people consenting and the 8-week retention rate will be reported with 95% confidence intervals, judged against the progression criteria in Table 3." The SAP should say that no hypothesis test of the treatment effect is planned, or that any estimate is descriptive. The how to write a clinical trial protocol guide covers the surrounding sections.
In Capture, a pilot can use the same system as the main trial, so what is learned is what will be used. The study team builds the eCRFs with edit checks, defines the visit schedule with windows, and tests every form and skip rule in the free sandbox before the first participant is enrolled. During the pilot, query rates, missing-data patterns and visit-window deviations show which parts of the data collection need to change, the monitoring dashboard reports progress against the thresholds you set, and ePRO diaries on participants' phones give completion rates for the questionnaire burden question. The field-level audit trail records each change with a reason, and CSV or Excel exports with a data dictionary give the statistician the SD and rates. See EDC for pilot clinical trials and EDC for feasibility studies for the setup, and the two-means and two-proportions calculators for the main trial that follows.
Before the pilot starts
SD estimate, feasibility rates, problem detection, or a mix, written as objectives.
Julious 12 per group, a precision target, or n = ln(1 - c) / ln(1 - π), with its source.
Go, amend and stop thresholds for recruitment, retention and adherence.
Treatment effect reported descriptively, with a confidence interval, if at all.
Main trial sized on the 80% upper limit or sensitivity range, not the point estimate.
External (separate) or internal (rolled into the main trial) before it starts.
It depends on the purpose. A common minimum for a two-arm pilot with no prior information is 12 per group (Julious, 2005). Estimating the SD well points to about 35 per group (Teare et al., 2014). Estimating a feasibility rate to within ±10 points at 95% confidence needs about 60 to 100 depending on the rate.
Julious (Pharmaceutical Statistics, 2005) argued that 12 per group balances feasibility, the precision of the mean and variance, and regulatory considerations, as a minimum when no other information exists. With 12 per group the 95% CI for the SD is about 0.77 to 1.42 times the observed SD.
No. A pilot is not designed to test the treatment effect. Justify its size by the precision it gives for a feasibility rate or a design parameter such as the SD, or by the chance of detecting a problem.
It is not recommended. A small pilot estimates the effect too imprecisely and selecting on a promising result biases it upward. Use a clinically important difference, and use the pilot for the SD and for feasibility.
Multiply the pilot SD by the upper-limit factor from the chi-square distribution (for example 1.16 for the upper 80% limit with 12 per group) and size the main trial on that value. The calculator shows the main-trial n under the point estimate, the 80% limit and the 95% limit.
A pre-specified threshold for a feasibility measure, such as consent rate, retention or adherence, that decides whether the main trial goes ahead, goes ahead after changes, or stops. The CONSORT extension for pilot and feasibility trials asks for them to be reported.
n = ln(1 − confidence) / ln(1 − π), where π is the chance a participant hits the problem. For π = 5% and 95% confidence it is 59; for 10% it is 29 (Viechtbauer et al., 2015).
No. It is a free planning aid. The feasibility-rate mode uses the simple Wald interval, which is rough for rates near 0% or 100%. A statistician should confirm the pilot design.
Keep exploring
How to run a pilot study
Objectives, progression criteria and reporting.
EDC for pilot clinical trials
Run the pilot in the system the main trial will use.
EDC for feasibility studies
Capture recruitment and retention rates.
Two-means sample size calculator
Size the main trial from the pilot SD.
Power analysis calculator
What a given n can detect.
Dropout-adjusted sample size calculator
Turn evaluable numbers into enrolment.
Build and test every form, visit window and ePRO diary in a free sandbox. No credit card. You pay only when you go live.