Free tool · Pilot and feasibility studiesUpdated October 9, 2026

Pilot study sample size calculator: precision, feasibility and problem detection

A pilot does not need 80% power, because it is not testing the treatment. Pick what it must establish: a usable SD for the main trial, a recruitment or retention rate, or a chance to spot a design problem, and get a justified n with the formula shown.

  • SD confidence limits for 12 per group and more
  • Feasibility-rate interval width
  • Chance of spotting a problem

Free sandbox · No credit card · 21 CFR Part 11 aligned

Pilot to main trial (demo plan)
  1. Pilot: 24 participants (12 per group)

    Feasibility and SD, not efficacy

  2. Progression criteria checked

    Recruitment, retention, adherence thresholds

  3. 3

    Main-trial SD set from the 80% upper limit

    Observed 10 becomes 11.6

  4. 4

    Main trial sized

    86 per group instead of 64

The pilot is judged on its pre-specified progression criteria, not on whether its treatment effect was significant.

What this calculator does

  • A pilot or feasibility study tests whether the main trial can be run and estimates design inputs such as the SD and the recruitment rate. It is not a small efficacy trial and should not be powered like one.
  • The best-known rule of thumb is 12 per group (Julious, 2005). With 12 per group the 95% CI for the SD runs from 0.77 to 1.42 times the observed SD, a precision that is often just good enough.
  • Teare et al. (2014) found that about 35 per group (70 in total) gave a much tighter SD estimate for continuous outcomes; if the pilot also estimates a rate, more are needed.
  • To estimate a feasibility rate (for example retention of 80%) to within ±10 points with 95% confidence needs 62 participants; ±5 points needs 246.
  • To have a 95% chance of seeing a problem at least once that affects 5% of participants, you need 59 (Viechtbauer et al., 2015); for 10% it is 29.

Free tool

Size a pilot for the question it has to answer

Choose what the pilot should establish. Each mode shows its formula and what it assumes. It runs in your browser; nothing is saved or sent.

What should the pilot establish?
95% CI for the true SD
7.7 to 14.2
12 per group, 22 df: ×0.77 to ×1.42 of the observed 10
Main-trial sample size per group by SD assumption
SD used for the main trialn per group
Observed (10)64
Upper 80% limit (11.6)86
Upper end of the 95% CI (14.2)127

Two-sided alpha 0.05, 1:1, t-distribution. The pilot SD is itself an estimate: plan the main trial on a pessimistic limit, not the point value.

Planning aid and reference only, not validated software. A pilot is for learning about feasibility and about design parameters such as the SD; it is not a small efficacy trial, and it should not be used to estimate the treatment effect for the main trial's sample size (Kraemer et al., Arch Gen Psychiatry 2006). The SD limits use the chi-square distribution with 2(n − 1) degrees of freedom (two groups, pooled SD). The rate formula is the Wald interval, which is rough for rates near 0% or 100%; the problem-detection formula is from Viechtbauer et al. (2015). Your statistician should confirm the choice.

The principle

Size the pilot for its purpose, not for power

A pilot has a job, and the sample size follows from the job. The CONSORT extension for pilot and feasibility trials (Eldridge et al., BMJ 2016) says a pilot should state its feasibility objectives and progression criteria, and should not present effect estimates as the main result. Three purposes cover most pilots, and each has a different formula.

The first is estimating a parameter for the main trial, usually the SD of a continuous outcome. The SD is estimated with degrees of freedom 2(n − 1) for two groups of n, and its precision is read from the chi-square distribution: the 95% limits are the observed SD multiplied by √(df / χ²) at the 97.5th and 2.5th percentiles. With 12 per group (22 df) that gives 0.77 and 1.42; with 6 per group the range is 0.70 to 1.75; with 35 per group, 0.86 to 1.20. The second is estimating a feasibility rate: the proportion who consent, stay in the study or take the intervention. The third is detecting design problems such as an unclear eligibility criterion or a questionnaire item people misread. Each purpose gives a defensible n, and none of them needs an effect size.

Why not use the pilot's effect size?

It is tempting to take the treatment effect from the pilot and size the main trial on it. Kraemer and colleagues (2006) showed this is unreliable: a small trial estimates the effect so imprecisely that the resulting sample size is likely to be too small, or, if the pilot got lucky, far too large, and selecting on a promising result makes the bias worse. Take the effect from clinical reasoning (the minimal clinically important difference is the usual anchor) and use the pilot for the SD and for feasibility.

Precision of the SD

How well does a pilot pin down the SD?

Pooled SD from two equal groups, df = 2(n - 1). Multipliers apply to the observed SD.

Pilot size per groupDegrees of freedom95% CI for the SDUpper 80% limit
610x0.70 to x1.75x1.27
1018x0.76 to x1.48x1.18
1222x0.77 to x1.42x1.16
1528x0.79 to x1.35x1.14
2038x0.82 to x1.29x1.12
3568x0.86 to x1.20x1.08

Values from the chi-square distribution, as in the calculator. The upper 80% limit is the inflation used by Browne (1995).

Worked example

What a 12-per-group pilot does to the main trial's size

A pilot with 12 per group observes an SD of 10 on the primary outcome. The main trial wants to detect a difference of 5 with 80% power at a two-sided alpha of 0.05. Taking 10 at face value, the two-means sample size calculator returns 64 per group. But the true SD could be as high as 14.2 (the top of the 95% interval), where the same trial needs 127 per group. Browne (1995) proposed using an upper confidence limit instead of the point estimate, and an 80% limit is the level usually cited; the upper 80% limit here is 10 x 1.16 = 11.6, which gives 86 per group. That is 34% more participants than the naive plan, in exchange for much lower odds of an underpowered main trial.

For a feasibility rate, suppose the pilot must show that at least 70% of those enrolled finish the primary assessment. Estimating a 70% retention rate to within ±10 points at 95% confidence needs n = 1.96² x 0.70 x 0.30 / 0.10² = 81 participants. If you have no prior view of the rate, the worst case (50%) needs 97. Narrowing the interval to ±5 points when you expect 80% takes 246. This is why pilots that try to estimate many rates precisely become larger than the main trial's first cohort: choose the one or two rates that decide whether to proceed, and set progression thresholds ahead of time.

For problem detection, the chance of seeing at least one participant with a problem of probability π among n is 1 − (1 − π)ⁿ, so n = ln(1 − confidence) / ln(1 − π). A problem that affects 1 in 20 participants needs 59 for 95% confidence; 1 in 10 needs 29; 1 in 5 needs 14; 1 in 100 needs 299. Seeing a problem needs an observer, so the pilot should record deviations, queries and participant comments systematically.

Rules of thumb

Published pilot sample size guidance

Different recommendations answer different questions. Always state which one you are using.

SourceRecommendationAim
Julious, Pharm Stat 200512 per group (24 in total)Minimum with usable precision for the mean and the SD, when no prior information exists.
Browne, Stat Med 1995Use an upper confidence limit of the pilot SD, rather than the point estimate (80% is the level commonly cited)Protect the main trial against an underestimated SD.
Teare et al., Trials 2014About 35 per group (70 in total) for a continuous outcomeEstimate the SD with little further gain from more participants. More are needed to estimate a binary outcome rate.
Whitehead et al., SMMR 2016Stepped by planned standardised effect size; for a 90% powered main trial, roughly 75, 25, 15 and 10 per arm for effects up to about 0.1, 0.1 to 0.3, 0.3 to 0.7 and above 0.7Minimise the combined sample of pilot and main trial. Check the paper's table before citing in a protocol.
Viechtbauer et al., J Clin Epidemiol 2015n = ln(1 - confidence) / ln(1 - π), for example 59 for π = 5% and 95% confidenceDetect design or procedural problems.

The Whitehead figures are summarised from secondary sources; confirm them in the original paper.

Common mistakes

What goes wrong with pilot sample sizes

These are the mistakes that most often undermine a pilot or the main trial that follows it.

Calling the study a pilot to avoid a sample size

A pilot still needs a justification: a stated purpose, a formula or a precision target, and a plan for what happens next. "Pilot, n = 30" with no reason invites the question from funders and ethics committees. Use the modes above and write the reasoning down, and see how to run a pilot study for the wider process.

Using significance as a progression criterion

Progression criteria should be about feasibility (recruitment, retention, adherence, data completeness), with thresholds set in advance, often as a traffic light with a go, amend or stop zone. Statistical significance of the treatment effect in a pilot is not one of them.

Treating the pilot SD as exact

The table above shows how wide the uncertainty is. If the main trial is expensive, spend more on the pilot or build in a sample-size re-estimation point, which is cheaper than an underpowered main trial. The power analysis calculator shows what a given n can detect if the SD turns out higher.

Dropping the pilot data from the main trial without a plan

An internal pilot (the first part of the main trial) can be rolled into the final analysis; an external pilot generally cannot. Decide before the pilot starts, because it changes the protocol, the consent forms and the data system.

Pilot progression criteria (demo data)

Consented / approached

41%

Retained at week 8

79%

Diary completion

92%

Decision

Amend

Recruitment: threshold 35%, stop below 20%41/100

Go

Retention: threshold 85%, stop below 70%79/100

Amend

Adherence: threshold 80%, stop below 60%92/100

Go

Thresholds were written into the protocol before the first participant was approached.

Pilot the data collection as well as the intervention

Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.

Build your study free

From the number to the protocol

How this feeds the protocol, the SAP and the study build

Write the pilot sample size as a justification of purpose: "Twenty-four participants (12 per group) will be randomized. This provides a 95% confidence interval for the SD of the primary outcome of 0.77 to 1.42 times the observed value, and the upper 80% limit will be used to size the main trial. The proportion of eligible people consenting and the 8-week retention rate will be reported with 95% confidence intervals, judged against the progression criteria in Table 3." The SAP should say that no hypothesis test of the treatment effect is planned, or that any estimate is descriptive. The how to write a clinical trial protocol guide covers the surrounding sections.

In Capture, a pilot can use the same system as the main trial, so what is learned is what will be used. The study team builds the eCRFs with edit checks, defines the visit schedule with windows, and tests every form and skip rule in the free sandbox before the first participant is enrolled. During the pilot, query rates, missing-data patterns and visit-window deviations show which parts of the data collection need to change, the monitoring dashboard reports progress against the thresholds you set, and ePRO diaries on participants' phones give completion rates for the questionnaire burden question. The field-level audit trail records each change with a reason, and CSV or Excel exports with a data dictionary give the statistician the SD and rates. See EDC for pilot clinical trials and EDC for feasibility studies for the setup, and the two-means and two-proportions calculators for the main trial that follows.

Before the pilot starts

Pilot study sample size checklist

Purpose stated

SD estimate, feasibility rates, problem detection, or a mix, written as objectives.

Formula or rule cited

Julious 12 per group, a precision target, or n = ln(1 - c) / ln(1 - π), with its source.

Progression criteria set

Go, amend and stop thresholds for recruitment, retention and adherence.

No efficacy test planned

Treatment effect reported descriptively, with a confidence interval, if at all.

Upper limit for the SD

Main trial sized on the 80% upper limit or sensitivity range, not the point estimate.

Fate of pilot data decided

External (separate) or internal (rolled into the main trial) before it starts.

FAQ

Questions teams ask before they switch

Something not covered here? Ask us directly.

How many participants do I need for a pilot study?

It depends on the purpose. A common minimum for a two-arm pilot with no prior information is 12 per group (Julious, 2005). Estimating the SD well points to about 35 per group (Teare et al., 2014). Estimating a feasibility rate to within ±10 points at 95% confidence needs about 60 to 100 depending on the rate.

Where does the 12 per group rule come from?

Julious (Pharmaceutical Statistics, 2005) argued that 12 per group balances feasibility, the precision of the mean and variance, and regulatory considerations, as a minimum when no other information exists. With 12 per group the 95% CI for the SD is about 0.77 to 1.42 times the observed SD.

Should a pilot study have a power calculation?

No. A pilot is not designed to test the treatment effect. Justify its size by the precision it gives for a feasibility rate or a design parameter such as the SD, or by the chance of detecting a problem.

Can I use the effect size from my pilot to size the main trial?

It is not recommended. A small pilot estimates the effect too imprecisely and selecting on a promising result biases it upward. Use a clinically important difference, and use the pilot for the SD and for feasibility.

How do I use the upper confidence limit of the SD?

Multiply the pilot SD by the upper-limit factor from the chi-square distribution (for example 1.16 for the upper 80% limit with 12 per group) and size the main trial on that value. The calculator shows the main-trial n under the point estimate, the 80% limit and the 95% limit.

What is a progression criterion?

A pre-specified threshold for a feasibility measure, such as consent rate, retention or adherence, that decides whether the main trial goes ahead, goes ahead after changes, or stops. The CONSORT extension for pilot and feasibility trials asks for them to be reported.

How many participants are needed to detect a problem?

n = ln(1 − confidence) / ln(1 − π), where π is the chance a participant hits the problem. For π = 5% and 95% confidence it is 59; for 10% it is 29 (Viechtbauer et al., 2015).

Is this calculator validated?

No. It is a free planning aid. The feasibility-rate mode uses the simple Wald interval, which is rough for rates near 0% or 100%. A statistician should confirm the pilot design.

Pilot your data collection before the main trial

Build and test every form, visit window and ePRO diary in a free sandbox. No credit card. You pay only when you go live.

Build your study free