When whole clinics, schools or villages are randomised, participants in the same cluster resemble each other and the usual sample size is too small. Enter the ICC and cluster size to get the design effect, clusters per arm and total participants.
Free sandbox · No credit card · 21 CFR Part 11 aligned
Clusters per arm
30
Analysed participants
1,200
Design effect
1.95
44 clusters per arm
30 clusters per arm
22 clusters per arm
19 clusters per arm
What this calculator does
Free tool
Choose a continuous or binary outcome, enter the effect you want to detect, the ICC and either the mean cluster size or the number of clusters you can afford. It runs in your browser; nothing is saved or sent.
Enrol about 23 per cluster to allow for 10% dropout, or 1,380 in total.
Design effect 1.95. The same trial randomised by individual needs 294 per arm (588 in total), so clustering multiplies the participants by about 2.04.
DE = 1 + ((0.00² + 1) × 20 − 1) × 0.05 = 1.950; k = n_ind × DE / m = 293.2 × 1.950 / 20
Adding participants to each cluster has diminishing returns: with this ICC the clusters per arm can never fall below 14.7.
Planning aid and reference only, not validated software. Two arms, equal numbers of clusters, a continuous or binary outcome analysed at cluster level or with a method that accounts for clustering. The design effect is 1 + {(CV² + 1) m − 1} ρ (Eldridge, Ashby and Kerry 2006; CV 0 gives Donner and Klar's 1 + (m − 1) ρ). The t correction uses 2(k − 1) degrees of freedom, exact in form for means and an approximation for proportions. Dropout inflates cluster size only; losing whole clusters is not modelled. Take the ICC from a similar trial with the same outcome and setting, and have a statistician confirm the final number.
The method
Participants in the same clinic, ward, school or village share staff, environment and often catchment, so their outcomes are more alike than outcomes of people picked at random. The intracluster correlation coefficient (ICC, ρ) measures that similarity: the share of the total variance in the outcome that lies between clusters. At ρ = 0 clusters are irrelevant and each participant contributes a full unit of information. At ρ = 1 everyone in a cluster is identical and a whole cluster is worth one participant. Real values for health outcomes are usually small, often between 0.01 and 0.05, but they matter because they are multiplied by the cluster size.
The design effect converts that into a number you can use: DE = 1 + (m − 1) ρ, a result in Kish’s survey work that Donner and Klar brought into cluster randomised trials. It is the factor by which the variance, and so the required sample size, grows compared with randomising individuals. With m = 20 and ρ = 0.05, DE = 1 + 19 × 0.05 = 1.95: nearly twice as many participants. With m = 100 the same ICC gives 5.95. Divide the inflated sample size by m and you have the number of clusters.
Take a trial comparing an event rate of 30% in the control arm with 20% in the intervention arm, two-sided alpha 0.05 and 80% power. Randomised by individual, the pooled-variance formula gives 293.2, so 294 per arm (588 in total). With 20 participants analysed per cluster and ρ = 0.05, you need 294 × 1.95 = 573 per arm, which is 573 / 20 = 28.7, so 29 clusters per arm (580 participants per arm, 1,160 in total). The calculator then adds a small-sample correction and shows 30 clusters per arm, 1,200 participants. If 10% of participants will not provide data, enrol about 23 per cluster, 1,380 in all.
The calculator first computes the per-arm size you would need if individuals were randomised, using the normal-approximation formula for two means, n = 2σ²(z₁₋α/₂ + z₁₋β)² / δ², or for two proportions the pooled-variance formula, the same ones as in the two-means and two-proportions calculators. Use those pages when you want unequal allocation, a continuity correction or a what-if table, and the sample size calculator for the common two-proportion case. The concepts are in the sample size calculation glossary entry.
Sensitivity
30% vs 20%, two-sided alpha 0.05, 80% power, equal cluster sizes, normal formula without the small-sample correction. Cells show the design effect and clusters per arm. Individual randomisation needs 294 per arm.
| ICC | 10 per cluster | 20 per cluster | 50 per cluster | 100 per cluster |
|---|---|---|---|---|
| 0.01 | DE 1.09, 32 | DE 1.19, 18 | DE 1.49, 9 | DE 1.99, 6 |
| 0.02 | DE 1.18, 35 | DE 1.38, 21 | DE 1.98, 12 | DE 2.98, 9 |
| 0.05 | DE 1.45, 43 | DE 1.95, 29 | DE 3.45, 21 | DE 5.95, 18 |
| 0.10 | DE 1.90, 56 | DE 2.90, 43 | DE 5.90, 35 | DE 10.90, 32 |
Illustrative values from the calculator. At ICC 0.05, going from 20 to 100 per cluster saves 11 clusters per arm but raises the analysed sample from 580 to 1,800 per arm. Several of these cells fall under 10 clusters per arm, where the analysis becomes fragile.
Unequal clusters
Real clusters are never the same size. Practices have different list sizes, wards different admissions, schools different rolls. Unequal sizes cost power, because the analysis gives each cluster a different weight. Eldridge, Ashby and Kerry (2006) showed that you only need the mean cluster size m and the coefficient of variation of cluster size (CV, the SD of cluster size divided by the mean): DE = 1 + {(CV² + 1) m − 1} ρ. With CV = 0 this is the equal-size formula.
The same paper found that when the CV is below about 0.23 the adjustment is negligible, and that for many trials randomising UK general practices the CV is expected to be about 0.65. In our example (m = 20, ρ = 0.05), moving the CV from 0 to 0.25, 0.5, 0.65 and 1 raises the design effect from 1.95 to 2.01, 2.20, 2.37 and 2.95 and the clusters per arm from 29 to 30, 33, 35 and 44. Ignoring variation in a setting with a CV of 1 would leave the trial about a third short of clusters.
You can estimate the CV from the list of clusters you intend to approach: take their expected sizes, compute the SD and divide by the mean. If you can only guess, run the calculator at two or three plausible values. You can also reduce the problem in the design, by setting a minimum and a maximum for cluster size, or by recruiting a fixed number per cluster, which is easier when participants are recruited prospectively than when whole populations are included.
The real constraint
The number of clusters, not the number of participants, usually limits a cluster trial. The design effect formula shows why: clusters per arm = n × (1 + (m − 1) ρ) / m = n × ρ + n × (1 − ρ) / m. As m grows the second term vanishes, leaving a floor of n × ρ clusters per arm. For the example, 294 × 0.05 = 14.7, so 15 clusters per arm is the best case with infinitely large clusters, and getting close needs a huge m. With the CV adjustment the floor is n × ρ × (CV² + 1).
Choose “Number of clusters” in the calculator to work the other way round: if you can recruit 20 clusters per arm, the required mean size is m = n(1 − ρ) / (k − nρ) = 294 × 0.95 / (20 − 14.7) = 52 per cluster (64 with the small-sample correction). With 15 clusters per arm no cluster size works with the correction, and with 30 clusters you need only about 20 per cluster. The calculator says when your number of clusters is below the floor instead of returning an absurd cluster size.
There is no single rule for the minimum number of clusters. Four per arm is the theoretical floor for a test based on re-randomisation to reach p < 0.05. In practice, standard analyses that rely on large-sample approximations are unreliable below about 10 clusters per arm and small-sample bias fades from about 14 per arm, so many methodologists recommend at least 10 to 15 per arm or a small-sample method such as a cluster-level analysis with a t-test, a permutation test or a corrected mixed model. The calculator’s optional t correction replaces the normal quantiles with t quantiles on 2(k − 1) degrees of freedom, which is the usual first adjustment (Hayes and Moulton). It is exact in form for means and an approximation for proportions. Treat it as a minimum, not a guarantee.
| What it means | |
|---|---|
| Fewer than 4 | No valid test possible |
| 4 to 9 | Fragile; plan a small-sample analysis |
| 10 to 14 | Workable with a t correction or permutation test |
| 15 or more | Standard cluster-adjusted analysis is usually reasonable |
Build the eCRFs, visit schedule and site structure in the free sandbox. No credit card, pay only when you go live.
The input that matters most
The ICC is the hardest input to know. Take it from a previous trial or a routine data set that uses the same outcome in a similar type of cluster, and prefer an estimate from a large study: ICCs from small trials are noisy and published ones are biased towards what was convenient to report. Report the source in the protocol. The CONSORT extension for cluster randomised trials asks for the ICC used in the sample size calculation to be stated, and for the observed ICC to be reported afterwards.
Because the design effect rises linearly with the ICC, an estimate that is too low is expensive. A sensible habit is to calculate at your best estimate and again at one that is double, or at the upper end of a confidence interval, and to decide beforehand what you would do if the lower-ICC world turns out to be the true one. An ICC measured at the cluster level of your own outcome, even in a pilot of a few clusters, is more reliable than one borrowed from a different outcome.
Outcomes that are processes of care measured on a whole clinic, such as vaccination or prescribing, tend to have larger ICCs than patient-level clinical measurements. Binary outcomes with a rate near 0 or 1 behave differently from a continuous measure. If you are unsure, bring a statistician in early, and use the pilot study sample size calculator to plan a feasibility phase, since cluster designs with few clusters, stratification, matching, stepped-wedge or cross-over structure need more than this formula. A pragmatic trial often randomises clusters, which is where this question usually starts.
In Capture
Cluster trials are won or lost on clean cluster identifiers and complete outcome data. Capture is built for multi-site studies, with site-level participant numbering, a separate site coordinator portal and by-site filtering and exports, so each row of the exported data carries the site it came from. Edit checks raise an auto-query when a value breaks a range or custom rule, and every change is held in the field-level audit trail. Allocating the clusters is a design decision for your statistician, so generate and document the allocation list of clusters separately, for example with the free randomization list generator, and keep it with the protocol.
Clusters open
26 / 30
Participants entered
412
Open queries
9
13 of 15 open
13 of 15 open
recruitment lagging
For the protocol and the SAP
A reviewer should be able to rebuild your number from the protocol. State each of these.
Cluster-level or individual-level analysis that accounts for clustering; never an unadjusted individual analysis.
Difference or rates, SD, alpha, power and the n per arm that individual randomisation would need.
Value, outcome, setting and why it applies; a sensitivity at a higher value.
Expected recruited or analysed size per cluster, and the CV if sizes vary.
The formula used and the resulting number of clusters and participants.
Participant dropout inflates cluster size; plan separately for the loss of whole clusters.
The analysis you will use if clusters per arm is below about 15.
Restriction, stratification or matching, and who allocates and when.
The design effect is the factor by which the sample size of an individually randomised trial must be multiplied to allow for clustering. For equal cluster sizes it is DE = 1 + (m − 1) × ICC, where m is the cluster size and ICC the intracluster correlation coefficient. For unequal sizes it is 1 + {(CV² + 1) m − 1} × ICC.
Work out the per-arm sample size you would need for individual randomisation, multiply by the design effect, divide by the mean cluster size and round up. For 294 per arm, 20 per cluster and an ICC of 0.05 that is 294 × 1.95 / 20 = 28.7, so 29 clusters per arm before any small-sample correction.
For many health outcomes, ICCs reported in cluster trials are small, often between 0.01 and 0.05, but they vary widely by outcome and setting and process-of-care outcomes tend to be higher. Take the value from a similar study with the same outcome, state its source and test a higher value.
Four per arm is the theoretical floor for a valid randomisation-based test, and standard analyses are unreliable below about 10 per arm. Many methodologists aim for at least 10 to 15 per arm, or plan a small-sample analysis such as a permutation test. It is a rule of thumb, not a regulatory requirement.
Because the design effect grows with cluster size. As m increases, the clusters per arm approach a floor of n × ICC × (CV² + 1) and never go below it. Past a certain size, extra participants per cluster add cost and save almost no clusters.
Enter the mean cluster size and the coefficient of variation (SD divided by the mean). Eldridge, Ashby and Kerry found the adjustment negligible below a CV of about 0.23 and expected about 0.65 for many UK general practice trials. Ignoring a large CV leaves the trial short of clusters.
No. It covers a two-arm parallel cluster design with equal numbers of clusters per arm. Stepped-wedge, cross-over, matched-pair and multi-level designs have different design effects and need a statistician or specialised software.
No. It is a free planning aid that applies published closed-form formulas and a standard small-sample adjustment. A statistician should confirm the ICC, the design effect and the final number before the protocol is submitted.
Keep exploring
Two-proportions sample size calculator
The individual-level number for a binary outcome.
Two-means sample size calculator
The individual-level number for a continuous outcome.
Sample size calculator for clinical trials
The common two-proportion case.
Power analysis calculator
What a fixed number of clusters can detect.
Sample size calculation (glossary)
The concepts behind the formulas.
Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.