Free tool · Sample size · Cluster randomisedUpdated October 10, 2026

Cluster randomized trial sample size calculator with ICC

When whole clinics, schools or villages are randomised, participants in the same cluster resemble each other and the usual sample size is too small. Enter the ICC and cluster size to get the design effect, clusters per arm and total participants.

  • Design effect from ICC and cluster size
  • Unequal cluster sizes (CV)
  • Clusters per arm or cluster size, your choice

Free sandbox · No credit card · 21 CFR Part 11 aligned

Demo: 30% vs 20% event rate, ICC 0.05, 20 per cluster, 80% power

Clusters per arm

30

Analysed participants

1,200

Design effect

1.95

10 per cluster44/60

44 clusters per arm

20 per cluster30/60

30 clusters per arm

50 per cluster22/60

22 clusters per arm

100 per cluster19/60

19 clusters per arm

Randomised by individual the same trial needs 294 per arm. Bigger clusters save clusters, not participants.

What this calculator does

  • In a cluster randomised trial the design effect is DE = 1 + (m − 1) × ICC, where m is the cluster size and ICC the intracluster correlation coefficient. Multiply the sample size of the same trial randomised by individual by DE.
  • Clusters per arm = n(individual, per arm) × DE / m, rounded up. Worked example: 30% vs 20% at 80% power needs 294 per arm by individual; with 20 per cluster and an ICC of 0.05, DE is 1.95 and you need 30 clusters per arm (29 before the small-sample correction), about 1,200 analysed participants.
  • If cluster sizes vary, use DE = 1 + {(CV² + 1) m − 1} × ICC, where CV is the coefficient of variation of cluster size (Eldridge, Ashby and Kerry 2006). With a CV of 0.65, common in general practice trials, the same example needs 36 clusters per arm instead of 30.
  • Larger clusters give diminishing returns: with an ICC of 0.05 you can never get below about 15 clusters per arm for this example, whatever the cluster size. Past roughly 50 per cluster you add participants and save almost no clusters.
  • Few clusters is the real risk. Four per arm is the floor for any valid test, and below about 10 per arm standard analyses are unreliable, so a trial with 6 large clusters per arm is usually weaker than it looks.

Free tool

Calculate clusters per arm and total participants

Choose a continuous or binary outcome, enter the effect you want to detect, the ICC and either the mean cluster size or the number of clusters you can afford. It runs in your browser; nothing is saved or sent.

Primary outcome
%
%
You know
%
Few-clusters t correction
Clusters per arm
30
Total clusters
60
Participants
1,200
analysed, both arms

Enrol about 23 per cluster to allow for 10% dropout, or 1,380 in total.

Design effect 1.95. The same trial randomised by individual needs 294 per arm (588 in total), so clustering multiplies the participants by about 2.04.

DE = 1 + ((0.00² + 1) × 20 − 1) × 0.05 = 1.950; k = n_ind × DE / m = 293.2 × 1.950 / 20

Adding participants to each cluster has diminishing returns: with this ICC the clusters per arm can never fall below 14.7.

Planning aid and reference only, not validated software. Two arms, equal numbers of clusters, a continuous or binary outcome analysed at cluster level or with a method that accounts for clustering. The design effect is 1 + {(CV² + 1) m − 1} ρ (Eldridge, Ashby and Kerry 2006; CV 0 gives Donner and Klar's 1 + (m − 1) ρ). The t correction uses 2(k − 1) degrees of freedom, exact in form for means and an approximation for proportions. Dropout inflates cluster size only; losing whole clusters is not modelled. Take the ICC from a similar trial with the same outcome and setting, and have a statistician confirm the final number.

The method

Why clustering inflates the sample size, and the formula

Participants in the same clinic, ward, school or village share staff, environment and often catchment, so their outcomes are more alike than outcomes of people picked at random. The intracluster correlation coefficient (ICC, ρ) measures that similarity: the share of the total variance in the outcome that lies between clusters. At ρ = 0 clusters are irrelevant and each participant contributes a full unit of information. At ρ = 1 everyone in a cluster is identical and a whole cluster is worth one participant. Real values for health outcomes are usually small, often between 0.01 and 0.05, but they matter because they are multiplied by the cluster size.

The design effect converts that into a number you can use: DE = 1 + (m − 1) ρ, a result in Kish’s survey work that Donner and Klar brought into cluster randomised trials. It is the factor by which the variance, and so the required sample size, grows compared with randomising individuals. With m = 20 and ρ = 0.05, DE = 1 + 19 × 0.05 = 1.95: nearly twice as many participants. With m = 100 the same ICC gives 5.95. Divide the inflated sample size by m and you have the number of clusters.

Take a trial comparing an event rate of 30% in the control arm with 20% in the intervention arm, two-sided alpha 0.05 and 80% power. Randomised by individual, the pooled-variance formula gives 293.2, so 294 per arm (588 in total). With 20 participants analysed per cluster and ρ = 0.05, you need 294 × 1.95 = 573 per arm, which is 573 / 20 = 28.7, so 29 clusters per arm (580 participants per arm, 1,160 in total). The calculator then adds a small-sample correction and shows 30 clusters per arm, 1,200 participants. If 10% of participants will not provide data, enrol about 23 per cluster, 1,380 in all.

Where the individual sample size comes from

The calculator first computes the per-arm size you would need if individuals were randomised, using the normal-approximation formula for two means, n = 2σ²(z₁₋α/₂ + z₁₋β)² / δ², or for two proportions the pooled-variance formula, the same ones as in the two-means and two-proportions calculators. Use those pages when you want unequal allocation, a continuity correction or a what-if table, and the sample size calculator for the common two-proportion case. The concepts are in the sample size calculation glossary entry.

Sensitivity

Clusters per arm by ICC and cluster size

30% vs 20%, two-sided alpha 0.05, 80% power, equal cluster sizes, normal formula without the small-sample correction. Cells show the design effect and clusters per arm. Individual randomisation needs 294 per arm.

ICC10 per cluster20 per cluster50 per cluster100 per cluster
0.01DE 1.09, 32DE 1.19, 18DE 1.49, 9DE 1.99, 6
0.02DE 1.18, 35DE 1.38, 21DE 1.98, 12DE 2.98, 9
0.05DE 1.45, 43DE 1.95, 29DE 3.45, 21DE 5.95, 18
0.10DE 1.90, 56DE 2.90, 43DE 5.90, 35DE 10.90, 32

Illustrative values from the calculator. At ICC 0.05, going from 20 to 100 per cluster saves 11 clusters per arm but raises the analysed sample from 580 to 1,800 per arm. Several of these cells fall under 10 clusters per arm, where the analysis becomes fragile.

Unequal clusters

When cluster sizes vary: the coefficient of variation

Real clusters are never the same size. Practices have different list sizes, wards different admissions, schools different rolls. Unequal sizes cost power, because the analysis gives each cluster a different weight. Eldridge, Ashby and Kerry (2006) showed that you only need the mean cluster size m and the coefficient of variation of cluster size (CV, the SD of cluster size divided by the mean): DE = 1 + {(CV² + 1) m − 1} ρ. With CV = 0 this is the equal-size formula.

The same paper found that when the CV is below about 0.23 the adjustment is negligible, and that for many trials randomising UK general practices the CV is expected to be about 0.65. In our example (m = 20, ρ = 0.05), moving the CV from 0 to 0.25, 0.5, 0.65 and 1 raises the design effect from 1.95 to 2.01, 2.20, 2.37 and 2.95 and the clusters per arm from 29 to 30, 33, 35 and 44. Ignoring variation in a setting with a CV of 1 would leave the trial about a third short of clusters.

You can estimate the CV from the list of clusters you intend to approach: take their expected sizes, compute the SD and divide by the mean. If you can only guess, run the calculator at two or three plausible values. You can also reduce the problem in the design, by setting a minimum and a maximum for cluster size, or by recruiting a fixed number per cluster, which is easier when participants are recruited prospectively than when whole populations are included.

The real constraint

Minimum number of clusters, and why more participants cannot fix too few

The number of clusters, not the number of participants, usually limits a cluster trial. The design effect formula shows why: clusters per arm = n × (1 + (m − 1) ρ) / m = n × ρ + n × (1 − ρ) / m. As m grows the second term vanishes, leaving a floor of n × ρ clusters per arm. For the example, 294 × 0.05 = 14.7, so 15 clusters per arm is the best case with infinitely large clusters, and getting close needs a huge m. With the CV adjustment the floor is n × ρ × (CV² + 1).

Choose “Number of clusters” in the calculator to work the other way round: if you can recruit 20 clusters per arm, the required mean size is m = n(1 − ρ) / (k − nρ) = 294 × 0.95 / (20 − 14.7) = 52 per cluster (64 with the small-sample correction). With 15 clusters per arm no cluster size works with the correction, and with 30 clusters you need only about 20 per cluster. The calculator says when your number of clusters is below the floor instead of returning an absurd cluster size.

There is no single rule for the minimum number of clusters. Four per arm is the theoretical floor for a test based on re-randomisation to reach p < 0.05. In practice, standard analyses that rely on large-sample approximations are unreliable below about 10 clusters per arm and small-sample bias fades from about 14 per arm, so many methodologists recommend at least 10 to 15 per arm or a small-sample method such as a cluster-level analysis with a t-test, a permutation test or a corrected mixed model. The calculator’s optional t correction replaces the normal quantiles with t quantiles on 2(k − 1) degrees of freedom, which is the usual first adjustment (Hayes and Moulton). It is exact in form for means and an approximation for proportions. Treat it as a minimum, not a guarantee.

How many clusters per arm? A practical reading
What it means
Fewer than 4No valid test possible
4 to 9Fragile; plan a small-sample analysis
10 to 14Workable with a t correction or permutation test
15 or moreStandard cluster-adjusted analysis is usually reasonable
Rules of thumb from the methods literature, not a regulatory threshold

Run the cluster trial in one place

Build the eCRFs, visit schedule and site structure in the free sandbox. No credit card, pay only when you go live.

Build your study free

The input that matters most

Where the ICC comes from, and how to be cautious about it

The ICC is the hardest input to know. Take it from a previous trial or a routine data set that uses the same outcome in a similar type of cluster, and prefer an estimate from a large study: ICCs from small trials are noisy and published ones are biased towards what was convenient to report. Report the source in the protocol. The CONSORT extension for cluster randomised trials asks for the ICC used in the sample size calculation to be stated, and for the observed ICC to be reported afterwards.

Because the design effect rises linearly with the ICC, an estimate that is too low is expensive. A sensible habit is to calculate at your best estimate and again at one that is double, or at the upper end of a confidence interval, and to decide beforehand what you would do if the lower-ICC world turns out to be the true one. An ICC measured at the cluster level of your own outcome, even in a pilot of a few clusters, is more reliable than one borrowed from a different outcome.

Outcomes that are processes of care measured on a whole clinic, such as vaccination or prescribing, tend to have larger ICCs than patient-level clinical measurements. Binary outcomes with a rate near 0 or 1 behave differently from a continuous measure. If you are unsure, bring a statistician in early, and use the pilot study sample size calculator to plan a feasibility phase, since cluster designs with few clusters, stratification, matching, stepped-wedge or cross-over structure need more than this formula. A pragmatic trial often randomises clusters, which is where this question usually starts.

In Capture

Collect the cluster-level data the analysis needs

Cluster trials are won or lost on clean cluster identifiers and complete outcome data. Capture is built for multi-site studies, with site-level participant numbering, a separate site coordinator portal and by-site filtering and exports, so each row of the exported data carries the site it came from. Edit checks raise an auto-query when a value breaks a range or custom rule, and every change is held in the field-level audit trail. Allocating the clusters is a design decision for your statistician, so generate and document the allocation list of clusters separately, for example with the free randomization list generator, and keep it with the protocol.

  • Site-level participant numbering and by-site exports.
  • Edit checks and auto-queries on outcome fields.
  • CSV and Excel exports with a data dictionary, and SDTM export.
EDC with randomization
Cluster trial status by arm (demo data)

Clusters open

26 / 30

Participants entered

412

Open queries

9

Control clusters13/15

13 of 15 open

Intervention clusters13/15

13 of 15 open

Clusters under 10 participants8/26

recruitment lagging

Demo numbers only. Watching cluster size and cluster count during the trial tells you early whether the design effect you assumed still holds.

For the protocol and the SAP

What to write down for a cluster trial sample size

A reviewer should be able to rebuild your number from the protocol. State each of these.

Unit of randomisation and unit of analysis

Cluster-level or individual-level analysis that accounts for clustering; never an unadjusted individual analysis.

Effect size and individual sample size

Difference or rates, SD, alpha, power and the n per arm that individual randomisation would need.

ICC with its source

Value, outcome, setting and why it applies; a sensitivity at a higher value.

Mean cluster size and CV

Expected recruited or analysed size per cluster, and the CV if sizes vary.

Design effect and clusters per arm

The formula used and the resulting number of clusters and participants.

Dropout and cluster loss

Participant dropout inflates cluster size; plan separately for the loss of whole clusters.

Small-sample method

The analysis you will use if clusters per arm is below about 15.

Randomisation of clusters

Restriction, stratification or matching, and who allocates and when.

FAQ

Questions teams ask before they switch

Something not covered here? Ask us directly.

What is the design effect in a cluster randomized trial?

The design effect is the factor by which the sample size of an individually randomised trial must be multiplied to allow for clustering. For equal cluster sizes it is DE = 1 + (m − 1) × ICC, where m is the cluster size and ICC the intracluster correlation coefficient. For unequal sizes it is 1 + {(CV² + 1) m − 1} × ICC.

How do I calculate the number of clusters per arm?

Work out the per-arm sample size you would need for individual randomisation, multiply by the design effect, divide by the mean cluster size and round up. For 294 per arm, 20 per cluster and an ICC of 0.05 that is 294 × 1.95 / 20 = 28.7, so 29 clusters per arm before any small-sample correction.

What is a typical ICC?

For many health outcomes, ICCs reported in cluster trials are small, often between 0.01 and 0.05, but they vary widely by outcome and setting and process-of-care outcomes tend to be higher. Take the value from a similar study with the same outcome, state its source and test a higher value.

What is the minimum number of clusters per arm?

Four per arm is the theoretical floor for a valid randomisation-based test, and standard analyses are unreliable below about 10 per arm. Many methodologists aim for at least 10 to 15 per arm, or plan a small-sample analysis such as a permutation test. It is a rule of thumb, not a regulatory requirement.

Why can I not just make the clusters bigger?

Because the design effect grows with cluster size. As m increases, the clusters per arm approach a floor of n × ICC × (CV² + 1) and never go below it. Past a certain size, extra participants per cluster add cost and save almost no clusters.

What if my cluster sizes are unequal?

Enter the mean cluster size and the coefficient of variation (SD divided by the mean). Eldridge, Ashby and Kerry found the adjustment negligible below a CV of about 0.23 and expected about 0.65 for many UK general practice trials. Ignoring a large CV leaves the trial short of clusters.

Does this calculator handle stepped-wedge, crossover or matched cluster designs?

No. It covers a two-arm parallel cluster design with equal numbers of clusters per arm. Stepped-wedge, cross-over, matched-pair and multi-level designs have different design effects and need a statistician or specialised software.

Is this calculator validated?

No. It is a free planning aid that applies published closed-form formulas and a standard small-sample adjustment. A statistician should confirm the ICC, the design effect and the final number before the protocol is submitted.

Plan the number, then build the study

Build the eCRFs, visit schedule and randomization in the free sandbox with every feature. No credit card. You pay only when you go live.

Build your study free