ANALYSIS CLINIC · CASE 006

CASE STATUS: DIAGNOSED

I may have chosen the wrong analysis · Mixed models · R/lme4

My ICC Is Small. Do I Still Need Multilevel Modeling?

A small ICC does not automatically justify treating observations as independent. The defensible choice depends on the design, cluster sizes, level of the predictor or treatment, and what you want to estimate.

01

Symptoms

Students are nested in classrooms, employees in organizations, or repeated observations within people. You fit an empty random-intercept model and obtain an ICC such as .01 or .03. Someone suggests ordinary regression because the ICC is below .05, or asks why you need a multilevel model when the between-cluster variance is small.

An ICC describes one aspect of the fitted covariance structure. It is not a universal pass/fail test for independence, an adjustment strategy, or the scientific relevance of the higher level.

02

What This Usually Means

For a continuous outcome in a two-level random-intercept model, the ICC is between-cluster variance / (between-cluster variance + residual variance). Under that model it is the correlation between two observations from the same cluster. A small ICC indicates modest modeled similarity for an individual pair; many such pairs can still affect uncertainty about an average or a cluster-level contrast.

For equal cluster size m and a common ICC rho, the design effect for the grand mean is 1 + (m - 1) × rho, relative to independent observations with the same marginal variance. With rho = .02 and m = 50, it is 1.98, corresponding to an SE multiplier of about 1.41.

This is a result for the mean under an exchangeable, equal-size model. It is not a universal multiplier for regression slopes, a sample-size rule for every analysis, or a threshold that selects a method. Unequal or informative cluster sizes, weights, covariate patterns, and different effects require additional analysis.

03

Common Causes

  1. A small pairwise correlation across large clusters. Modest dependence can accumulate across many observations within each cluster.
  2. An uncertain variance estimate. Few independent clusters or weak between-cluster information can make a small estimate imprecise; a boundary estimate does not prove a population variance is zero.
  3. The ICC belongs to a different model. An empty-model ICC and a covariate-adjusted residual ICC describe different decompositions. Both depend on the outcome, scale, sample, and specification.
  4. The scientific question concerns levels. Within-classroom and between-classroom associations, classroom predictors, or cross-level moderation require explicit level interpretation even when the outcome ICC is small.
  5. The dependence is more complicated. A random intercept cannot summarize slope variation, serial correlation, crossed memberships, or all forms of heterogeneity.
  6. The design determines the independent units. A classroom-randomized treatment has classroom-level assignment. A small outcome ICC does not turn students into independently randomized treatment units.

04

Run These Checks

  1. Map the sampling and assignment structure

    Identify the independent units, nesting or crossing, repeated measurements, treatment assignment level, strata, and weights. Check that cluster IDs identify the intended units uniquely.

    What it suggests: the design determines which dependence must be considered. A multilevel model, marginal model, cluster-aware regression, or design-based analysis may be appropriate; a point ICC cannot choose among them.

  2. Count clusters and inspect their sizes

    analysis_dat <- dat[complete.cases(dat[c("y", "cluster")]), ]
    analysis_dat$cluster <- droplevels(factor(analysis_dat$cluster))
    sizes <- table(analysis_dat$cluster)
    length(sizes)
    summary(as.numeric(sizes))
    

    What it suggests: report both observations and independent clusters. Thousands of observations in a handful of clusters do not supply thousands of independent cluster-level comparisons. Unequal sizes can change weighting and precision; average size alone can conceal the problem. This sample construction is illustrative, not a recommended final missing-data strategy.

  3. Identify what the ICC actually estimates

    library(lme4)
    empty <- lmer(y ~ 1 + (1 | cluster), data = analysis_dat)
    between <- as.numeric(VarCorr(empty)$cluster[1, 1])
    within <- sigma(empty)^2
    icc <- between / (between + within)
    icc
    isSingular(empty)
    

    What it suggests: this is an estimated empty-model ICC for a continuous outcome with an exchangeable random intercept. Review convergence and singularity separately. For logistic, count, or ordinal outcomes, state the scale and ICC definition rather than reusing this residual-variance calculation.

  4. Evaluate uncertainty and relevant covariance structure

    Use profile likelihood or a model-appropriate bootstrap for variance-component uncertainty when feasible. A confidence interval for the random-intercept standard deviation alone is not an ICC interval: the ICC also depends on residual variance. Investigate any singularity and whether the design supports random slopes, serial correlation, or additional levels.

    What it suggests: a small point estimate is not proof that dependence is negligible. A random-intercept ICC near zero does not rule out slope heterogeneity. In a random-slope model, within-cluster covariance depends on predictor values and cannot generally be represented by one constant ICC.

  5. State the focal comparison and separate levels

    For a numeric individual-level predictor, a within/between decomposition may help clarify the question:

    # Use a common analysis sample containing y, x, and cluster.
    analysis_x <- dat[complete.cases(dat[c("y", "x", "cluster")]), ]
    analysis_x$cluster <- droplevels(factor(analysis_x$cluster))
    analysis_x$x_mean <- ave(analysis_x$x, analysis_x$cluster, FUN = mean)
    analysis_x$x_within <- analysis_x$x - analysis_x$x_mean
    fit_levels <- lmer(y ~ x_within + x_mean + (1 | cluster),
                       data = analysis_x)
    

    What it suggests: the within-cluster and between-cluster coefficients answer different questions. This specification is a diagnostic starting point, not a universal causal model. Cluster means can be noisy, cluster size may matter, and random slopes or contextual variables may be needed.

  6. Compare uncertainty across defensible analyses

    Compare the focal estimate, SE or interval, assumptions, and inferential target using the same analysis sample. Consider a justified multilevel model, a marginal approach such as GEE, cluster-robust inference, or an appropriate survey/design-based analysis. For few clusters, use methods with suitable finite-sample treatment and examine effective degrees of freedom and leverage.

    What it suggests: agreement can support a sensitivity conclusion; disagreement needs explanation. A sandwich estimator changes uncertainty under a fitted mean model. It does not correct an inappropriate adjustment set, confounding, or a misspecified scientific comparison.

05

What Not to Do

  • Do not use .01, .05, a design effect of 2, or another cutoff as automatic permission to ignore clustering.
  • Do not require a significant random-intercept variance test before accounting for the sampling or assignment design.
  • Do not treat the number of individual records as the number of independent treatment units in a cluster-randomized study.
  • Do not multiply every regression SE by the grand-mean design-effect formula.
  • Do not assume cluster fixed effects resolve all residual dependence. They also absorb cluster-constant predictors, preventing separate estimation of those predictors in that specification.
  • Do not describe a zero estimated intercept variance as proof that random slopes or temporal dependence are absent.
  • Do not choose a method because it produces the smallest p-value.

06

Treatment Options

Model the multilevel structure

Use when: variance components, partial pooling, level-specific associations, random slopes, or cluster predictions matter to the question.

Tradeoff: specify and assess the mean and covariance model. Random-effects assumptions, sparse clusters, and limited cluster-level information can affect inference. Adding a random intercept alone does not guarantee that within/between differences or confounding are addressed.

Use cluster-robust regression inference

Use when: the regression mean model answers the question and within-cluster dependence is a nuisance rather than the object of study.

Tradeoff: inference depends on independent clusters and adequate information at that level. Small-sample corrections and effective degrees of freedom matter; they cannot manufacture independent comparisons or fix bias in the mean model.

Use a marginal or design-based approach

Use when: the target is a population-average association, or sampling weights, strata, and cluster assignment are central.

Tradeoff: use assumptions and finite-sample inference appropriate to the design. For nonlinear outcomes, marginal and conditional model coefficients can target different quantities.

Use cluster fixed effects where they fit the target

Use when: the question concerns within-cluster variation and controlling cluster-specific intercepts is defensible.

Tradeoff: cluster-constant effects are absorbed; serial correlation or other residual dependence may still require appropriate uncertainty estimation. Fixed effects are not interchangeable with partial pooling.

Justify a simpler independence analysis

Use when: design review, the target comparison, and sensitivity evidence support treating remaining dependence as practically negligible for the focal inference.

Tradeoff: document the evidence and uncertainty rather than relying on an ICC cutoff. A small ICC alone is insufficient justification, especially for cluster-level assignment or poorly estimated covariance structure.

Treatment principle: account for the design and target inference; the choice is not simply “small ICC versus multilevel model.”

07

Worked Example

This simulation generates a continuous outcome for 80 independent clusters, each with 50 observations. The generating random-intercept variance is .02 and residual variance is .98, so the population ICC is .02. The example estimates the grand mean, keeping the model comparison deliberately simple.

See it in R and Python

The R script simulates the study and fits both models. The Python script reads the same synthetic observations and uses the exact balanced-model REML formulas to reproduce the variance components and SEs. It requires only Python's standard library. Download the CSV beside the Python script before running it.

Using the same numeric seed in R and Python does not generally produce the same random sample. The shared CSV makes this a direct comparison. The Python formulas apply to this balanced, intercept-only model with a positive variance estimate; use a suitable mixed-model library for more general models.

suppressPackageStartupMessages(library(lme4))
set.seed(20261006)
J <- 80L
m <- 50L
rho <- 0.02
cluster <- factor(rep(seq_len(J), each = m))
u <- rnorm(J, sd = sqrt(rho))
y <- 10 + u[cluster] + rnorm(J * m, sd = sqrt(1 - rho))
dat <- data.frame(y, cluster)
naive <- lm(y ~ 1, data = dat)
clustered <- lmer(y ~ 1 + (1 | cluster), data = dat, REML = TRUE)
tau2 <- as.numeric(VarCorr(clustered)$cluster[1, 1])
sigma2 <- sigma(clustered)^2
icc <- tau2 / (tau2 + sigma2)
print(c(J = J, m = m, n = nrow(dat), generating_ICC = rho,
        fitted_ICC = icc, fitted_between_variance = tau2,
        fitted_within_variance = sigma2,
        naive_mean = coef(naive)[1], mixed_mean = fixef(clustered)[1],
        naive_SE = sqrt(vcov(naive)[1, 1]),
        mixed_SE = sqrt(vcov(clustered)[1, 1]),
        generating_design_effect = 1 + (m - 1) * rho,
        fitted_design_effect = 1 + (m - 1) * icc))
print(isSingular(clustered))
print(packageVersion('lme4'))
sessionInfo()

# Numerical and structural checks for the controlled example.
stopifnot(nrow(dat) == 4000L, nlevels(dat$cluster) == 80L,
          abs(coef(naive)[1] - fixef(clustered)[1]) < 1e-8,
          sqrt(vcov(clustered)[1, 1]) > sqrt(vcov(naive)[1, 1]),
          !isSingular(clustered), clustered@optinfo$conv$opt == 0)
"""Case 006: reproduce the R comparison using its shared synthetic CSV.

Download case-006-small-icc-data.csv beside this script, then run:
    python case-006-small-icc-multilevel.py
Or pass the CSV path as the first argument. Python standard library only.

For this balanced, intercept-only random-intercept model, interior REML
variance estimates have closed-form ANOVA expressions. This is not a
general mixed-model fitter for covariates, unequal sizes, or random slopes.
"""
import csv
import math
from pathlib import Path
import statistics
import sys

path = Path(sys.argv[1]) if len(sys.argv) > 1 else Path(__file__).with_name(
    "case-006-small-icc-data.csv"
)
groups = {}
with path.open(newline="", encoding="utf-8-sig") as source:
    for row in csv.DictReader(source):
        groups.setdefault(row["cluster"], []).append(float(row["y"]))
J = len(groups)
sizes = {len(values) for values in groups.values()}
if J < 2 or len(sizes) != 1:
    raise ValueError("This example requires balanced clusters and J > 1.")
m = sizes.pop()
if m < 2:
    raise ValueError("This example requires at least two observations per cluster.")
values = [y for group in groups.values() for y in group]
n = len(values)
mean = statistics.mean(values)
means = [statistics.mean(group) for group in groups.values()]
ms_within = sum(
    sum((y - statistics.mean(group)) ** 2 for y in group)
    for group in groups.values()
) / (n - J)
ms_between = m * sum((value - mean) ** 2 for value in means) / (J - 1)
tau2 = (ms_between - ms_within) / m
if tau2 <= 0:
    raise ValueError("Interior REML formula not applicable: variance is on boundary.")
sigma2 = ms_within
icc = tau2 / (tau2 + sigma2)
naive_se = math.sqrt(statistics.variance(values) / n)
mixed_se = math.sqrt((tau2 + sigma2 / m) / J)
for name, value in {
    "clusters": J, "cluster_size": m, "observations": n,
    "mean_both_models": mean, "between_variance_REML": tau2,
    "within_variance_REML": sigma2, "fitted_ICC": icc,
    "independence_SE": naive_se, "random_intercept_SE": mixed_se,
    "fitted_design_effect": 1 + (m - 1) * icc,
    "SE_ratio": mixed_se / naive_se,
}.items():
    print(f"{name}: {value:.8f}")
assert n == 4000 and J == 80 and m == 50
assert math.isclose(icc, 0.02612601, abs_tol=1e-7)
assert math.isclose(naive_se, 0.01606634, abs_tol=1e-8)
assert math.isclose(mixed_se, 0.02426445, abs_tol=1e-8)
print("All checks passed. Python", sys.version.split()[0])

Verified output — R 4.6.0, lme4 2.0-1: 4,000 observations in 80 clusters; fitted ICC = .02612601; between variance = .02698398; residual variance = 1.005856. Both models estimate the mean as 9.982148. The independence SE is .01606634 and the mixed-model SE is .02426445. The fit is not singular. The Python standard-library calculation reproduces the displayed ICC and SEs from the shared dataset; all Python assertions pass.

The generating design effect is 1.98. The fitted ICC gives a mean design effect of approximately 2.280175. These differ because the fitted variance components are estimates from one simulated sample. The estimated SE ratio is approximately 1.51; slight differences from the square root of the fitted design effect reflect finite-sample variance estimation.

Interpretation: equal point estimates do not imply equal uncertainty. A fitted ICC below .03 coexists with materially greater model-based uncertainty for the mean. This example does not establish the inflation for a within-cluster regression slope, make multilevel models mandatory for every small ICC, or validate a universal cutoff. It assumes independent clusters and a correctly specified random-intercept data-generating process.

Download the complete verified R script · Download Python script · Download shared synthetic CSV.

08

What to Tell Your Committee

Explain the sampling and assignment units, cluster count and size distribution, outcome and ICC scale, and whether the ICC came from an empty or adjusted model. Identify the focal within-cluster, between-cluster, or population-average comparison and why the chosen model and uncertainty method fit that comparison.

Report relevant variance-component uncertainty and model diagnostics, then describe sensitivity to defensible alternatives. A useful explanation is: “Although the estimated ICC was [value], we selected [approach] because [design and target]. There were [clusters] independent clusters with [size distribution]. The focal estimate and interval were [results], and [sensitivity findings] qualified the conclusion.” Replace each field with actual evidence.

Do not claim “multilevel modeling was unnecessary because ICC < .05” or “the large individual sample solved the small cluster count.”

09

When You Need More Help

The first checks often clarify the cluster units, size distribution, ICC definition, and target effect. Crossed memberships, cluster-level treatment, very few clusters, informative sizes, nonlinear outcomes, or random slopes may require closer design review.

Bring your model to the Analysis Clinic. The Multilevel and Hierarchical Models for Education workflow supports cluster audits, ICCs, centering decisions, random-intercept and random-slope models, and reproducible diagnostics.

Use the multilevel workflow.

Related case: Why Am I Getting a Singular Fit in My Mixed Model?. For an unusual design, discuss the analysis in a consultation.

Can't find a time that works? Email bookingrequest@dissertationstatshelper.com with a few times you're available and your time zone, and we'll find a time. It's a new address, so if you don't see our reply, please check your spam folder.

Technical references: