ANALYSIS CLINIC · CASE 011

CASE STATUS: DIAGNOSED

Something looks wrong with my data · Group comparisons · Precision and assumptions

My Groups Are Extremely Unequal

Unequal groups change precision and can expose fragile assumptions. Check the outcome, group-specific information, and target comparison before discarding observations to balance the counts.

01

Symptoms

You have 20 participants in one group and 100 in another, or a small subgroup alongside several large groups. Someone recommends deleting participants from the larger group, switching to a nonparametric test, or abandoning the comparison.

Unequal counts alone do not determine whether the analysis is valid. The consequences depend on variances, distributions, independence, selection, and the effect you want to estimate.

02

What This Usually Means

For two independent means, the estimated variance of their difference is s1²/n1 + s2²/n2 when variances are estimated separately. A large group cannot eliminate uncertainty contributed by a small group. With equal variances and a fixed total sample, balanced allocation gives greater precision for this simple difference; unequal costs or variances can change the optimal recruitment allocation.

A pooled t test assumes a common population variance. Welch’s test estimates group variances separately and uses an approximate degrees-of-freedom calculation. Imbalance combined with different variances can make the common-variance assumption consequential. Choose the analysis for its assumptions and target, rather than whichever p-value is smaller.

A mean difference, an adjusted contrast, a distributional comparison, and a prediction target are different questions. Unequal group sizes also do not demonstrate that selection or confounding is absent.

03

Common Causes

  • One group is rare, access is limited, or recruitment followed naturally occurring proportions.
  • Dropout, exclusions, or missing outcomes differ across groups.
  • Group-specific variability or outliers differ, with little information in the smaller group.
  • An adjusted analysis has poor overlap, sparse cells, or interactions supported by few observations.
  • Participant counts conceal clustered assignment, repeated records, weights, or unequal numbers of independent units.

04

Run These Checks

  1. Define the target. Identify the group comparison, outcome scale, population, and whether the goal is description, prediction, or a causal contrast. Distinguish sample proportions from target-population proportions.
  2. Count usable information by group. Report recruitment, exclusions, missingness, final counts, and independent clusters or events where relevant. Different missingness can change the population represented by the comparison.
  3. Inspect group distributions. Plot observations or suitable summaries and examine means, SDs, skewness, floor/ceiling effects, and influential values. Investigate possible errors without deleting valid values solely to improve significance.
  4. Check assumptions and design. For independent continuous-group means, consider Welch inference when a common variance is not justified. Do not use a preliminary variance-test p-value as the sole rule for choosing a pooled or Welch analysis. Severe skewness, tiny groups, clustering, or dependence require further design-specific assessment.
  5. Audit adjusted comparisons. Examine covariate overlap, group-specific information, variance structure, and sparse interaction cells. Robust standard errors do not create overlap or guarantee reliable inference with very few independent units.
  6. Report the estimate and uncertainty. Use an interval appropriate to the design and model, and show group-specific summaries. If performing multiple comparisons, specify the contrasts and multiplicity procedure. A large total N does not establish precision for every subgroup.

05

What Not to Do

  • Do not randomly delete valid observations merely to equalize group counts.
  • Do not duplicate observations or use synthetic oversampling as though it added independent participants for inferential group comparisons.
  • Do not assume a rank test estimates the same mean difference or automatically solves unequal variability.
  • Do not assume ordinary label permutation is valid when exchangeability fails; permutation inference must match the design and hypothesis.
  • Do not choose the method that restores statistical significance, or call nonsignificance evidence of no difference.
  • Do not combine substantively different groups only to enlarge a sparse subgroup.

06

Treatment Options

Retain all eligible observations and use inference that matches the comparison and variance structure. For independent continuous outcomes, Welch’s two-group test or a suitable multi-group procedure may address heterogeneous variances; neither fixes dependence or selection bias.

Recruit additional participants in the smaller or less informative group when feasible and consistent with the protocol. Use a design-specific precision or power calculation instead of targeting equal counts automatically.

Use a justified adjusted model when the research question requires adjustment. Check overlap and account for the design and variance pattern. Preserving a causal interpretation requires assumptions beyond balancing counts.

Consider a robust or distribution-focused analysis when it answers the intended question. State whether the target becomes a trimmed mean, quantile, or distributional contrast. Bootstrap and permutation methods also need suitable assumptions.

Use sampling or population weights only when their design and target justify them. Matching or weighting for causal questions requires additional assumptions and can reduce effective information; neither is a cosmetic balancing step.

Qualify subgroup conclusions when the available data cannot estimate them precisely. A narrower question may be defensible if the change and its timing are reported transparently.

07

Worked Example

Assume two independent groups with continuous normal outcomes. Hypothetical sample summaries are n1 = 20, mean1 = 53, SD1 = 12, and n2 = 100, mean2 = 50, SD2 = 6. The observed difference is 3 points. These summaries are stipulated, not empirical data or a simulated dropout process.

Welch uses SE = sqrt(144/20 + 36/100) = 2.749545 and approximately 20.937454 degrees of freedom. The smaller group contributes 7.2 of the total 7.56 variance, or 95.24%. Despite 120 total observations, uncertainty about this comparison is dominated by that group.

The pooled calculation uses a common variance estimate dominated by the larger, less variable group. Its SE is 1.789802, producing a narrower interval. This comparison illustrates the consequence of assuming a common variance; these summaries alone do not establish the true population variances or demonstrate the long-run error rate of either test.

Both intervals include zero and substantial positive differences. Neither establishes that the population difference is zero. Welch’s procedure addresses unequal variances under its assumptions, not confounding, informative dropout, or clustered observations.

A separate allocation-only calculation assumes common population SD = 10 and fixes total N at 120. Allocation 20/100 gives SE 2.449490; 60/60 gives SE 1.825742. The ratio is 1.341641. This is a planning comparison under equal variance and equal participant costs, not a reason to delete 80 existing observations.

Scope: the code implements standard summary-statistic formulas. It does not assess distribution shape, handle missing data, fit adjusted models, or provide multi-group multiplicity corrections.

Verified results for the hypothetical mean difference of 3 points
MethodSEdf95% intervalTwo-sided p
Welch2.74954520.937454−2.719033 to 8.719033.287631
Pooled1.789802118−0.544294 to 6.544294.096354

See it in R and Python

Base R; Python requires NumPy and SciPy: python -m pip install numpy scipy. No CSV is needed; both scripts use the same summaries.

# Case 011: hypothetical independent-group summaries, base R.
# Continuous normal outcomes; no sampled observations or attrition simulation.
n1 <- 20; n2 <- 100; mean1 <- 53; mean2 <- 50; sd1 <- 12; sd2 <- 6
difference <- mean1-mean2
v1 <- sd1^2/n1; v2 <- sd2^2/n2
se_w <- sqrt(v1+v2)
df_w <- (v1+v2)^2/(v1^2/(n1-1)+v2^2/(n2-1))
pooled_variance <- ((n1-1)*sd1^2+(n2-1)*sd2^2)/(n1+n2-2)
se_p <- sqrt(pooled_variance*(1/n1+1/n2))
summarize <- function(se,df) {
  statistic <- difference/se
  half_width <- qt(.975,df)*se
  c(difference=difference,SE=se,df=df,t=statistic,p=2*pt(-abs(statistic),df),
    CI_low=difference-half_width,CI_high=difference+half_width)
}
results <- rbind(Welch=summarize(se_w,df_w),pooled=summarize(se_p,n1+n2-2))
print(round(results,6))
# Allocation-only comparison: assumed common population SD=10 and fixed total N=120.
se_unequal <- 10*sqrt(1/20+1/100)
se_balanced <- 10*sqrt(1/60+1/60)
print(c(equal_variance_SE_20_100=se_unequal,equal_variance_SE_60_60=se_balanced,
        allocation_SE_ratio=se_unequal/se_balanced,small_group_variance_share=v1/(v1+v2)))
stopifnot(se_w>se_p,abs(v1/(v1+v2)-20/21)<1e-10,
 abs(se_unequal/se_balanced-sqrt(1.8))<1e-10,
 results["Welch","CI_low"]<0,results["Welch","CI_high"]>0)
sessionInfo()
"""Case 011: hypothetical independent-group summaries. NumPy + SciPy.
No sampled observations; formulas assume continuous normal outcomes.
"""
import numpy as np
import scipy
from scipy.stats import t
n1,n2=20,100
mean1,mean2=53.,50.
sd1,sd2=12.,6.
difference=mean1-mean2
v1,v2=sd1**2/n1,sd2**2/n2
se_w=np.sqrt(v1+v2)
df_w=(v1+v2)**2/(v1**2/(n1-1)+v2**2/(n2-1))
pooled_variance=((n1-1)*sd1**2+(n2-1)*sd2**2)/(n1+n2-2)
se_p=np.sqrt(pooled_variance*(1/n1+1/n2))
def summarize(se,df):
    statistic=difference/se
    half_width=t.ppf(.975,df)*se
    return np.array([difference,se,df,statistic,2*t.sf(abs(statistic),df),
                     difference-half_width,difference+half_width])
welch,pooled=summarize(se_w,df_w),summarize(se_p,n1+n2-2)
print("difference SE df t p CI_low CI_high")
print("Welch:",np.round(welch,6))
print("pooled:",np.round(pooled,6))
# Separate allocation-only comparison: common population SD=10, total N=120.
se_unequal=10*np.sqrt(1/20+1/100)
se_balanced=10*np.sqrt(1/60+1/60)
print("equal_variance_SE_20_100:",round(se_unequal,6))
print("equal_variance_SE_60_60:",round(se_balanced,6))
print("allocation_SE_ratio:",round(se_unequal/se_balanced,6))
print("small_group_variance_share:",round(v1/(v1+v2),6))
assert se_w>se_p
assert np.isclose(v1/(v1+v2),20/21,atol=1e-10)
assert np.isclose(se_unequal/se_balanced,np.sqrt(1.8),atol=1e-10)
assert welch[5]<0<welch[6]
print("All assertions passed; NumPy",np.__version__,"SciPy",scipy.__version__)

Verified: R 4.6.0 and Python 3.12.1 / NumPy 1.26.4 / SciPy 1.15.3 agree at the displayed precision. Both scripts passed all assertions.

Download verified R script · Download verified Python script

08

What to Tell Your Committee

“The final groups contained [counts and independent units]. We estimated [target contrast] using [method] because [design and variance assumptions]. We examined [distributions, missingness, and overlap] and retained [eligible observations]. The estimate was [effect and interval]. The smaller group limits [specific conclusion], and the comparison depends on [selection, measurement, and model assumptions].”

Include a group summary table, the sample flow, the primary contrast and interval, and a clear rationale for the uncertainty calculation. Explain which assumption matters rather than defending a group-size ratio against a universal rule.

09

When You Need More Help

Bring the research question, group definitions, outcome type, counts and group-specific summaries, missingness, design and clustering information, proposed covariates, and current results. A consultation can help distinguish loss of precision from a design or modeling problem.

Book a free consultation

Can't find a time that works? Email bookingrequest@dissertationstatshelper.com with a few times you're available and your time zone, and we'll find a time. It's a new address, so if you don't see our reply, please check your spam folder.

Related case: My Sample Is Smaller Than Planned

Technical references: