01
Symptoms
Your proposal called for 360 participants, but only 120 completed the study. A committee member asks whether the dissertation is still viable, whether you should run a new power analysis, or whether a nonsignificant finding simply means the sample was too small.
A shortfall does not automatically invalidate every result. It also does not justify treating the original precision or power calculation as if the planned sample had been obtained. The consequences depend on the design, the target effect, the remaining information, and why observations are missing.
02
What This Usually Means
With the same design and variability, fewer independent observations usually mean wider intervals and less ability to detect a specified effect. The sample count is only part of the story: imbalance, loss of entire clusters, fewer outcome events, selective dropout, or little predictor variation can matter more than the total number of rows.
Power is a probability of rejection under a specified alternative and analysis procedure. A calculation at the achieved sample can describe sensitivity to a prespecified meaningful effect under explicit assumptions. It is not the probability that your observed finding is true or that a nonsignificant result is a false negative.
Interpret the observed effect estimate and its interval in the substantive units of the question. An interval spanning both negligible and important effects signals limited discrimination. A nonsignificant test by itself does not establish equivalence or no meaningful effect.
03
Common Causes
- Recruitment assumptions were optimistic, access changed, or the recruitment window closed.
- Dropout, missing outcomes, or exclusions reduced the analysis sample after enrollment.
- The final allocation is uneven, key subgroups are sparse, or too few clusters or events remain.
- The original calculation used a different estimand, variance, model, or analysis unit from the final study.
- A request for a binary adequate/inadequate verdict obscures distinct concerns about bias, precision, estimability, and generalizability.
04
Run These Checks
- Reconstruct the sample flow. Record eligible, invited, enrolled, retained, and analyzed counts, with documented reasons for losses. Distinguish participants from repeated rows, events, and independent clusters.
- Revisit the original justification. Preserve the planned effect, variability, allocation, alpha, target power or precision, and design assumptions. Identify which assumptions changed. A broad rule such as “30 participants is enough” cannot replace this review.
- Investigate selection and missingness. Compare available baseline information for retained and lost participants. These comparisons cannot establish an ignorable missingness mechanism. Consider methods and sensitivity analyses appropriate to the design; imputation does not create new independent participants.
- Check whether the model is estimable. Examine sparse cells, outcome events, parameter count, predictor variation, overlap, convergence, and covariance boundaries. Do not drop scientifically required confounders solely to improve significance.
- Quantify the information actually available. Report the primary estimate and an appropriate interval. If useful, calculate achieved-design power across externally justified effects, or the effect needed for a target power, with assumptions clearly stated. Do not plug the observed effect into “observed power” to explain its p-value.
- Document the decision and its timing. Retain the prespecified primary analysis when defensible. Label changed outcomes, simplified models, subgroups, and additional analyses accurately. If recruitment continues, use an approved stopping plan that handles any interim examination of outcomes.
05
What Not to Do
- Do not announce that the study is adequately powered because the observed result is significant.
- Do not use observed-effect post hoc power as evidence for the null or an explanation of a nonsignificant result.
- Do not relax alpha, switch to a one-sided test, combine incompatible groups, or repeatedly try outcomes to obtain significance.
- Do not assume a nonparametric test, bootstrap, Bayesian model, or imputation automatically fixes the information shortfall. Each has its own assumptions.
- Do not collect until p < .05 without a valid sequential procedure.
- Do not present an unplanned reduction of the model or a revised effect threshold as though it were the original plan.
06
Treatment Options
Continue recruitment when additional enrollment is feasible and consistent with ethics approval, resources, eligibility, and the stopping plan. More participants can improve precision but cannot by itself repair selective recruitment or measurement errors.
Keep the primary question and report uncertainty when the model remains defensible. State the shortfall, its causes, and which effects the interval can or cannot distinguish.
Use a justified missing-data approach and sensitivity analyses when missingness is the main source of loss. Explain the assumptions and how conclusions change under plausible departures.
Simplify an overambitious analysis only when the revised model answers a defensible question. Reduce unnecessary complexity rather than required adjustment, and report the change as planned or exploratory according to its actual timing.
Limit the claims or revise the study aim when the achieved information cannot support the intended conclusion. An estimation-focused or exploratory contribution can still be useful, but a small sample does not automatically become a pilot study.
For a question about negligible effects, use equivalence testing only with substantively justified bounds and an appropriate design and procedure. Ordinary nonsignificance is not a substitute.
07
Worked Example
Consider two independent groups with normal outcomes, equal allocation, common population SD = 10, a prespecified meaningful difference of 3 points, and a two-sided equal-variance t test at alpha = .05. The plan is 180 per group; the achieved design has 60 per group. Power calculations retain the same externally specified effect and population SD.
To isolate the precision change, suppose the observed mean difference is 3 and both sample SDs are 10 in each hypothetical scenario. The intervals below are calculated from those summaries. They are not actual study results, simulated attrition, or a claim that a real shortfall leaves the estimate unchanged.
The standard error increases by sqrt(180 / 60) = 1.732. The achieved-sample interval includes zero and effects exceeding the meaningful three-point difference, so it does not establish no effect. Its upper limit also shows that some sufficiently large effects are incompatible with these summaries under the model.
At 60 per group, a difference of about 5.157 points would provide 80% power under these assumptions. That is a sensitivity calculation, not a new definition of scientific importance, a guarantee of detection, or an equivalence bound. It does not replace the prespecified three-point target.
Scope: the pooled t calculation does not directly apply to unequal variances, clustered data, repeated measures, sparse binary outcomes, or adjusted models. Use a calculation or simulation reflecting the actual design.
| Scenario | N per group (total) | SE | 95% interval | Power for 3 points |
|---|---|---|---|---|
| Planned | 180 (360) | 1.054 | 0.927 to 5.073 | 81.0% |
| Achieved | 60 (120) | 1.826 | −0.615 to 6.615 | 37.1% |
See it in R and Python
Base R; Python requires NumPy and SciPy: python -m pip install numpy scipy. No CSV is needed. Both scripts use the same assumptions and summaries.
# Case 007: planned versus achieved sample, base R only.
# Independent normal outcomes, equal allocation and common SD.
# n is PER GROUP. Hypothetical summaries, not simulated attrition.
n <- c(planned=180L, achieved=60L)
sd_assumed <- 10
important_difference <- 3
alpha <- .05
summarize <- function(n) {
df <- 2*n-2
se <- sd_assumed*sqrt(2/n)
half_width <- qt(1-alpha/2,df)*se
power <- power.t.test(n=n,delta=important_difference,sd=sd_assumed,
sig.level=alpha,type="two.sample",alternative="two.sided",strict=TRUE)$power
# CI below assumes sample SDs = 10 and observed mean difference = 3.
c(per_group=n,total=2*n,SE=se,CI_low=3-half_width,CI_high=3+half_width,
power_for_prespecified_difference=power)
}
results <- t(sapply(n,summarize))
print(round(results,6))
mde <- power.t.test(n=n["achieved"],sd=sd_assumed,sig.level=alpha,
power=.80,tol=1e-10,type="two.sample",alternative="two.sided",strict=TRUE)$delta
print(c(difference_for_80_percent_power=mde,
SE_ratio=results["achieved","SE"]/results["planned","SE"]))
stopifnot(abs(results["achieved","SE"]/results["planned","SE"]-sqrt(3))<1e-10,
results["planned","CI_low"]>0, results["achieved","CI_low"]<0,
results["achieved","power_for_prespecified_difference"]<results["planned","power_for_prespecified_difference"],
mde>important_difference)
sessionInfo()
"""Case 007: same planning assumptions and hypothetical summaries as R.
Requires NumPy and SciPy. n is per group, not total sample size.
"""
import numpy as np
import scipy
from scipy.stats import t, nct
from scipy.optimize import brentq
sd_assumed=10.0
important_difference=3.0
alpha=.05
def design_power(n,delta):
df=2*n-2
critical=t.ppf(1-alpha/2,df)
noncentrality=delta/(sd_assumed*np.sqrt(2/n))
return nct.sf(critical,df,noncentrality)+nct.sf(critical,df,-noncentrality)
def summarize(n):
se=sd_assumed*np.sqrt(2/n)
half_width=t.ppf(1-alpha/2,2*n-2)*se
# Assumes sample SDs = 10 and observed mean difference = 3.
return np.array([n,2*n,se,3-half_width,3+half_width,design_power(n,important_difference)])
planned,achieved=summarize(180),summarize(60)
print("per_group total SE CI_low CI_high power_for_prespecified_difference")
print("planned:",np.round(planned,6))
print("achieved:",np.round(achieved,6))
mde=brentq(lambda delta:design_power(60,delta)-.80,0,10)
print("difference_for_80_percent_power:",round(mde,6))
print("SE_ratio:",round(achieved[2]/planned[2],6))
assert np.isclose(achieved[2]/planned[2],np.sqrt(3),atol=1e-10)
assert planned[3]>0 and achieved[3]<0
assert achieved[5]<planned[5] and mde>important_difference
print("All assertions passed; NumPy",np.__version__,"SciPy",scipy.__version__)
Verified: R 4.6.0 and Python 3.12.1 with NumPy 1.26.4 / SciPy 1.15.3 agree at the displayed precision. Both scripts passed all assertions. The difference for 80% power is 5.157066 and the SE ratio is 1.732051.
Download verified R script · Download verified Python script
08
What to Tell Your Committee
“We planned [independent units and allocation] and analyzed [achieved counts]. The shortfall arose from [documented reasons]. We assessed [selection, missingness, and model feasibility]. For [primary question], the estimate was [estimate] with [interval]. The data distinguish [supported conclusions] but remain compatible with [important alternatives]. Our sensitivity calculation retained [prespecified effect and assumptions]. We [retained/revised] the analysis because [reason], and any changes are labeled [timing and status].”
Attach a sample-flow table, the original justification, the primary estimate and interval, and a concise explanation of any revisions. The aim is to make the limits reviewable rather than defend the sample using a universal cutoff.
09
When You Need More Help
Bring the original sample-size justification, recruitment and exclusion counts, allocation or cluster structure, outcome type, missingness summary, planned model, and primary results. A consultation can help identify what remains estimable and which claims need qualification.
Can't find a time that works? Email bookingrequest@dissertationstatshelper.com with a few times you're available and your time zone, and we'll find a time. It's a new address, so if you don't see our reply, please check your spam folder.
Related cases: Small ICC and multilevel modeling · Choosing covariates
Technical references: