ANALYSIS CLINIC · CASE 010

CASE STATUS: DIAGNOSED

Something looks wrong with my data · Missing data · Assumptions and information

Missing Data: Deletion, Multiple Imputation, or FIML?

Choose a missing-data strategy for your estimand, analysis model, and defensible assumptions—not a universal percentage cutoff.

01

Symptoms

Your software drops participants, different models use different samples, or your committee asks whether to delete incomplete records, impute them, or use FIML. A percentage alone cannot choose the method. Start with what is missing, why, and which quantity you want to estimate.

02

What This Usually Means

MCAR means missingness does not depend on observed or unobserved values. MAR means that, conditional on the observed information used in the model, missingness does not additionally depend on the missing values. MNAR means that dependence remains. MAR is an assumption, not a label established by a nonsignificant test.

Complete-case deletion changes the available sample and loses information. MCAR is a sufficient condition for many complete-case analyses, but not a necessary condition for every regression coefficient: validity depends on the model and selection process. A marginal mean and a conditional regression slope can behave differently.

Multiple imputation creates several plausible completed datasets, fits the analysis separately, and pools estimates and uncertainty. FIML fits a specified model to each record’s observed variables by integrating over missing components; it does not create a filled-in dataset. Both require appropriate models and assumptions. Standard MAR implementations do not automatically address MNAR.

03

Common Causes

Participants skip sensitive items or discontinue follow-up. Missingness differs by observed baseline status, group, or prior responses.

Administrative errors, planned missingness, or instrument routing create patterns with different substantive meanings. A structurally inapplicable question is not automatically a value to impute.

Special numeric codes are mistaken for real observations, or scores are computed with inconsistent item rules. Missing predictors, outcomes, and repeated measurements have different consequences.

04

Run These Checks

Define the estimand and analysis first: population mean, conditional association, treatment contrast, or another target. Identify the variables and independent units required.

Audit missing codes, item scoring rules, skip logic, exclusions, and missingness by variable, group, and time. Report complete-case counts for the actual analysis, not just overall missing percentages.

Compare observed information across missingness patterns and investigate recorded reasons. Such comparisons may challenge MCAR but cannot verify MAR versus MNAR from observed data alone.

Identify observed predictors of missingness and of the incomplete variables. Consider justified auxiliary information. Protect the design: clusters, repeated measurements, weights, interactions, and nonlinear terms require compatible models.

For MI, inspect convergence and observed/imputed distributions, logged problems, plausibility, fraction of missing information, and Monte Carlo stability. Choose enough imputations for the desired stability; do not regard five as a universal default.

For FIML, check convergence, identification, distributional assumptions, estimator support, and which cases the software actually retains. Missing exogenous covariates can still trigger deletion under some settings. Plan sensitivity analyses for credible departures from MAR.

05

What Not to Do

Do not choose deletion merely because missingness is below 5% or 10%. Do not choose the method that produces a preferred p-value.

Do not replace missing values with a single mean and analyze the result as if every value were observed. A missingness indicator is not a general repair for incomplete covariates.

Do not pool completed datasets into one enlarged dataset, average p-values, or ignore between-imputation variation. Pool the analysis estimates with appropriate uncertainty.

Do not impute impossible structural values or assume a generic continuous imputation model suits ordinal items, counts, bounds, or nested data. Do not call agreement between MI and FIML proof that MAR holds.

06

Treatment Options

Deletion can be defensible for a particular estimand under a justified selection assumption, with the lost information and analyzed sample clearly reported. Check that model comparisons do not silently compare different populations.

MI is flexible when incomplete predictors, auxiliary variables, or different downstream analyses matter. Include the outcome and analysis-relevant structure in the imputation model. Use type-appropriate models and analysis-specific pooling; substantive compatibility matters.

FIML is often useful within supported likelihood models, including continuous SEM. In lavaan, missing="ML" enables observed-data likelihood for supported ML analyses; verify exogenous-variable treatment and estimator compatibility. Ordinary FIML is not available for every categorical estimator.

When MNAR is plausible, specify scientifically meaningful sensitivity assumptions, such as shifts in missing outcomes, and report how conclusions change. Recovering data or documenting reasons may be more valuable than switching software. Neither method restores information that was never collected.

07

Worked Example

Consider a fixed sample of 200 participants: 100 in each fully observed group. Constructed outcome values have group means 10 and 14. We retain 80 symmetric outcome values in group 0 and 40 in group 1. This deliberately arranged dataset illustrates composition and uncertainty calculations; it is not empirical data, a random MAR simulation, or a demonstration of repeated-sampling bias.

The target is the equally weighted mean across these known groups. Deletion weights the observed groups 80/120 and 40/120, giving 11.333333. A saturated two-group normal observed-data likelihood estimates the two observed group means and weights them 1/2 each, giving 12.000000 with model-based SE 0.222158. The formula is a special closed-form FIML calculation, not a general SEM fitting routine.

The likelihood and MI calculations assume ignorable missingness within group and an appropriate independent normal group model. The construction does not verify those assumptions. Fully observed group membership supplies the target weights; unknown group memberships or missing predictors would require a different model.

Proper normal-model MI draws each group’s variance and mean, then missing outcomes. For each of 2,000 completed datasets, it estimates the same equally weighted target and its complete-data variance. Rubin pooling adds within-imputation variance to (1+1/M) times between-imputation variance.

R gives MI mean 11.992030 and pooled SE 0.228009; Python gives 12.005420 and SE 0.227084. Different random generators produce different draws. Both recover the same exact deletion and likelihood calculations; the MI difference is Monte Carlo variation. The scripts use a reference prior and a large-complete-sample Rubin degrees-of-freedom approximation, not a small-sample correction or general chained-equations implementation. Normal-model inference on these constructed values is illustrative.

For your own data, use a validated implementation suited to the variable types and design, inspect diagnostics, and assess missingness assumptions. A retained group with no observed outcomes would not support this calculation without additional assumptions.

See it in R and Python

# Case 010: constructed data; not empirical observations. Base R only.
set.seed(20261007)
x <- rep(0:1, each=100)
e <- rep(seq(-3,3,length.out=100),2)
y_full <- 10 + 4*x + e
observed <- rep(FALSE,200)
for(g in 0:1) {
  k <- if(g==0) 80L else 40L
  ix <- which(x==g)
  observed[ix[c(seq_len(k/2),seq.int(101-k/2,100))]] <- TRUE
}
y <- y_full; y[!observed] <- NA_real_
means <- sapply(0:1,function(g) mean(y[x==g],na.rm=TRUE))
counts <- sapply(0:1,function(g) sum(observed & x==g))
variances <- sapply(0:1,function(g) var(y[x==g],na.rm=TRUE))
# Saturated two-group normal observed-data likelihood, fixed target weights 1/2.
fiml_mean <- mean(means)
fiml_se <- sqrt(sum(.25*variances*(counts-1)/counts^2))
# Proper normal-model MI with independent reference priors p(mu,sigma^2) ~ 1/sigma^2.
M <- 2000L; Q <- U <- numeric(M)
for(i in seq_len(M)) {
  completed <- y
  for(g in 0:1) {
    j <- g+1L
    variance_draw <- (counts[j]-1)*variances[j]/rchisq(1,counts[j]-1)
    mean_draw <- rnorm(1,means[j],sqrt(variance_draw/counts[j]))
    missing <- x==g & !observed
    completed[missing] <- rnorm(sum(missing),mean_draw,sqrt(variance_draw))
  }
  Q[i] <- mean(completed)
  U[i] <- sum(sapply(0:1,function(g) .25*var(completed[x==g])/100))
}
W <- mean(U); B <- var(Q); T <- W+(1+1/M)*B
mi_se <- sqrt(T)
# Large-complete-sample Rubin df; no small-sample adjustment here.
df <- (M-1)*(1+W/((1+1/M)*B))^2
interval <- mean(Q)+c(-1,1)*qt(.975,df)*mi_se
print(counts)
print(c(deletion_mean=mean(y,na.rm=TRUE),fiml_mean=fiml_mean,fiml_SE=fiml_se,
        MI_mean=mean(Q),MI_SE=mi_se,within=W,between=B,total=T))
print(interval)
stopifnot(identical(as.integer(counts),c(80L,40L)),
          abs(mean(y,na.rm=TRUE)-34/3)<1e-10,abs(fiml_mean-12)<1e-10,
          T>W,abs(mean(Q)-12)<.1)
print(R.version.string)
# Case 010: constructed data; not empirical observations.
# Normal outcomes within two fully observed groups; missingness depends on group.
import numpy as np
from scipy.stats import t
rng = np.random.default_rng(20261007)
x = np.repeat([0, 1], 100)
e = np.tile(np.linspace(-3, 3, 100), 2)
y_full = 10 + 4*x + e
# Retain symmetric positions: 80 in group 0 and 40 in group 1.
observed = np.zeros(200, dtype=bool)
for g, k in [(0, 80), (1, 40)]:
    ix = np.flatnonzero(x == g)
    keep = np.r_[np.arange(k//2), np.arange(100-k//2, 100)]
    observed[ix[keep]] = True
y = np.where(observed, y_full, np.nan)
means = np.array([np.nanmean(y[x == g]) for g in [0, 1]])
counts = np.array([np.sum(observed & (x == g)) for g in [0, 1]])
variances = np.array([np.nanvar(y[x == g], ddof=1) for g in [0, 1]])
# Saturated two-group normal observed-data likelihood: group means and ML variances.
# Fixed target weights are 1/2, because group membership is known for all 200.
fiml_mean = means.mean()
fiml_se = np.sqrt(np.sum(.25 * variances*(counts-1)/counts**2))
# Proper normal-model MI: draw variance, then mean, then missing outcomes.
# Independent reference prior p(mu_g, sigma_g^2) proportional to 1/sigma_g^2.
M = 2000
Q, U = [], []
for _ in range(M):
    completed = y.copy()
    for g in [0, 1]:
        variance_draw = (counts[g]-1)*variances[g]/rng.chisquare(counts[g]-1)
        mean_draw = rng.normal(means[g], np.sqrt(variance_draw/counts[g]))
        missing = (x == g) & ~observed
        completed[missing] = rng.normal(mean_draw, np.sqrt(variance_draw), missing.sum())
    Q.append(completed.mean())
    U.append(sum(.25*np.var(completed[x == g], ddof=1)/100 for g in [0, 1]))
Q, U = np.array(Q), np.array(U)
W, B = U.mean(), Q.var(ddof=1)
T = W + (1 + 1/M)*B
mi_se = np.sqrt(T)
# Large-complete-sample Rubin df; small-sample adjustments are not implemented.
df = (M-1)*(1 + W/((1+1/M)*B))**2
interval = Q.mean() + np.array([-1, 1])*t.ppf(.975, df)*mi_se
print('Observed counts:', counts)
print('Deletion mean:', np.nanmean(y))
print('Likelihood mean / model-based SE:', fiml_mean, fiml_se)
print('MI mean / pooled SE / 95% interval:', Q.mean(), mi_se, interval)
print('Within / between / total variance:', W, B, T)
assert np.array_equal(counts, [80, 40])
assert np.isclose(np.nanmean(y), 34/3)
assert np.isclose(fiml_mean, 12)
assert T > W and abs(Q.mean()-12) < .1
print('All checks passed. NumPy', np.__version__)

Verified in R 4.6.0 and Python 3.12.1 / NumPy 1.26.4 / SciPy 1.15.3. All structural and numerical checks passed.

Download R script · Download Python script

08

What to Tell Your Committee

“Our target was [estimand]. Missingness affected [variables, counts, and patterns]. We selected [method] because [analysis and missingness rationale], using [predictors, auxiliary information, and model structure]. We assumed [specific conditions]. We checked [diagnostics] and evaluated [sensitivity assumptions]. The estimate was [effect and interval]; conclusions were [stable or sensitive] across the prespecified analyses.”

Provide a missingness table, sample flow, model specifications, MI settings and pooling details or FIML settings, diagnostics, and sensitivity results. Distinguish an assumption supported by substantive knowledge from one that the observed data cannot establish.

09

When You Need More Help

Bring the research question, data dictionary, missing codes and reasons, missingness patterns, design, proposed analysis, and current software settings. Complex clustering, interactions, incomplete categorical predictors, or credible MNAR need a tailored strategy.

Book a free consultation

Can't find a time that works? Email bookingrequest@dissertationstatshelper.com with a few times you're available and your time zone, and we'll find a time. It's a new address, so if you don't see our reply, please check your spam folder.

Technical references: