ANALYSIS CLINIC · CASE 012

CASE STATUS: DIAGNOSED

My committee or reviewer challenged my analysis · Covariates · R and Python

My Committee Wants More Covariates

Choose covariates for the question and design. A longer variable list can address confounding, change the target effect, or introduce bias.

01

Symptoms

Your committee asks you to add age, income, baseline scores, or every variable in the dataset. The request may be reasonable, but “more controls” does not specify what the expanded model should estimate. Ask what concern each proposed variable addresses and when it was measured.

02

What This Usually Means

A covariate is an additional variable in a model. Its role depends on the question and design. For a descriptive analysis, adjustment defines a conditional comparison. For prediction, judge candidate predictors using validation that reflects future use and information available at prediction time. For a causal analysis, choose adjustments using a defensible account of how the exposure, outcome, and other variables relate.

A larger adjustment set is not automatically more credible. A plausible common cause may need adjustment. A baseline outcome predictor may improve precision. A mediator may change a total-effect question into a conditional or direct-effect question. Conditioning on a collider—a common effect of two variables—can create an association along a previously closed path. These roles cannot be established by correlations or p-values alone.

03

Common Causes

The research question does not distinguish a total effect, a direct effect, a conditional association, and prediction.

A demographic checklist substitutes for an explanation of timing and causal roles.

Covariates are chosen because they predict the outcome, change the focal coefficient, or reach significance.

Adding terms consumes information, increases collinearity, or excludes observations with missing values.

The committee wants reassurance about alternative explanations, but the proposed controls do not address those explanations.

04

Run These Checks

1. Write the target comparison in one sentence. Define exposure, outcome, population, time ordering, and whether the goal is association, prediction, or causation.

2. Make a covariate decision table. Record each variable’s timing, proposed role, evidence for that role, missingness, and include/exclude rationale. A pretreatment measurement alone does not guarantee a valid control.

3. For causal questions, sketch the assumed causal structure. Identify pathways you want to retain and noncausal paths you need to block. Check whether the proposed set actually addresses confounding under those assumptions. An unmeasured common cause cannot be fixed by simply expanding a variable list.

4. Check information and functional form. Count model parameters, including factor levels, interactions, and nonlinear terms. Examine overlap and independent variation. There is no universal observations-per-covariate rule for all designs.

5. Separate sample changes from adjustment. Compare models on the same identifiable rows as a diagnostic, then choose a justified missing-data strategy. Equal sample sizes alone do not prove equal observations.

6. Fit a small set of scientifically defensible specifications. Report estimates and uncertainty, sample composition, assumptions, and sensitivity. For prediction use appropriate validation; do not use a favorable sign or p-value to pick the winning model.

05

What Not to Do

Do not add every available variable automatically.

Do not screen confounders using univariable significance or a percentage-change rule alone.

Do not label a mediator-adjusted coefficient a causal direct effect without additional identification assumptions.

Do not assume controlling for a post-exposure measure improves a total-effect analysis.

Do not use lower residual error, higher R-squared, or lower VIF as proof that the adjustment set is valid.

Do not hide dropped observations or choose the specification that preserves significance.

06

Treatment Options

Add a justified control when the design and assumed causal structure support its role. Explain the alternative explanation it addresses and account for its measurement quality and functional form.

Add a precision variable when it is appropriate for the design and target. In randomized studies, prespecified baseline adjustment can improve precision; it is not needed to make randomization remove baseline confounding in expectation. Follow the assignment design when estimating uncertainty.

Retain the total-effect question when a proposed variable lies on the causal pathway. Consider a separate mediation analysis only if it answers a substantive question and its assumptions are supportable.

Exclude a proposed collider when adjustment would open a noncausal path under the stated structure. Its role requires substantive knowledge, not a diagnostic software flag.

Report defensible sensitivity analyses when alternative structures or specifications are plausible. Explain their assumptions rather than presenting a larger model as inherently stronger.

07

Worked Example

Two controlled causal worlds illustrate different consequences of adjustment. The 32 rows enumerate independent centered factors. They are an algebra teaching device, not a recommended dissertation sample size or observational evidence that identifies a real-world causal structure.

World A: c causes x and y; x affects y directly and through m. The equations are x = c + a, m = x + b, and y = x + 2c + 2m + e. Under this stipulated structure the total effect is 3 and the direct effect is 1. The unadjusted x slope is 4, adjusting for c gives 3, and adding m gives 1. The last coefficient answers a different question; in actual data a mediator-adjusted regression does not automatically identify a direct effect.

World B is separate: x2 = a, y2 = x2 + b + e, and k = x2 + b + d. The graph contains x2 → k ← b → y2. The unadjusted slope is 1; adjusting for the collider k changes it to 0.5 despite the stipulated effect remaining 1. The additional covariate introduces bias for this target.

Both R and Python reproduce these exact slopes on the same constructed data. No CSV is required. Python needs NumPy. We omit inferential intervals because this deliberately enumerated dataset is used to verify algebra, not to represent a random sample.

See it in R and Python

# Case 012: controlled teaching examples; base R only.
dat <- expand.grid(c = c(-1,1), a = c(-1,1), b = c(-1,1), e = c(-1,1), d = c(-1,1))
dat$x <- dat$c + dat$a
dat$m <- dat$x + dat$b
dat$y <- dat$x + 2*dat$c + 2*dat$m + dat$e
fits <- list(unadjusted=lm(y~x,dat), confounder=lm(y~x+c,dat), mediator=lm(y~x+c+m,dat))
print(sapply(fits, function(f) coef(f)["x"]))
# A separate causal world: conditioning on a common effect.
dat$x2 <- dat$a
dat$y2 <- dat$x2 + dat$b + dat$e
dat$k <- dat$x2 + dat$b + dat$d
collider_fits <- list(unadjusted=lm(y2~x2,dat), collider=lm(y2~x2+k,dat))
print(sapply(collider_fits, function(f) coef(f)["x2"]))
stopifnot(all(abs(sapply(fits,function(f) coef(f)["x"])-c(4,3,1))<1e-10),
 all(abs(sapply(collider_fits,function(f) coef(f)["x2"])-c(1,.5))<1e-10),
 all(sapply(c(fits,collider_fits),nobs)==32L))
sessionInfo()

Verified in R 4.6.0 and NumPy 1.26.4: World A: 4 → 3 → 1. World B: 1 → 0.5. All numerical assertions passed.

Download R script · Download Python script

08

What to Tell Your Committee

“Our primary question is [target comparison]. We included [variables] because [timing, design, and proposed role]. We did not include [variable] in the primary model because [reason relevant to that target]. We compared [defensible specifications], documented changes in observations, and reported [estimates and uncertainty]. The conclusions depend on [assumptions and limitations].”

A useful response to the request is a short covariate decision table plus a clearly labeled sensitivity analysis. Acknowledge a justified new control when it improves the analysis; explain a questionable control in relation to the question rather than dismissing the committee’s concern.

09

When You Need More Help

Bring the research question, study design, variable definitions and measurement times, proposed controls, missingness summary, and existing model output. A consultation can help distinguish adjustment choices, changes in the target comparison, and limits of what the design supports.

Book a free consultation

Technical reference: Cinelli, Forney & Pearl: A Crash Course in Good and Bad Controls