01
Symptoms
You fit y ~ x, then add covariates with y ~ x + z. The coefficient for x switches from positive to negative. You may also see a larger standard error, a reversal after one particular covariate enters, or different signs across nearby specifications.
The sign change itself is not a software warning. A small positive estimate and a small negative estimate with wide intervals may both be compatible with considerable uncertainty. Record the estimates and intervals rather than treating the signs alone as opposing conclusions.
02
What This Usually Means
In an additive ordinary least-squares model, the unadjusted slope describes the linear association between x and y. The adjusted slope describes the association between the parts of x and y left after their linear relationships with the included covariates are removed. These are different statistical comparisons.
For the same observations, an intercept, two predictors, and no interactions, the fitted slopes satisfy:
unadjusted x slope = adjusted x slope + adjusted z slope × slope from regressing z on x.
Shared variation can therefore outweigh an adjusted coefficient and reverse the marginal sign. This identity explains the algebra; it does not identify which adjustment is scientifically appropriate or establish a causal effect.
Scope: This case focuses on additive linear regression. With interactions, a coefficient is conditional on the other variable's reference value. With logistic regression, adjusted and unadjusted odds ratios are also affected by noncollapsibility; the linear identity above does not apply.
03
Common Causes
- Adjustment changes the comparison. The added variable is associated with both the focal predictor and outcome, so the marginal and conditional slopes can differ substantially.
- Suppression or shared predictor variation. Adjustment may reveal a relationship obscured in the marginal association. Suppression terminology varies; demonstrate the coefficient pattern before assigning a label.
- The analysis sample changed. Missing covariate values or new exclusions can change the observations used, mixing sample changes with adjustment changes.
- The estimate is weakly determined. Highly related predictors, little independent predictor variation, or influential observations can make the conditional coefficient sensitive to small changes.
- Coding or model meaning changed. Reversing a scale, changing a factor reference, modifying a transformation, or adding an interaction can alter the coefficient's direction or interpretation.
- The adjustment set is inappropriate for the target. A mediator, collider, or post-outcome variable may change the causal question or introduce bias. A sign reversal alone cannot diagnose any of these roles.
04
Run These Checks
Preserve the two specifications and their uncertainty
formula(fit_unadjusted) formula(fit_adjusted) nobs(fit_unadjusted) nobs(fit_adjusted) coef(summary(fit_unadjusted)) coef(summary(fit_adjusted)) confint(fit_unadjusted) confint(fit_adjusted)What it suggests: identify the exact coefficient, scale, adjustment set, exclusions, weights, and uncertainty that changed. Matching counts alone do not establish identical observations.
Separate sample changes from adjustment changes
For this simple demonstration, create one complete-case sample for the variables in both models:
common <- dat[complete.cases(dat[c("y", "x", "z")]), ] fit_same_unadjusted <- lm(y ~ x, data = common, na.action = na.fail) fit_same_adjusted <- lm(y ~ x + z, data = common, na.action = na.fail)What it suggests: compare the original unadjusted fit with the common-sample unadjusted fit, then compare the two common-sample fits. The first comparison isolates a sample difference; the second isolates adjustment within that sample. This is a diagnostic comparison, not a recommendation that complete-case analysis is the final missing-data strategy. Adapt it to the actual terms, weights, filters, and missing-data procedure.
Audit coding and the comparison being made
str(common[c("y", "x", "z")]) summary(common[c("y", "x", "z")]) model.matrix(fit_same_adjusted)What it suggests: check scale direction, category references, impossible values, interactions, and transformations against the codebook. Positive rescaling alone cannot reverse a slope's sign. Centering alone does not reverse the slope in an additive model with an intercept; it can change a lower-order coefficient in an interaction model because that coefficient answers a different conditional question.
Examine shared variation and independent information
cor(common[c("y", "x", "z")]) # numeric variables in this example r2_x <- summary(lm(x ~ z, data = common))$r.squared vif_x <- 1 / (1 - r2_x) x_res <- resid(lm(x ~ z, data = common)) y_res <- resid(lm(y ~ z, data = common)) plot(x_res, y_res) coef(lm(y_res ~ x_res))What it suggests: the residualized plot shows the adjusted linear association. Little residual variation in
xwarns that the conditional comparison has limited information. VIF measures variance inflation from predictor relationships; no VIF cutoff by itself identifies the cause of a reversal. Use model-appropriate diagnostics for factors or multiple covariates. The residual regression illustrates the slope; use the full model for inferential standard errors and intervals.Investigate influence and model form
influence.measures(fit_same_adjusted) dfbetas(fit_same_adjusted) plot(fit_same_adjusted)What it suggests: examine whether individual observations drive the focal slope and whether the additive linear form is plausible. Review flagged observations for errors and substantively unusual cases; a diagnostic flag is not permission to delete an observation. Clustered or repeated data require diagnostics and uncertainty that respect the design.
Return to the research question
Identify whether the target is descriptive association, prediction, or a causal effect. Explain each covariate's timing and proposed role. For causal questions, use the design and a defensible causal structure to choose adjustment; correlations and coefficient changes alone cannot determine which controls are appropriate.
05
What Not to Do
- Do not choose the model because its sign agrees with your hypothesis.
- Do not call the adjusted result causal merely because it includes more covariates.
- Do not diagnose confounding, mediation, or suppression from a sign change alone.
- Do not drop a required covariate solely to reduce VIF or restore the preferred sign.
- Do not compare models fitted to different observations and attribute the entire change to adjustment.
- Do not delete influential observations without a documented reason and sensitivity assessment.
- Do not treat a sign change near zero as decisive evidence that the substantive relationship reversed.
06
Treatment Options
Correct coding or sample construction
Use when: The audit identifies a verifiable error
Tradeoff: Document the correction and rerun affected analyses.
Retain a justified adjustment set
Use when: The conditional comparison answers the stated question
Tradeoff: Explain why the marginal and adjusted associations differ; report uncertainty.
Reconsider an inappropriate control
Use when: Design and timing indicate the variable does not belong in the target adjustment set
Tradeoff: Removing it changes the comparison and must follow the scientific rationale.
Examine a small set of defensible models
Use when: Several specifications address plausible assumptions
Tradeoff: Report sensitivity; selecting the preferred sign conceals instability.
Qualify the conclusion
Use when: Independent predictor variation or precision is insufficient
Tradeoff: The dataset may not distinguish the direction of the target relationship.
Revise form or target
Use when: Interactions, nonlinearity, or aggregation make a single additive slope inadequate
Tradeoff: State the new question and distinguish planned analyses from exploratory changes.
Treatment principle: choose the comparison that answers the question and report what the data support.
07
Worked Example
This deliberately small deterministic teaching dataset contains eight observations. It is constructed so that x and z share variation, while the residual error is orthogonal to both. It illustrates the regression algebra without relying on a lucky random seed.
See it in R and Python
Both examples construct the same deterministic eight observations. Python uses NumPy and SciPy for least squares, t-based intervals, residualization, and VIF. No CSV is needed.
Python dependencies: NumPy and SciPy. Install them with python -m pip install numpy scipy.
# Analysis Clinic Case 005: deterministic same-sample sign reversal.
# Base R only. Run this script to verify the page example before publication.
dat <- expand.grid(a = c(-1, 1), b = c(-1, 1), e = c(-1, 1))
dat$x <- dat$a
dat$z <- dat$a + dat$b
dat$y <- 10 - dat$x + 2 * dat$z + dat$e
unadjusted <- lm(y ~ x, data = dat)
adjusted <- lm(y ~ x + z, data = dat)
x_res <- resid(lm(x ~ z, data = dat))
y_res <- resid(lm(y ~ z, data = dat))
partial <- lm(y_res ~ x_res)
print(coef(summary(unadjusted)))
print(coef(summary(adjusted)))
print(confint(unadjusted))
print(confint(adjusted))
print(c(n_unadjusted = nobs(unadjusted), n_adjusted = nobs(adjusted),
correlation_x_z = cor(dat$x, dat$z),
vif_x = 1 / (1 - summary(lm(x ~ z, dat))$r.squared),
partial_slope = unname(coef(partial)["x_res"])))
stopifnot(nobs(unadjusted) == 8L, nobs(adjusted) == 8L,
abs(coef(unadjusted)["x"] - 1) < 1e-10,
abs(coef(adjusted)["x"] + 1) < 1e-10,
abs(coef(adjusted)["z"] - 2) < 1e-10,
abs(coef(partial)["x_res"] + 1) < 1e-10)
sessionInfo()
"""Case 005: same-sample sign reversal, NumPy + SciPy.
Self-contained deterministic dataset; no download needed.
"""
import itertools
import numpy as np
import scipy
from scipy.stats import t
rows=np.array(list(itertools.product([-1.,1.],repeat=3)))
a,b,e=rows.T
x=a
z=a+b
y=10-x+2*z+e
def ols(X):
beta=np.linalg.lstsq(X,y,rcond=None)[0]
df=len(y)-X.shape[1]
cov=np.sum((y-X@beta)**2)/df*np.linalg.inv(X.T@X)
se=np.sqrt(np.diag(cov))
ci=np.column_stack([beta-t.ppf(.975,df)*se,beta+t.ppf(.975,df)*se])
return beta,se,ci,2*t.sf(np.abs(beta/se),df)
simple=ols(np.column_stack([np.ones(len(y)),x]))
adjusted=ols(np.column_stack([np.ones(len(y)),x,z]))
print('NumPy',np.__version__,'SciPy',scipy.__version__)
for label,result in [('unadjusted',simple),('adjusted',adjusted)]:
beta,se,ci,p=result
print(label,'x estimate:',beta[1],'SE:',se[1],'95% CI:',ci[1],'p:',p[1])
Z=np.column_stack([np.ones(len(y)),z])
x_res=x-Z@np.linalg.lstsq(Z,x,rcond=None)[0]
y_res=y-Z@np.linalg.lstsq(Z,y,rcond=None)[0]
partial=x_res@y_res/(x_res@x_res)
correlation=np.corrcoef(x,z)[0,1]
vif=1/(1-correlation**2)
print('Residualized slope:',partial,'correlation:',correlation,'VIF:',vif)
print('Both x intervals include zero; this is an algebra illustration.')
assert np.isclose(simple[0][1],1) and np.isclose(adjusted[0][1],-1)
assert np.isclose(adjusted[0][2],2) and np.isclose(partial,-1)
assert np.isclose(vif,2)
assert simple[2][1,0]<0<simple[2][1,1]
assert adjusted[2][1,0]<0<adjusted[2][1,1]
print('All checks passed.')
The exact constructed slopes are +1 for x in the unadjusted model and −1 for x in the adjusted model; the adjusted z slope is +2. Both fits use the same eight observations. The slope of z on x is +1, so the identity is +1 = −1 + 2 × 1. The correlation of x and z is approximately .707 and the VIF for x is 2: a reversal does not require extreme collinearity.
Interpretation: the marginal relationship mixes the negative conditional x slope with a positive contribution from shared z variation. The sign change follows from the constructed data and specification. This example does not establish which covariate should be included in a real dissertation or identify a causal effect. Eight observations here are a teaching device, not a recommended sample size.
Download the complete verified R script.
Verified R output: Verified under R 4.6.0 (2026-04-24 ucrt), Windows x64, base stats. All script assertions passed.
Unadjusted x: +1; SE 0.9128709; 95% CI −1.233715 to 3.233715; p = .3153.
Adjusted x: −1; SE 0.6324555; 95% CI −2.625779 to 0.625779; p = .1747.
Both intervals include zero. This demonstrates a sign reversal in the fitted coefficients, not strong evidence of opposite population effects. The dataset is constructed to illustrate algebra; these model-based intervals are teaching output rather than empirical evidence.
Python verification: the x slope changes from +1 to −1; both intervals include zero, and VIF = 2. The intervals and residualized slope reproduce the R results. All assertions passed.
08
What to Tell Your Committee
Explain the original comparison, the exact adjustment, whether the observations changed, how the focal scale was coded, and what the diagnostic checks showed. Justify the final covariates using the question and design, then report estimates, intervals, sensitivity, and causal limits.
Adaptable wording, to be completed with actual results:
“The unadjusted model estimated [comparison], whereas the adjusted model estimated [conditional comparison]. We compared the specifications on [same analysis sample / explained sample difference] and reviewed [coding, shared variation, influence, and model form]. We retained [covariates] because [design-based rationale]. The focal estimate changed from [estimate and interval] to [estimate and interval]. We interpret this as [supported explanation], with [limitations], rather than selecting a specification on the basis of its sign.”
09
When You Need More Help
A consultation can help distinguish sample changes, coding, shared variation, weak information, and adjustment choices. Bring the research question, two model specifications, coefficient tables with intervals, sample counts and exclusions, variable definitions, and covariate timing.
Can't find a time that works? Email bookingrequest@dissertationstatshelper.com with a few times you're available and your time zone, and we'll find a time. It's a new address, so if you don't see our reply, please check your spam folder.
.Related case: My Interaction Is Significant, but the Main Effects Aren't.
Technical references: