ANALYSIS CLINIC · CASE 009

CASE STATUS: DIAGNOSED

My model won’t run · SEM · Global and local fit

Why Is My SEM Model Fit Poor?

Check the data, estimation, and pattern of residuals before changing paths. Improving a fit index is useful only when the revised model answers a defensible question.

01

Symptoms

Your CFA or SEM converges, but the chi-square test rejects exact fit, CFI or TLI is low, RMSEA or SRMR is high, or the indices disagree. Software lists large modification indices, and someone suggests adding correlated errors until the fit looks acceptable.

Poor fit is different from failure to converge. First establish that the solution is numerically stable, identified, and admissible. A model that stops without an error message can still have negative variances, implausible estimates, or a worse local solution.

02

What This Usually Means

A covariance SEM imposes restrictions on relationships among the observed variables; a mean structure adds restrictions on means. Fit statistics summarize discrepancies between the model-implied quantities and those represented by the data under the estimator and assumptions.

Chi-square tests exact fit under its relevant sampling assumptions and is sensitive to sample size. CFI and TLI compare the proposed model with a baseline model; their interpretation depends partly on that baseline. RMSEA summarizes discrepancy per degree of freedom and has sampling uncertainty. SRMR summarizes standardized residuals; it can hide a localized pattern. None identifies the cause of misfit.

Conventional cutoffs are guides from particular settings, not universal acceptance rules. Inspect the statistic, degrees of freedom, uncertainty where applicable, estimator, sample, and local residuals together. A just-identified model can reproduce the sample moments by construction and provides no global test of its unrestricted parts.

Good fit does not establish the correct theory, construct validity, or causal direction. Alternative models may fit similarly, and data-driven modifications can capitalize on chance.

03

Common Causes

  • The data or scoring do not match the intended analysis: reverse keys, missing codes, item order, groups, or covariance input may be wrong.
  • The measurement model omits a factor, cross-loading, or substantively justified method effect, or constrains relationships too strongly.
  • The structural model imposes unsupported zero paths, equality constraints, or functional relationships.
  • The estimator or uncertainty calculation does not suit the item type, distribution, missingness, clustering, or study design.
  • The solution is unstable or inadmissible, or limited information makes estimates and fit indices imprecise.
  • Pooling heterogeneous groups, influential observations, or an unsupported mean structure produces a pattern the model cannot represent.

04

Run These Checks

  1. Preserve the original specification and output. Record the estimator, model syntax, analysis sample, degrees of freedom, fit statistics, and any warnings. Verify convergence and admissibility before interpreting fit or modification indices.
  2. Audit the input. Check scoring, item ranges, missing-value treatment, distributions, covariance/correlation input, sample size, and variable ordering. Examine flagged observations rather than deleting them to improve fit.
  3. Match estimation to the design. Ordered categorical indicators may require an ordinal measurement approach; nonnormal continuous responses may motivate robust methods. Missingness and clustering also need appropriate handling. Robust inference changes the relevant statistics; it does not automatically repair a misspecified structure.
  4. Separate measurement from structural restrictions. Assess whether the proposed measurement model can represent the indicators before attributing all misfit to structural paths. Compare defensible models on compatible outcomes and samples; changing the item set changes what is being modeled.
  5. Examine local discrepancies. Inspect observed-minus-implied covariances or correlations, and estimator-appropriate residual diagnostics. Look for blocks or other systematic patterns. Raw residuals, standardized covariance residuals, and residual z statistics are different quantities.
  6. Use modification indices as diagnostic clues. Review the proposed parameter, its expected parameter change, and the substantive explanation. A modification index approximates improvement from freeing a fixed parameter in the current model; it does not verify that the parameter belongs in the theory. Testing many changes and retaining the best requires transparent reporting and independent validation.
  7. Test a small set of justified revisions. Document why a loading, constraint, path, or residual covariance changes. For nested models, use the difference test appropriate to the estimator; subtracting robust chi-square statistics is not generally the correct robust difference test. Seek replication or a held-out evaluation when the sample supports it.

05

What Not to Do

  • Do not add every residual covariance or path with a large modification index.
  • Do not use a residual covariance to conceal an omitted dimension without evaluating that possibility.
  • Do not remove indicators solely because their deletion improves fit; content coverage and score meaning can change.
  • Do not switch estimators only to obtain more favorable indices or mix conventional and robust statistics without explanation.
  • Do not treat good global fit, nonsignificant chi-square, or a saturated model as proof of causal identification.
  • Do not report a heavily modified model as though it were the original confirmatory specification.

06

Treatment Options

Correct a verified input or scoring error and rerun all affected analyses. Preserve the original output and document the correction.

Revise the measurement structure when item content and measurement evidence justify it. Distinct constructs, cross-loadings, and method effects imply different interpretations, even if they each improve a fit index.

Relax a structural restriction when the revised relation answers the scientific question and has a defensible rationale. Fit alone cannot establish a causal path or resolve omitted confounding.

Use a suitable estimator and design-aware inference when assumptions warrant it. Report the relevant robust or categorical fit statistics and comparison method accurately.

Report remaining misfit and qualify interpretation when the model remains useful for a limited purpose but does not represent all relevant relationships. Evaluate whether the particular estimates you interpret are sensitive to the unresolved discrepancy.

Reconsider the proposed model or question when substantive restrictions are not supported. Exploratory work can guide a future model, but revisions selected on these data need validation.

07

Worked Example

This example supplies a controlled six-variable covariance matrix rather than sampled observations. Every variance is 2; covariances within items 1–3 and within items 4–6 are 1; covariances between those groups are .4. It is exactly representable by two correlated unit-variance factors, unit loadings, independent residual variances of 1, and factor correlation .4.

We assign a nominal N = 400 for conventional normal-ML fit calculations. That number illustrates how software computes fit from covariance input; these are not empirical hypothesis-test results or a recommended sample size. Means, missing data, ordinal items, clustering, and robust corrections are outside this example.

The one-factor model converges but cannot reproduce the two blocks of relationships. At the selected solution it underpredicts within-block covariances and overpredicts between-block covariances. This pattern offers a measurement explanation for the global misfit. A large residual-covariance modification index does not by itself tell us whether a method effect or a missing factor is the substantive explanation.

The two-factor model reproduces the matrix because that matrix was constructed from it. Its perfect fit verifies the implementation and the stipulated covariance structure; it is not independent evidence that a real instrument has two factors.

Python uses a compact positive-loading normal-ML implementation with multiple starts. Two equally good one-factor solutions can exchange the roles of the item blocks; fit statistics agree, while within-block residuals may exchange positions. Python does not reproduce lavaan modification indices, standard errors, or robust statistics. Both implementations use unit factor variances, matching covariance input, and matching fit-statistic conventions.

Verified conventional normal-ML fit for controlled covariance input
ModelChi-square (df)CFIRMSEASRMR
One factor201.810 (9).665980.231427.129959
Two correlated factors≈0 (8)10≈0

See it in R and Python

R requires lavaan (install.packages("lavaan")). Python requires NumPy and SciPy: python -m pip install numpy scipy. Both scripts construct the same covariance matrix; no CSV is needed.

# Case 009: controlled covariance CFA example, not empirical observations.
suppressPackageStartupMessages(library(lavaan))
S <- matrix(.4,6,6)
S[1:3,1:3] <- 1
S[4:6,4:6] <- 1
diag(S) <- 2
dimnames(S) <- list(paste0("x",1:6),paste0("x",1:6))
N <- 400L  # nominal sample size for the illustrative fit calculations
one <- cfa("f =~ x1+x2+x3+x4+x5+x6",sample.cov=S,sample.nobs=N,
 sample.cov.rescale=FALSE,std.lv=TRUE,meanstructure=FALSE,estimator="ML")
two <- cfa("f1 =~ x1+x2+x3
f2 =~ x4+x5+x6",sample.cov=S,sample.nobs=N,
 sample.cov.rescale=FALSE,std.lv=TRUE,meanstructure=FALSE,estimator="ML")
measures <- c("chisq","df","cfi","rmsea","srmr")
print(rbind(one_factor=fitMeasures(one,measures),two_factor=fitMeasures(two,measures)))
print(round(S-fitted(one)$cov,6))
print(modindices(one,sort=TRUE,maximum.number=5))
# Better fit under the stipulated construction is not independent validation.
stopifnot(lavInspect(one,"converged"),lavInspect(two,"converged"),
 lavInspect(one,"post.check"),lavInspect(two,"post.check"),
 fitMeasures(one,"chisq")>100,fitMeasures(two,"chisq")<1e-6,
 max(abs(S-fitted(two)$cov))<1e-6)
print(packageVersion("lavaan"))
sessionInfo()
"""Case 009: normal-ML covariance CFA, NumPy + SciPy.
Controlled covariance matrix; nominal N=400, not sampled observations.
Compact positive-loading example, not a general SEM package.
"""
import numpy as np
import scipy
from scipy.optimize import minimize
S=np.full((6,6),.4)
S[:3,:3]=1
S[3:,3:]=1
np.fill_diagonal(S,2)
N=400
logdetS=np.linalg.slogdet(S)[1]
def implied(p,two):
    loadings=np.exp(p[:6])
    residual=np.exp(p[6:12])
    if two:
        L=np.zeros((6,2))
        L[:3,0]=loadings[:3]
        L[3:,1]=loadings[3:]
        rho=np.tanh(p[12])
        phi=np.array([[1,rho],[rho,1]])
    else:
        L=loadings[:,None]
        phi=np.ones((1,1))
    return L@phi@L.T+np.diag(residual)
def discrepancy(p,two):
    sigma=implied(p,two)
    return np.linalg.slogdet(sigma)[1]+np.trace(np.linalg.solve(sigma,S))-logdetS-6
baseline=N*(np.log(np.diag(S)).sum()-logdetS)
results=[]
for two in [False,True]:
    start=np.r_[np.log(np.repeat(.9,6)),np.zeros(6)]
    if two: start=np.r_[start,np.arctanh(.3)]
    starts=[start]
    if not two:
        starts += [np.r_[np.log([1,1,1,.6,.6,.6]),np.zeros(6)],
                   np.r_[np.log([.6,.6,.6,1,1,1]),np.zeros(6)]]
    candidates=[minimize(discrepancy,initial,args=(two,),method="L-BFGS-B",jac="3-point",
        options=dict(ftol=1e-14,gtol=1e-8,maxiter=2000,maxls=50)) for initial in starts]
    fit=min(candidates,key=lambda result:result.fun)
    if not two:
        print("One-factor multistart discrepancies:",np.round([result.fun for result in candidates],9))
    sigma=implied(fit.x,two)
    chi=max(N*discrepancy(fit.x,two),0)
    df=21-len(start)
    cfi=1-max(chi-df,0)/max(chi-df,baseline-15,0)
    rmsea=np.sqrt(max(chi/df-1,0)/N)
    residual=(S-sigma)/np.sqrt(np.outer(np.diag(S),np.diag(S)))
    srmr=np.sqrt(np.mean(residual[np.tril_indices(6)]**2))
    print("two_factor" if two else "one_factor", "chisq df CFI RMSEA SRMR:",
          np.round([chi,df,cfi,rmsea,srmr],6))
    print("Optimizer success:",fit.success,"max numerical gradient:",np.max(np.abs(fit.jac)))
    assert fit.success and np.max(np.abs(fit.jac))<1e-5
    assert np.linalg.eigvalsh(sigma).min()>0
    results.append((chi,sigma))
print("One-factor raw covariance residuals:\n",np.round(S-results[0][1],6))
assert results[0][0]>100 and results[1][0]<1e-6
assert np.max(np.abs(S-results[1][1]))<1e-6
print("All assertions passed; NumPy",np.__version__,"SciPy",scipy.__version__)
print("Does not calculate lavaan modification indices or robust fit statistics.")

Verified: R 4.6.0 / lavaan 0.6-21 and Python 3.12.1 / NumPy 1.26.4 / SciPy 1.15.3 reproduce the displayed fit values. Both scripts passed all assertions. Python’s largest numerical gradient was below 1e−8 for both selected fits.

Download verified R script · Download verified Python script

08

What to Tell Your Committee

“The original [model and estimator] produced [fit statistics with degrees of freedom and relevant uncertainty]. We confirmed [numerical and input checks] and found [local discrepancy pattern]. We evaluated [limited revisions] because [substantive rationale]. The final specification is [original/revised/exploratory], with [remaining misfit, sensitivity, and validation limits]. Global fit alone does not establish [causal or measurement claim].”

Include the original and final model specifications, a fit comparison, the rationale for each change, and evidence about the estimates central to your question. Explain residual patterns rather than presenting a checklist of passing cutoffs.

09

When You Need More Help

Bring the research question, diagram and syntax, item definitions and scoring, estimator, sample and design information, complete output, and any revisions already attempted. Avoid sending raw participant data when model output and a deidentified summary are sufficient for the initial discussion.

Book a free consultation

Can't find a time that works? Email bookingrequest@dissertationstatshelper.com with a few times you're available and your time zone, and we'll find a time. It's a new address, so if you don't see our reply, please check your spam folder.

Related cases: SEM nonconvergence · Lower reliability than expected

Technical references: