01
Symptoms
You fit a binary logistic regression and see an implausibly large coefficient, enormous standard error, fitted probabilities extremely close to 0 or 1, or a zero cell in a predictor-by-outcome table. Software may print:
glm.fit: algorithm did not converge
You may instead see glm.fit: fitted probabilities numerically 0 or 1 occurred—or no warning at all. A routine can report convergence and still print an extreme coefficient from separated data. In a simple pattern, one predictor group contains only non-events and another only events; separation can also come from a continuous threshold or a multivariable combination hidden from two-way tables.
02
What This Usually Means
Complete separation occurs when a linear combination of predictors perfectly divides the observed events from the non-events. The corresponding ordinary logistic maximum-likelihood estimate is not finite: the likelihood keeps improving as the coefficient moves toward positive or negative infinity.
A large printed coefficient is therefore not a stable finite MLE. It does not prove a deterministic population relationship, and a software convergence flag does not make it reportable. Quasi-complete separation allows observations on the separating boundary; overlap is the geometry under which a finite unique ordinary MLE exists, subject to the model's rank conditions.
03
Common Causes
1. A sparse zero cell
One categorical level or combination contains events but no non-events, or the reverse. Rare outcomes and heavily stratified models make this more likely.
2. A continuous threshold
All observed events fall above a predictor value and all non-events below it, even though the predictor remains continuous.
3. Interactions or combinations
Each predictor may overlap alone while a linear combination, interaction, or finely divided category pattern separates the classes.
4. Too many sparse patterns
Many parameters, sparse combinations, or a small sample can leave coefficients supported only by deterministic cells. No universal sample-size cutoff diagnoses this.
5. Coding or construction errors
A reversed outcome, implausible category, accidental filter, duplicate, or inappropriate missing-data step can manufacture the pattern.
6. Outcome information leaks in
A predictor derived from the outcome or measured afterward can leak the answer into the model. That is a design problem, not exceptional prediction.
04
Run These Checks
Preserve the model and numerical symptoms.
formula(fit) R.version.string coef(summary(fit)) fit$converged fit$iter range(fitted(fit))What it suggests: Extreme coefficients, standard errors, probabilities, or iteration behavior justify investigation. None is a universal detector, and
converged = TRUEdoes not rule separation out.Cross-tabulate categorical predictors and the outcome.
with(dat, table(y, x, useNA = "ifany"))What it suggests: A predictor level or combination containing only one outcome reveals a simple separating pattern. Overlap in individual tables does not rule out multivariable separation.
Inspect continuous predictors by outcome.
boxplot(score ~ y, data = dat) with(dat, tapply(score, y, range, na.rm = TRUE))What it suggests: Nonoverlapping ranges can reveal a threshold. Marginal overlap still cannot exclude separation created jointly by several predictors.
Audit the data and predictor timing.
Verify the analysis population, outcome direction, reference categories, filters, missing-data handling, transformations, interactions, duplicates, and when each predictor was measured.
What it suggests: Correct miscoding, leakage, or unintended restriction before changing estimators. A method change must not conceal a data defect.
Inspect the complete model.
Review the model matrix and the combinations produced by factor coding and interactions. Use a formal detector such as
detectseparationwhen the pattern is not obvious.What it suggests: Locating the separating direction clarifies whether the source is an error, unsupported parameterization, or genuine sparse-data limitation.
Return to the intended estimand.
State the effect or prediction the dissertation is meant to estimate and which adjustment variables were prespecified.
What it suggests: If merging categories, dropping a term, or changing estimators changes the scientific question, it requires substantive and methodological justification.
05
What Not to Do
Do not raise the iteration limit and report the resulting larger coefficient. Further iteration follows an unbounded direction; it does not recover a missing finite MLE.
Do not report the ordinary Wald odds ratio, interval, or p-value as trustworthy inference. Transforming an unstable printed coefficient does not repair it.
Do not delete observations or predictors simply to force overlap. That can change the population, estimand, or adjustment set while hiding the limitation.
Do not combine categories only because a cell is empty or add arbitrary constants. Any consolidation must have substantive meaning; an ad hoc correction is not a principled model.
Do not treat Firth, exact, Bayesian, and ordinary fits as interchangeable. Name and justify the method actually used.
Do not claim causation or perfect population prediction. Separation describes the observed sample under the fitted specification.
06
Treatment Options
Correct an error first
Use when: coding, leakage, filtering, duplicates, or construction produced the pattern.
Tradeoff: the corrected analysis may answer a different question from the mistaken fit, and the correction must be documented.
Use Firth bias reduction
Use when: the intended logistic model is justified and penalized-likelihood inference fits the goal.
Tradeoff: this is a Firth penalized-likelihood estimate—not the nonexistent ordinary MLE—and must be reported as such.
Use an exact method
Use when: a small, sparse conditional analysis matches the question and data structure.
Tradeoff: computation and the conditional estimand may differ from the original analysis.
Use justified Bayesian regularization
Use when: Bayesian analysis fits the plan and priors can be justified and examined.
Tradeoff: posterior results depend on the prior where the data do not bound a coefficient and require sensitivity analysis.
Simplify for substantive reasons
Use when: merged categories truly represent one construct or a term is not needed for the estimand.
Tradeoff: the contrast, adjustment set, or target population may change; outcome-driven simplification can bias conclusions.
Add information or qualify the claim
Use when: genuinely informative observations are feasible or the current design cannot support the intended effect.
Tradeoff: more of the same deterministic pattern does not guarantee overlap; descriptive reporting may be the honest endpoint.
Treatment principle: answer the research question defensibly—not merely produce a finite coefficient or remove a warning.
07
Worked Example
This deterministic example contains 40 observations. All 20 non-events have x = 0, and all 20 events have x = 1. It contrasts iteration-sensitive ordinary logistic regression with Firth penalized likelihood.
suppressPackageStartupMessages(library(logistf))
dat <- data.frame(
y = c(rep(0L, 20), rep(1L, 20)),
x = c(rep(0L, 20), rep(1L, 20))
)
fit_10 <- glm(y ~ x, binomial(), dat,
control = glm.control(maxit = 10))
fit_default <- glm(y ~ x, binomial(), dat)
fit_firth <- logistf(
y ~ x, data = dat,
plcontrol = logistpl.control(maxit = 1000)
)
table(dat$y, dat$x)
coef(fit_10)["x"]
coef(fit_default)["x"]
fit_default$converged
coef(fit_firth)
confint(fit_firth)["x", ]
Observed table: all 20 cases with x = 0 are non-events; all 20 with x = 1 are events; the other two cells are zero.
Ordinary logistic regression: slope = 23.13211 after 10 iterations and 53.13213 after the default 25; default converged = TRUE; slope SE = 112616.2.
Firth fit: intercept = −3.713572; slope = 7.427144; 95% profile penalized-likelihood interval = 4.379495 to 13.177302.
Interpretation: The zero cells establish complete separation. The growing ordinary slope and enormous standard error are numerical symptoms, not a finite ordinary MLE—even though the default run reports convergence. Firth's method returns finite penalized-likelihood results for this constructed model. That demonstration does not make Firth the correct treatment for every empirical case.
08
What to Tell Your Committee
Defend the reasoning, not just the final command:
- Original target: state the outcome, focal effect or prediction, adjustment set, interactions, and role of each term.
- Diagnosis: report the numerical symptoms and show the tables, ranges, model-matrix combination, or formal result that located separation.
- Data audit: explain checks for coding, reference levels, missingness, filtering, duplicates, predictor timing, leakage, and sparse interactions.
- Selected response: justify the correction, Firth fit, exact or Bayesian approach, specification change, or qualified reporting against the estimand.
- Method change: name the estimator and interval actually used rather than describing a penalized, conditional, or posterior result as ordinary maximum likelihood.
- Limitations: state that sample separation does not establish causation or deterministic population prediction and report sensitivity to defensible assumptions.
A defensible explanation shows why the ordinary MLE was not finite, what produced the separating pattern, which inferential target mattered, and why the chosen method fits that target.