ANALYSIS CLINIC · CASE 013

CASE STATUS: DIAGNOSED

Something looks wrong with my data · Data preparation · R and Python

My Results Are Null. Is It My Data or My Hypothesis?

Before you conclude the effect isn't there, rule out a data preparation error. A coding or scoring mistake can hide a real effect and produce a convincing null result.

01

Symptoms

Your key analysis returns a small, nonsignificant effect, and the result is not what the literature, your pilot data, or your own expectations predicted. Your committee suggests exploratory analysis. Before you search for something else to report, ask a prior question: did the data enter the analysis the way you think it did?

02

What This Usually Means

A null result has two broad explanations. The effect may be small or absent in your population, or the analysis may not be measuring what you intended. The first is a finding. The second is a data problem, and it can look identical in the output.

Preparation errors are common because they are quiet. The software runs without a warning, and the number it returns looks like any other result. Many errors also push estimates toward zero. A scale that mixes items scored in opposite directions, a missing-value code treated as a real score, or a merge that pairs the wrong rows all add noise or cancel signal. Treat a surprising null as a reason to audit the data before you interpret it.

03

Common Causes

Reverse-worded items were not reverse-scored before summing or averaging them into a scale score.

Missing-value codes such as -99, 99, 999, or 9 were left in the data and treated as real values.

A merge or join paired the wrong rows, duplicated participants, or dropped many cases without notice.

A categorical variable was coded or labeled in the wrong direction, or its levels were collapsed incorrectly.

A scale score used the wrong items, the wrong number of items, or a subscale mix-up.

Rows were filtered or excluded by a rule that differs from the one in the analysis plan.

Variables were misaligned after reshaping between wide and long format.

A measure was truncated or mis-scaled, for example with a ceiling or floor, after recoding.

04

Run These Checks

1. Compare every variable against the codebook. For each variable, list its range, its labels, and its missing-value codes. A score of 99 on a 1 to 5 scale is a missing code, not a high score.

2. Run frequency tables or a summary for every analysis variable. Look for impossible values, unexpected spikes, and categories that are empty or too large.

3. Check the direction of your items. Compute the inter-item correlation matrix for each scale. If some correlations are negative, some items are probably reverse-worded and still need reversing. The Scale Scoring and Item Analysis workflow automates this check.

4. Recompute each scale score from the raw items with code you can read, and compare it with the score you used. Differences point to the step that failed.

5. Check row counts and IDs at every merge, filter, and reshape. Record the number of rows before and after, confirm each participant ID appears the expected number of times, and check that no variable was silently dropped.

6. Test the pipeline on a relationship you already know should exist, such as two measures of the same construct or a well-documented correlation. If a known relationship does not appear, suspect the data before the hypothesis.

7. Compare descriptive statistics with published norms or your pilot data. A mean far from the usual range for a validated scale deserves an explanation.

8. Write down every fix as it happens and keep the script. You will need the log to explain what changed and why.

05

What Not to Do

Do not switch to exploratory analysis before you have audited the data. A null result produced by a coding error is not evidence for a different research question.

Do not try a new set of covariates, subgroups, or outcomes until something reaches significance. That inflates false positives and hides the original problem.

Do not fix an error and then quietly report only the version that worked. Document the original problem, the correction, and the corrected analysis.

Do not delete cases or items to make a scale behave without a stated rule and a rationale.

Do not interpret the corrected result as a finding the original analysis plan did not intend. A corrected confirmatory test remains confirmatory.

06

Treatment Options

Correct the preparation error and rerun the planned analysis when the audit identifies a coding, scoring, or merge mistake. Report the corrected results as your confirmatory analysis and keep the correction in your documentation.

Report the null result as a result when the audit finds nothing wrong. Report the estimate with its confidence interval, state the power or sensitivity of the design, and consider an equivalence test when you want to claim the effect is negligible. A well-documented null is a legitimate contribution.

Add exploratory analyses only after the confirmatory result is settled, and label them as exploratory. Describe how many analyses you ran and which hypotheses they were not planned to test.

Revisit measurement when the audit shows a scale does not behave as expected. The right response may be a change to the planned scoring rule, with a justification, rather than a change to the research question.

07

Worked Example

A four-item well-being scale has two reverse-worded items (wb2 and wb4), and the data contain a real relationship between social support and well-being. The 120 rows are deterministic teaching data built from repeatable arithmetic, not a random sample, a recommended sample size, or evidence about how often this error occurs. The items correlate more strongly than most real scales do, which keeps the mechanism easy to see.

The raw correlation matrix is the warning sign: items that measure one construct should correlate positively, and wb2 and wb4 correlate negatively with the others. Summing the raw items mixes opposite directions, so the signal cancels. The estimate is 0.105 (SE 0.079, t = 1.34), which looks like no effect. After reverse-scoring items 2 and 4, every correlation is positive and the same data give 3.042 (SE 0.147, t = 20.74). The data and the hypothesis did not change. The preparation step did.

Both R and Python reproduce these numbers from the same constructed data. No CSV is required. Python needs NumPy. In real data, the codebook and the instrument's manual, not the correlation matrix alone, tell you which items to reverse.

See it in R and Python

# Case 013: a reverse-scoring error can hide a real effect. Base R only.
# Deterministic teaching data (no random numbers), so R and Python give identical results.
i <- 1:120
unit <- function(a, m) (((i * a) %% m) - (m - 1) / 2) / ((m - 1) / 2)   # repeatable values in [-1, 1]

support <- ((i * 37) %% 11 - 5) / 2.5
trait   <- 0.8 * support + 0.6 * unit(53, 13)
item    <- function(a, m) pmin(5, pmax(1, floor(3 + trait + 0.8 * unit(a, m) + 0.5)))

# Four-item well-being scale (1-5). Items 2 and 4 are reverse-worded:
# a HIGH raw score on them means LOW well-being.
d <- data.frame(support,
                wb1 = item(17, 7),
                wb2 = 6 - item(19, 7),
                wb3 = item(23, 7),
                wb4 = 6 - item(29, 7))
items <- c("wb1", "wb2", "wb3", "wb4")

# Check: inter-item correlations. Negative values show items that point the other way.
round(cor(d[items]), 2)

# The mistake: sum the raw items without reverse-scoring items 2 and 4.
d$wb_wrong <- rowSums(d[items])
wrong <- summary(lm(wb_wrong ~ support, data = d))$coefficients["support", ]

# The fix: reverse-score items 2 and 4 (minimum + maximum - response), then sum.
d$wb2r <- 6 - d$wb2
d$wb4r <- 6 - d$wb4
d$wb_right <- d$wb1 + d$wb2r + d$wb3 + d$wb4r
round(cor(d[c("wb1", "wb2r", "wb3", "wb4r")]), 2)
right <- summary(lm(wb_right ~ support, data = d))$coefficients["support", ]

print(round(rbind(unreversed = wrong, reversed = right), 3))
stopifnot(abs(wrong[["Estimate"]] - 0.105) < 5e-4, abs(right[["Estimate"]] - 3.042) < 5e-4, nrow(d) == 120L)
sessionInfo()

Verified in R 4.5.1 and NumPy 1.26.4: unreversed estimate 0.105 (SE 0.079, t = 1.342); reversed estimate 3.042 (SE 0.147, t = 20.740). All numerical assertions passed.

Download R script · Download Python script

08

What to Tell Your Committee

“Before interpreting the null result, I audited the data preparation against the codebook. I checked [ranges, missing-value codes, scale direction, merges, and row counts]. I found [the problem] and corrected it by [the specific fix]. I reran the pre-specified analysis, and the corrected estimate was [estimate with confidence interval]. The earlier result reflected [the error], and I have documented both in [the analysis log].”

If the audit finds nothing, say that directly: “I checked the preparation steps, and I found no errors. The estimate was [estimate with confidence interval], and the study had [power or sensitivity] to detect effects of [size]. Here is what the results do and do not rule out.”

Showing the audit is part of the answer. A committee that suggests exploratory analysis is usually trying to help you salvage something. A documented data check shows that you are reporting the result, not a mistake.

09

When You Need More Help

If the audit leaves you unsure whether a result is real, a data preparation review can help. Bring your codebook, your preparation script, your analysis plan, and the output that surprised you. You do the preparation, and a consultation helps you check it step by step so you can explain every decision. The pricing page describes how this support is scoped.

Book a free consultation

Prefer to run these checks in a documented workflow? The Scale Scoring and Item Analysis workflow audits item direction, reverse-scores, and computes scale scores with an explicit missing-item rule. Data Cleaning and Audit covers missing codes, ranges, and duplicate IDs before you score.

Want the steps on one page? Get the free Data Preparation Audit Checklist (PDF).

Related Cases: My Reliability Is Lower Than Expected · Missing Data: Deletion, Multiple Imputation, or FIML?