Purpose

This workflow turns raw questionnaire items into documented scale scores. It checks the data before it scores them, so a quiet scoring error does not turn into a misleading or null result.

The workflow covers:

  1. Importing item responses and a codebook
  2. Screening items for missing codes and out-of-range responses
  3. Auditing item direction against the codebook
  4. Reverse-scoring items
  5. Computing scale scores with an explicit missing-item rule
  6. Reviewing reliability and item behavior
  7. Checking scores against relationships you already expect
  8. Exporting scored data, audit tables, and a scoring log

The included dataset is synthetic. Replace it with your own CSV only after the example runs successfully.

Reverse-scoring and scale scoring are common places for silent errors. The software does not warn you when a reverse-worded item is left unreversed. This workflow makes each decision visible and shows what the error would have cost.

User settings

Edit this section before using the workflow with another dataset.

Importing data and the codebook

Read the files

The codebook is the single source of truth for item direction and range. Every item needs a scale, a reverse flag (TRUE or FALSE), and a minimum and maximum. Take these from the instrument’s manual, not from the data.

Codebook
item scale reverse min max wording
wb1 wellbeing FALSE 1 5 I feel satisfied with my life.
wb2 wellbeing TRUE 1 5 I often feel discouraged.
wb3 wellbeing FALSE 1 5 I am optimistic about my future.
wb4 wellbeing TRUE 1 5 I feel worn out by my work.
wb5 wellbeing FALSE 1 5 I enjoy my daily activities.
wb6 wellbeing FALSE 1 5 I feel hopeful most days.
st1 stress FALSE 1 5 I feel overwhelmed by my workload.
st2 stress FALSE 1 5 I have trouble relaxing.
st3 stress TRUE 1 5 I feel in control of my schedule.
st4 stress FALSE 1 5 I worry about meeting deadlines.
st5 stress FALSE 1 5 I feel tense during the week.
so1 support FALSE 1 5 My advisor supports my progress.
so2 support FALSE 1 5 My peers encourage me.
so3 support FALSE 1 5 I can ask for help when I need it.
so4 support FALSE 1 5 My family understands my goals.

Import summary

measure value
Participants 200
Items in the codebook 15
Scales 3
Missing IDs 0
Duplicate IDs 0

Item screening

Numeric parsing

Item responses are converted to numbers. Any response that cannot be read as a number is listed for review.

Every nonmissing response was numeric.

Responses outside the codebook range

A response outside the codebook range is an entry or coding error, not a score. The default action sets it to missing and records it.

source_row participant_id item value expected_range
31 P031 wb3 7 1 to 5
44 P044 st4 0 1 to 5

Missingness by item

item missing_n missing_percent
wb1 2 1.0
wb2 1 0.5
wb3 2 1.0
wb4 0 0.0
wb5 1 0.5
wb6 0 0.0
st1 0 0.0
st2 1 0.5
st3 0 0.0
st4 1 0.5
st5 0 0.0
so1 0 0.0
so2 0 0.0
so3 1 0.5
so4 0 0.0

Direction audit

Raw inter-item correlations

Items measuring the same construct should relate in a consistent direction. Negative correlations among items in one scale mean some items point the other way. The table below shows each scale before any reversal.

wellbeing

wb1 wb2 wb3 wb4 wb5 wb6
wb1 1.00 -0.51 0.50 -0.52 0.57 0.45
wb2 -0.51 1.00 -0.59 0.47 -0.41 -0.45
wb3 0.50 -0.59 1.00 -0.53 0.44 0.49
wb4 -0.52 0.47 -0.53 1.00 -0.50 -0.46
wb5 0.57 -0.41 0.44 -0.50 1.00 0.49
wb6 0.45 -0.45 0.49 -0.46 0.49 1.00

stress

st1 st2 st3 st4 st5
st1 1.00 0.47 -0.51 0.54 0.55
st2 0.47 1.00 -0.49 0.48 0.51
st3 -0.51 -0.49 1.00 -0.49 -0.52
st4 0.54 0.48 -0.49 1.00 0.48
st5 0.55 0.51 -0.52 0.48 1.00

support

so1 so2 so3 so4
so1 1.00 0.45 0.38 0.45
so2 0.45 1.00 0.52 0.46
so3 0.38 0.52 1.00 0.50
so4 0.45 0.46 0.50 1.00

Reverse-worded items correlate negatively with the others. The codebook flags those items for reversal.

Does the codebook match the data?

After the codebook’s reversals are applied, every item should correlate positively with the rest of its scale. A negative value means the codebook flag, the item wording, or the data coding needs review.

scale item codebook_reverse item_total_r audit_result
wellbeing wb1 FALSE 0.67 Consistent with the codebook
wellbeing wb2 TRUE 0.63 Consistent with the codebook
wellbeing wb3 FALSE 0.67 Consistent with the codebook
wellbeing wb4 TRUE 0.64 Consistent with the codebook
wellbeing wb5 FALSE 0.62 Consistent with the codebook
wellbeing wb6 FALSE 0.60 Consistent with the codebook
stress st1 FALSE 0.66 Consistent with the codebook
stress st2 FALSE 0.61 Consistent with the codebook
stress st3 TRUE 0.63 Consistent with the codebook
stress st4 FALSE 0.63 Consistent with the codebook
stress st5 FALSE 0.65 Consistent with the codebook
support so1 FALSE 0.52 Consistent with the codebook
support so2 FALSE 0.60 Consistent with the codebook
support so3 FALSE 0.58 Consistent with the codebook
support so4 FALSE 0.59 Consistent with the codebook

What a codebook error looks like

This demonstration changes one codebook entry on purpose: item wb2, which is reverse-worded, is marked as not reversed. The audit flags it.

scale item codebook_reverse item_total_r audit_result
wellbeing wb2 FALSE -0.63 REVIEW: negative after applying the codebook

Correct the codebook against the instrument’s wording and manual. Do not change the codebook just to make the audit pass.

Reverse-scoring and scale scores

Reverse-scored items

Reverse-scored items use minimum + maximum - response. Only the items flagged in the codebook change.

item scale min max wording
wb2 wellbeing 1 5 I often feel discouraged.
wb4 wellbeing 1 5 I feel worn out by my work.
st3 stress 1 5 I feel in control of my schedule.

Scores and the missing-item rule

A scale score is computed when at least 80% of its items were answered. Otherwise the score is set to missing.

scale items participants_scored all_items_answered some_items_missing_scored too_few_items_not_scored
wellbeing 6 199 196 3 1
stress 5 200 198 2 0
support 4 199 199 0 1

Scale descriptives

scale n mean sd minimum maximum
wellbeing 199 2.96 0.89 1.00 5
stress 200 3.02 0.91 1.00 5
support 199 2.90 0.87 1.25 5

Item analysis

Reliability and item statistics use participants with every item answered, after reversal. Alpha describes internal consistency under strong assumptions. It does not show that a scale is unidimensional or valid.

wellbeing

item n mean sd pct_floor pct_ceiling corrected_item_total_r alpha_if_deleted
wb1 198 2.94 1.23 13.6 13.6 0.67 0.83
wb2 199 3.04 1.17 10.1 12.1 0.63 0.83
wb3 198 2.96 1.20 11.1 13.1 0.67 0.83
wb4 200 2.94 1.15 9.5 12.0 0.65 0.83
wb5 199 2.91 1.12 10.6 9.0 0.62 0.83
wb6 200 2.94 1.20 13.0 10.5 0.61 0.84

Cronbach’s alpha = 0.85

stress

item n mean sd pct_floor pct_ceiling corrected_item_total_r alpha_if_deleted
st1 200 2.99 1.21 12.0 12.0 0.66 0.80
st2 199 3.04 1.16 10.1 11.1 0.61 0.81
st3 200 2.99 1.13 12.0 7.5 0.63 0.80
st4 199 3.00 1.20 12.1 12.6 0.63 0.80
st5 200 3.07 1.17 11.0 13.0 0.65 0.80

Cronbach’s alpha = 0.84

support

item n mean sd pct_floor pct_ceiling corrected_item_total_r alpha_if_deleted
so1 200 2.86 1.14 11.0 8.5 0.53 0.74
so2 200 2.84 1.17 15.0 10.5 0.60 0.70
so3 199 2.90 1.12 11.1 9.0 0.58 0.72
so4 200 3.02 1.10 8.5 10.0 0.59 0.71

Cronbach’s alpha = 0.77

Review items with a low corrected item-total correlation, high floor or ceiling percentages, or an alpha-if-deleted above the scale alpha. Remove an item only with a stated rationale. Do not drop items just to raise alpha.

Do the scores behave as expected?

Known relationships

If the scores do not show relationships you can defend in advance, suspect the data preparation before you suspect the hypothesis. The expected signs below come from the settings, not from the data.

scale_a scale_b expected_sign observed_r observed_sign check
wellbeing support positive 0.53 positive As expected
wellbeing stress negative -0.41 negative As expected
support stress negative -0.27 negative As expected

What skipping the reversal would have cost

This comparison scores the same items again, but ignores the codebook’s reversals. It shows how the error would have changed a regression.

scoring predictor estimate p_value
Reversed per codebook support_score 0.46 < .001
Reversed per codebook stress_score -0.28 < .001
Reversals ignored support_score 0.14 < .001
Reversals ignored stress_score -0.10 0.004

Ignoring the reversals shrinks the estimates because opposite-direction items cancel each other. The data and the hypothesis did not change. In a study with weaker effects, the same error can turn a real effect into a nonsignificant result. This comparison is a teaching display. Do not report it as a result.

Comparison with reference values

No reference values were supplied. Add published or pilot means to compare.

Export

step detail
Reverse-scored items wb2, wb4, st3
Responses set to missing (out of range) 2
Participants with a scale score set missing (too few items) 2
Scoring method mean
Minimum share of items answered 0.8

Preview

source_row participant_id wellbeing_score stress_score support_score
1 P001 2.17 3.4 1.50
2 P002 2.00 3.8 2.50
3 P003 3.17 2.6 2.75
4 P004 2.83 2.4 2.00
5 P005 3.00 3.6 1.50
6 P006 2.83 3.2 2.50
7 P007 2.83 3.0 2.75
8 P008 3.17 3.4 3.00
9 P009 3.67 3.8 4.00
10 P010 4.33 1.6 3.75

Methods write-up template

Item responses were screened against the codebook’s response ranges, and responses outside those ranges were set to missing. Item direction was checked by confirming that, after the codebook’s reversals, every item correlated positively with the rest of its scale. Reverse-worded items (wb2, wb4, st3) were reverse-scored as the minimum plus the maximum minus the response. Scale scores were computed as the mean of the items, and a score was computed only when at least 80% of a scale’s items were answered. Internal consistency was estimated with Cronbach’s alpha (wellbeing = 0.85, stress = 0.84, support = 0.77). Scale scores were checked against relationships specified in advance, and the scoring decisions were logged so each score can be reproduced from the raw items.

Interpretation guidance

  • Check item direction against the instrument’s wording and manual, not only against the data.
  • A low alpha does not by itself mean a scale is poor. Examine the number of items, the item content, and whether the scale measures one construct.
  • Do not remove items to raise alpha without a stated reason. Document each change.
  • This workflow does not test factor structure, measurement invariance, or validity evidence. Use a factor-analytic workflow for those questions.
  • Keep the exported audit and scoring log with the project so your committee can reconstruct every scoring decision.

References

  • Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555
  • DeVellis, R. F., & Thorpe, C. T. (2021). Scale development: Theory and applications (5th ed.). SAGE.
  • Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach’s alpha. Psychometrika, 74(1), 107–120. https://doi.org/10.1007/s11336-008-9101-0
  • Wickham, H., et al. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https://doi.org/10.21105/joss.01686