This workflow turns raw questionnaire items into documented scale scores. It checks the data before it scores them, so a quiet scoring error does not turn into a misleading or null result.
The workflow covers:
The included dataset is synthetic. Replace it with your own CSV only after the example runs successfully.
Reverse-scoring and scale scoring are common places for silent errors. The software does not warn you when a reverse-worded item is left unreversed. This workflow makes each decision visible and shows what the error would have cost.
Edit this section before using the workflow with another dataset.
The codebook is the single source of truth for item direction and
range. Every item needs a scale, a reverse flag
(TRUE or FALSE), and a minimum and maximum.
Take these from the instrument’s manual, not from the data.
| item | scale | reverse | min | max | wording |
|---|---|---|---|---|---|
| wb1 | wellbeing | FALSE | 1 | 5 | I feel satisfied with my life. |
| wb2 | wellbeing | TRUE | 1 | 5 | I often feel discouraged. |
| wb3 | wellbeing | FALSE | 1 | 5 | I am optimistic about my future. |
| wb4 | wellbeing | TRUE | 1 | 5 | I feel worn out by my work. |
| wb5 | wellbeing | FALSE | 1 | 5 | I enjoy my daily activities. |
| wb6 | wellbeing | FALSE | 1 | 5 | I feel hopeful most days. |
| st1 | stress | FALSE | 1 | 5 | I feel overwhelmed by my workload. |
| st2 | stress | FALSE | 1 | 5 | I have trouble relaxing. |
| st3 | stress | TRUE | 1 | 5 | I feel in control of my schedule. |
| st4 | stress | FALSE | 1 | 5 | I worry about meeting deadlines. |
| st5 | stress | FALSE | 1 | 5 | I feel tense during the week. |
| so1 | support | FALSE | 1 | 5 | My advisor supports my progress. |
| so2 | support | FALSE | 1 | 5 | My peers encourage me. |
| so3 | support | FALSE | 1 | 5 | I can ask for help when I need it. |
| so4 | support | FALSE | 1 | 5 | My family understands my goals. |
| measure | value |
|---|---|
| Participants | 200 |
| Items in the codebook | 15 |
| Scales | 3 |
| Missing IDs | 0 |
| Duplicate IDs | 0 |
Item responses are converted to numbers. Any response that cannot be read as a number is listed for review.
Every nonmissing response was numeric.
A response outside the codebook range is an entry or coding error, not a score. The default action sets it to missing and records it.
| source_row | participant_id | item | value | expected_range |
|---|---|---|---|---|
| 31 | P031 | wb3 | 7 | 1 to 5 |
| 44 | P044 | st4 | 0 | 1 to 5 |
| item | missing_n | missing_percent |
|---|---|---|
| wb1 | 2 | 1.0 |
| wb2 | 1 | 0.5 |
| wb3 | 2 | 1.0 |
| wb4 | 0 | 0.0 |
| wb5 | 1 | 0.5 |
| wb6 | 0 | 0.0 |
| st1 | 0 | 0.0 |
| st2 | 1 | 0.5 |
| st3 | 0 | 0.0 |
| st4 | 1 | 0.5 |
| st5 | 0 | 0.0 |
| so1 | 0 | 0.0 |
| so2 | 0 | 0.0 |
| so3 | 1 | 0.5 |
| so4 | 0 | 0.0 |
Items measuring the same construct should relate in a consistent direction. Negative correlations among items in one scale mean some items point the other way. The table below shows each scale before any reversal.
wellbeing
| wb1 | wb2 | wb3 | wb4 | wb5 | wb6 | |
|---|---|---|---|---|---|---|
| wb1 | 1.00 | -0.51 | 0.50 | -0.52 | 0.57 | 0.45 |
| wb2 | -0.51 | 1.00 | -0.59 | 0.47 | -0.41 | -0.45 |
| wb3 | 0.50 | -0.59 | 1.00 | -0.53 | 0.44 | 0.49 |
| wb4 | -0.52 | 0.47 | -0.53 | 1.00 | -0.50 | -0.46 |
| wb5 | 0.57 | -0.41 | 0.44 | -0.50 | 1.00 | 0.49 |
| wb6 | 0.45 | -0.45 | 0.49 | -0.46 | 0.49 | 1.00 |
stress
| st1 | st2 | st3 | st4 | st5 | |
|---|---|---|---|---|---|
| st1 | 1.00 | 0.47 | -0.51 | 0.54 | 0.55 |
| st2 | 0.47 | 1.00 | -0.49 | 0.48 | 0.51 |
| st3 | -0.51 | -0.49 | 1.00 | -0.49 | -0.52 |
| st4 | 0.54 | 0.48 | -0.49 | 1.00 | 0.48 |
| st5 | 0.55 | 0.51 | -0.52 | 0.48 | 1.00 |
support
| so1 | so2 | so3 | so4 | |
|---|---|---|---|---|
| so1 | 1.00 | 0.45 | 0.38 | 0.45 |
| so2 | 0.45 | 1.00 | 0.52 | 0.46 |
| so3 | 0.38 | 0.52 | 1.00 | 0.50 |
| so4 | 0.45 | 0.46 | 0.50 | 1.00 |
Reverse-worded items correlate negatively with the others. The codebook flags those items for reversal.
After the codebook’s reversals are applied, every item should correlate positively with the rest of its scale. A negative value means the codebook flag, the item wording, or the data coding needs review.
| scale | item | codebook_reverse | item_total_r | audit_result |
|---|---|---|---|---|
| wellbeing | wb1 | FALSE | 0.67 | Consistent with the codebook |
| wellbeing | wb2 | TRUE | 0.63 | Consistent with the codebook |
| wellbeing | wb3 | FALSE | 0.67 | Consistent with the codebook |
| wellbeing | wb4 | TRUE | 0.64 | Consistent with the codebook |
| wellbeing | wb5 | FALSE | 0.62 | Consistent with the codebook |
| wellbeing | wb6 | FALSE | 0.60 | Consistent with the codebook |
| stress | st1 | FALSE | 0.66 | Consistent with the codebook |
| stress | st2 | FALSE | 0.61 | Consistent with the codebook |
| stress | st3 | TRUE | 0.63 | Consistent with the codebook |
| stress | st4 | FALSE | 0.63 | Consistent with the codebook |
| stress | st5 | FALSE | 0.65 | Consistent with the codebook |
| support | so1 | FALSE | 0.52 | Consistent with the codebook |
| support | so2 | FALSE | 0.60 | Consistent with the codebook |
| support | so3 | FALSE | 0.58 | Consistent with the codebook |
| support | so4 | FALSE | 0.59 | Consistent with the codebook |
This demonstration changes one codebook entry on purpose: item
wb2, which is reverse-worded, is marked as not reversed.
The audit flags it.
| scale | item | codebook_reverse | item_total_r | audit_result |
|---|---|---|---|---|
| wellbeing | wb2 | FALSE | -0.63 | REVIEW: negative after applying the codebook |
Correct the codebook against the instrument’s wording and manual. Do not change the codebook just to make the audit pass.
Reverse-scored items use minimum + maximum - response.
Only the items flagged in the codebook change.
| item | scale | min | max | wording |
|---|---|---|---|---|
| wb2 | wellbeing | 1 | 5 | I often feel discouraged. |
| wb4 | wellbeing | 1 | 5 | I feel worn out by my work. |
| st3 | stress | 1 | 5 | I feel in control of my schedule. |
A scale score is computed when at least 80% of its items were answered. Otherwise the score is set to missing.
| scale | items | participants_scored | all_items_answered | some_items_missing_scored | too_few_items_not_scored |
|---|---|---|---|---|---|
| wellbeing | 6 | 199 | 196 | 3 | 1 |
| stress | 5 | 200 | 198 | 2 | 0 |
| support | 4 | 199 | 199 | 0 | 1 |
| scale | n | mean | sd | minimum | maximum |
|---|---|---|---|---|---|
| wellbeing | 199 | 2.96 | 0.89 | 1.00 | 5 |
| stress | 200 | 3.02 | 0.91 | 1.00 | 5 |
| support | 199 | 2.90 | 0.87 | 1.25 | 5 |
Reliability and item statistics use participants with every item answered, after reversal. Alpha describes internal consistency under strong assumptions. It does not show that a scale is unidimensional or valid.
| item | n | mean | sd | pct_floor | pct_ceiling | corrected_item_total_r | alpha_if_deleted |
|---|---|---|---|---|---|---|---|
| wb1 | 198 | 2.94 | 1.23 | 13.6 | 13.6 | 0.67 | 0.83 |
| wb2 | 199 | 3.04 | 1.17 | 10.1 | 12.1 | 0.63 | 0.83 |
| wb3 | 198 | 2.96 | 1.20 | 11.1 | 13.1 | 0.67 | 0.83 |
| wb4 | 200 | 2.94 | 1.15 | 9.5 | 12.0 | 0.65 | 0.83 |
| wb5 | 199 | 2.91 | 1.12 | 10.6 | 9.0 | 0.62 | 0.83 |
| wb6 | 200 | 2.94 | 1.20 | 13.0 | 10.5 | 0.61 | 0.84 |
Cronbach’s alpha = 0.85
| item | n | mean | sd | pct_floor | pct_ceiling | corrected_item_total_r | alpha_if_deleted |
|---|---|---|---|---|---|---|---|
| st1 | 200 | 2.99 | 1.21 | 12.0 | 12.0 | 0.66 | 0.80 |
| st2 | 199 | 3.04 | 1.16 | 10.1 | 11.1 | 0.61 | 0.81 |
| st3 | 200 | 2.99 | 1.13 | 12.0 | 7.5 | 0.63 | 0.80 |
| st4 | 199 | 3.00 | 1.20 | 12.1 | 12.6 | 0.63 | 0.80 |
| st5 | 200 | 3.07 | 1.17 | 11.0 | 13.0 | 0.65 | 0.80 |
Cronbach’s alpha = 0.84
| item | n | mean | sd | pct_floor | pct_ceiling | corrected_item_total_r | alpha_if_deleted |
|---|---|---|---|---|---|---|---|
| so1 | 200 | 2.86 | 1.14 | 11.0 | 8.5 | 0.53 | 0.74 |
| so2 | 200 | 2.84 | 1.17 | 15.0 | 10.5 | 0.60 | 0.70 |
| so3 | 199 | 2.90 | 1.12 | 11.1 | 9.0 | 0.58 | 0.72 |
| so4 | 200 | 3.02 | 1.10 | 8.5 | 10.0 | 0.59 | 0.71 |
Cronbach’s alpha = 0.77
Review items with a low corrected item-total correlation, high floor or ceiling percentages, or an alpha-if-deleted above the scale alpha. Remove an item only with a stated rationale. Do not drop items just to raise alpha.
If the scores do not show relationships you can defend in advance, suspect the data preparation before you suspect the hypothesis. The expected signs below come from the settings, not from the data.
| scale_a | scale_b | expected_sign | observed_r | observed_sign | check |
|---|---|---|---|---|---|
| wellbeing | support | positive | 0.53 | positive | As expected |
| wellbeing | stress | negative | -0.41 | negative | As expected |
| support | stress | negative | -0.27 | negative | As expected |
This comparison scores the same items again, but ignores the codebook’s reversals. It shows how the error would have changed a regression.
| scoring | predictor | estimate | p_value |
|---|---|---|---|
| Reversed per codebook | support_score | 0.46 | < .001 |
| Reversed per codebook | stress_score | -0.28 | < .001 |
| Reversals ignored | support_score | 0.14 | < .001 |
| Reversals ignored | stress_score | -0.10 | 0.004 |
Ignoring the reversals shrinks the estimates because opposite-direction items cancel each other. The data and the hypothesis did not change. In a study with weaker effects, the same error can turn a real effect into a nonsignificant result. This comparison is a teaching display. Do not report it as a result.
No reference values were supplied. Add published or pilot means to compare.
| step | detail |
|---|---|
| Reverse-scored items | wb2, wb4, st3 |
| Responses set to missing (out of range) | 2 |
| Participants with a scale score set missing (too few items) | 2 |
| Scoring method | mean |
| Minimum share of items answered | 0.8 |
| source_row | participant_id | wellbeing_score | stress_score | support_score |
|---|---|---|---|---|
| 1 | P001 | 2.17 | 3.4 | 1.50 |
| 2 | P002 | 2.00 | 3.8 | 2.50 |
| 3 | P003 | 3.17 | 2.6 | 2.75 |
| 4 | P004 | 2.83 | 2.4 | 2.00 |
| 5 | P005 | 3.00 | 3.6 | 1.50 |
| 6 | P006 | 2.83 | 3.2 | 2.50 |
| 7 | P007 | 2.83 | 3.0 | 2.75 |
| 8 | P008 | 3.17 | 3.4 | 3.00 |
| 9 | P009 | 3.67 | 3.8 | 4.00 |
| 10 | P010 | 4.33 | 1.6 | 3.75 |
Item responses were screened against the codebook’s response ranges, and responses outside those ranges were set to missing. Item direction was checked by confirming that, after the codebook’s reversals, every item correlated positively with the rest of its scale. Reverse-worded items (wb2, wb4, st3) were reverse-scored as the minimum plus the maximum minus the response. Scale scores were computed as the mean of the items, and a score was computed only when at least 80% of a scale’s items were answered. Internal consistency was estimated with Cronbach’s alpha (wellbeing = 0.85, stress = 0.84, support = 0.77). Scale scores were checked against relationships specified in advance, and the scoring decisions were logged so each score can be reproduced from the raw items.