Purpose

Public preview: This page shows the analysis output, figures, diagnostics, and reporting guidance without exposing the R code. The paid version includes the complete editable R Markdown source, reusable functions, all analysis code, synthetic data, documentation, and exported results.

This workflow introduces educational measurement with dichotomous assessment items. It moves from response audits and dimensionality evidence to Rasch and two-parameter logistic IRT models, item and test information, person scores, local dependence, and differential item functioning. The synthetic assessment contains varied difficulty and discrimination, one locally dependent item pair, one uniform-DIF item, and one nonuniform-DIF item.

User settings

Simulate or import responses

Response audit

Item difficulty and missingness

Classical item audit
item p_value missing_n item_total
item01 0.822 0 0.408
item02 0.831 0 0.353
item03 0.833 0 0.403
item04 0.746 0 0.316
item05 0.608 0 0.311
item06 0.566 0 0.330
item07 0.656 0 0.368
item08 0.553 0 0.497
item09 0.531 0 0.424
item10 0.466 0 0.450
item11 0.419 0 0.406
item12 0.419 0 0.431
item13 0.388 0 0.309
item14 0.261 0 0.350
item15 0.272 0 0.341
item16 0.219 0 0.396
item17 0.222 0 0.308
item18 0.218 0 0.277

## Total scores by group # Dimensionality and local dependence {.tabset .tabset-pills} ## Tetrachoric evidence

## Parallel analysis suggests that the number of factors =  10  and the number of components =  NA
One-factor loadings
item loading
item01 0.690
item02 0.596
item03 0.675
item04 0.480
item05 0.439
item06 0.469
item07 0.534
item08 0.700
item09 0.600
item10 0.635
item11 0.575
item12 0.614
item13 0.443
item14 0.538
item15 0.519
item16 0.641
item17 0.453
item18 0.420

Parallel analysis, content evidence, residual dependence, and interpretability should be considered together. A dominant first factor does not prove strict unidimensionality.

Local dependence

Largest residual Q3 associations
item1 item2 Q3
item17 item18 0.201
item08 item12 0.036
item01 item16 0.033
item01 item11 0.030
item01 item08 0.023
item03 item08 0.018
item01 item02 0.016
item02 item14 0.011
item03 item09 0.008
item10 item11 0.007
item03 item12 0.007
item02 item16 0.005

Large residual associations can reflect shared wording, content bundles, speed, or an additional dimension. Do not automatically delete one item; first examine content and test design.

Rasch and 2PL models

Model comparison

IRT model comparison
model AIC BIC logLik
Rasch/1PL 31484.5 31586.7 -15723.2
2PL 31353.8 31547.4 -15640.9

Better relative fit does not alone justify a more flexible model. Rasch measurement imposes equal discrimination and supports a specific measurement philosophy; the 2PL allows items to discriminate differently. Choose based on purpose, evidence, sample size, and intended score use.

Item parameters

2PL item parameters
item a b g u
item01 1.699 -1.318 0 1
item02 1.355 -1.548 0 1
item03 1.694 -1.380 0 1
item04 0.932 -1.360 0 1
item05 0.798 -0.628 0 1
item06 0.869 -0.354 0 1
item07 1.063 -0.749 0 1
item08 1.667 -0.192 0 1
item09 1.240 -0.133 0 1
item10 1.369 0.136 0 1
item11 1.159 0.358 0 1
item12 1.299 0.334 0 1
item13 0.814 0.639 0 1
item14 1.086 1.180 0 1
item15 1.034 1.150 0 1
item16 1.426 1.210 0 1
item17 0.926 1.583 0 1
item18 0.820 1.771 0 1

Item characteristic curves

# Information and scoring {.tabset .tabset-pills} ## Test information ## Person estimates and person-item map Person estimates are model-dependent and uncertain, especially at score extremes. Report the estimator, scaling, conditional uncertainty, and any transformations used for stakeholders.

Differential item functioning

Uniform and nonuniform DIF

Logistic-regression DIF screening
item uniform_p nonuniform_p group_log_odds interaction uniform_adj nonuniform_adj flag
item01 0.438 0.024 NA NA 0.788 0.218 No flag
item02 0.600 0.177 NA NA 0.813 0.440 No flag
item03 0.210 0.684 NA NA 0.539 0.835 No flag
item04 0.980 0.267 NA NA 0.986 0.481 No flag
item05 0.885 0.044 NA NA 0.986 0.265 No flag
item06 0.000 0.507 NA NA 0.000 0.760 Uniform
item07 0.234 0.298 NA NA 0.539 0.488 No flag
item08 0.030 0.835 NA NA 0.179 0.835 No flag
item09 0.269 0.718 NA NA 0.539 0.835 No flag
item10 0.263 0.630 NA NA 0.539 0.835 No flag
item11 0.677 0.196 NA NA 0.813 0.440 No flag
item12 0.000 0.001 NA NA 0.003 0.022 Uniform
item13 0.558 0.829 NA NA 0.813 0.835 No flag
item14 0.095 0.802 NA NA 0.428 0.835 No flag
item15 0.986 0.253 NA NA 0.986 0.481 No flag
item16 0.564 0.102 NA NA 0.813 0.411 No flag
item17 0.657 0.137 NA NA 0.813 0.411 No flag
item18 0.228 0.124 NA NA 0.539 0.411 No flag

DIF is an item-level group difference conditional on the matching variable; it is not automatically bias. Review impact, effect size, content, translation, opportunity to learn, matching-score validity, and model choice. Iterative purification and independent replication are preferable to deleting every flagged item.

Dynamic reporting example

The 2PL model had the lower BIC in the synthetic data. Estimated item difficulties ranged from -1.55 to 1.77, and discrimination estimates ranged from 0.80 to 1.70. After false-discovery-rate adjustment at 0.01, the DIF screen flagged item06, item12. These flags require substantive and impact review rather than automatic removal.

Reporting checklist

  • Define the construct, intended score use, population, item format, administration, and stakes.
  • Audit missingness, speed, item exposure, floor/ceiling behavior, and response dependence.
  • Report dimensionality evidence and local-dependence diagnostics.
  • State the IRT model, estimator, identification, software, convergence, and parameterization.
  • Report item parameters, model fit, information, conditional precision, and score method.
  • Explain group definitions, matching criterion, DIF method, multiplicity control, effect size, and impact.
  • Separate statistical DIF from bias and document content-expert review.
  • Validate score interpretations in an independent sample before consequential use.

When this model is not enough

Use polytomous IRT for partial-credit or rating-scale items; multidimensional IRT when constructs are not essentially unidimensional; testlet or bifactor models for local dependence; multilevel IRT for students nested in schools; longitudinal IRT for growth and scale drift; and linking/equating designs for different forms. High-stakes testing requires stronger governance, security, accessibility, fairness, and validation evidence than this tutorial provides.

Export results