Not “did it drift?” but “when, how, and what should we do?”
Most drift checks compare two calibrations at equating time. Continuous testing programs and pre-equated banks need ongoing surveillance. driftwatch borrows sequential change detection from statistical process control.
library(driftwatch)
sim <- dw_simulate(seed = 1) # or your responses + bank
est <- dw_estimate(sim$responses, sim$bank) # item x window estimates
tu <- dw_tune(est, target = 0.01) # threshold for 1% false alarms/item
mon <- dw_monitor(est, h = tu$h) # CUSUM + change points + type
mon
dw_actions(mon, anchors = my_anchor_ids) # recommendations + audit log
dw_impact(form_ids, sim$bank, current_b, flagged) # score / pass-rate impactFrom CRAN (once released):
install.packages("driftwatch")Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/driftwatch")z = (b_t - b_bank) / sqrt(SE_t^2 + SE_bank^2).z (reference value
k, threshold h).dw_tune() regenerates responses under no drift for the
program’s actual items, windows, sample sizes and examinees, then reruns
estimation and CUSUM. This matters: each item’s bank error is shared by
every window and accumulates in the CUSUM. The design-based threshold (h
≈ 9.4) is far above the iid-normal one (≈ 6.9), and only the former hits
the false-alarm target.undetermined
rather than guessed.inst/validation/known_truth.R)300 items, 40 windows, ~80 responses per item per window. 10% of items drift gradually (0.02–0.06 logits/window) and 5% jump (0.4–1.0 logits).
| continuous (driftwatch) | two-point (first vs last 5 windows) | |
|---|---|---|
| false alarms, stable items (target 1%) | 1.3% | — |
| abrupt detected | 100%, median 3.4 windows after onset | 74%, at the end |
| gradual detected | 89%, median 12 windows after onset | 78%, at the end |
| abrupt onset error (windows) | 0.6 | — |
Drift-type classification:
| data used | abrupt: typed / accuracy | gradual: typed / accuracy |
|---|---|---|
| alarm + 3 windows | 79% / 98% | 47% / 65% |
| all 40 windows | 93% / 100% | 74% / 88% |
Gradual drift is hard to type soon after an alarm, because a short ramp looks like a step. Reclassify as windows accumulate.
Score impact (60-item form with 12 drifted items, cut at theta = 0.5): keeping banked values mis-states the pass rate by -1.3 points; recalibrating flagged items cuts that to -0.2 points, and removing them to -0.15.
Done: dw_simulate, dw_estimate,
dw_tune, dw_monitor, dw_impact,
dw_actions, dw_twopoint. Rasch difficulty
only; examinee ability treated as known. Next: 2PL discrimination drift,
ability uncertainty, anchor-set re-linking after removals, and a
per-item run-length (ARL) view.