--- title: "Calibrating generated items with predicted priors" output: markdown::html_format vignette: > %\VignetteIndexEntry{Calibrating generated items with predicted priors} %\VignetteEngine{knitr::knitr} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` Automatic item generation produces items faster than pretesting can calibrate them. coldstart predicts each new item's difficulty from its features, uses the prediction as a robust prior, and finds template families where the prediction cannot be trusted. ## A bank with features Legacy items have calibrated difficulties; new generated items have only features (any numeric representation works, such as embeddings from any text model or cognitive-attribute codes) and a template family. In the simulation, one family's template drifted, making its new items harder than its history suggests. ```{r} library(coldstart) sim <- cs_simulate(n_train = 400, n_new = 150, seed = 3) it <- sim$items tr <- it$set == "train" table(it$set) ``` ## Predict difficulty, with honest uncertainty ```{r} pr <- cs_predictor(it$b_legacy[tr], sim$features[tr, ], it$family[tr], seed = 1) pr pred <- predict(pr, sim$features[!tr, ], it$family[!tr]) head(pred) ``` The predictive SD is estimated out of sample. It is larger for a family never seen in training, because the family's own effect is then unknown. ## How many pretest responses do we need? ```{r} plan <- cs_plan(pred, target_sd = 0.3) summary(plan[c("n_with_prior", "n_without_prior")]) ``` ## Calibrate from a small pretest ```{r} resp <- cs_responses(sim, n_per_item = 25, seed = 4) cal <- cs_calibrate(resp, pred) truth <- it$b_true[match(cal$item, it$item)] c(baseline = sqrt(mean((cal$base_mean - truth)^2)), with_prior = sqrt(mean((cal$post_mean - truth)^2))) ``` ## Which families can we trust? ```{r} fam <- setNames(it$family[!tr], it$item[!tr]) chk <- cs_check(cs_calibrate(cs_responses(sim, 100, seed = 5), pred), fam) chk unique(it$family[it$rogue]) ``` Priors for families that fail the check are withdrawn before the final calibration: ```{r} resp100 <- cs_responses(sim, 100, seed = 5) final <- cs_calibrate(resp100, cs_distrust(pred, chk)) head(final[c("item", "n", "post_mean", "post_sd")]) ```