Power / sample-size planning

Reliability convergence simulator

Fabricate Best–Worst responses under assumptions you choose, run them through the real scoring pipeline, and see how many annotators each survey needs before split-half reliability plateaus.

Assumptions

Pick a rater-quality preset (or custom), then run. Simulation uses the same B−W scoring + split-half pipeline as live analysis.

N sweep: 10, 20, 40, 80, 160, 300 · target SHR ≥ 0.80 · ~few seconds in-browser

Run a simulation to see per-survey reliability vs sample size, and the N where SHR first crosses 0.80.