Annotation workbench
Keyboard-first token tagging against the guideline, with per-item confidence and a blind-shuffle test-retest mode that computes per-label self-agreement kappa. This is how a real human annotation pass gets made — the machine-drafted labels shown elsewhere on this site are NOT this pass; see /limitations.