Use a review band between positive and negative decisions. Set
thresholds with representative cases; do not copy a numeric
threshold as a default safety guarantee.
Measure false positives, false negatives, and calibration on
labeled statements. Recheck results when input language or the
incident definition changes.