Write the rule before testing the model
Specify the behavior that violates your policy, permitted exceptions, and what counts as insufficient context. Avoid a single vague instruction such as “detect bad content.” Quotation, reporting abuse, and satire can change the meaning of the same phrase.
Use three outcomes
Allow covers clearly permitted text. Review covers ambiguity and missing context. Block should be reserved for clear violations supported by your policy and evaluation results. A confidence threshold adds a second review condition; it does not replace the policy.
This workflow accepts text. Do not use it as an image, audio, or video moderation system.
Audit errors in both directions
False positives silence legitimate users; false negatives expose others to harmful content. Review both, compare performance across supported languages, and provide an appeal path. Do not represent a model score as a legal finding or certainty.