Keep a held-out set
Write down the labels and their definitions. Have a knowledgeable person label representative records. Resolve disagreements before treating those labels as the reference. Keep a separate set that you do not use to tune instructions.
Count the errors that matter
Overall accuracy can hide failure in a small, important category. Count errors for each label, review ambiguous cases, and measure how much work remains manual at the chosen threshold. Include language and channel differences present in your real traffic.
Roll out in stages
Start in suggestion mode. Compare model decisions with existing outcomes, then automate only a reversible step. Set a clear stop condition if failure rate or review volume rises. Save the request identifiers and model name with evaluation results so future comparisons are meaningful.