Open your logs and collect 10–20 real outputs you would be embarrassed to ship. Cluster them: wrong facts, wrong format, wrong tone, unsafe content, broken tool calls. Each cluster becomes a scored category — this taxonomy is the foundation everything else stands on.
- Factual errors: invented figures, wrong citations, hallucinated APIs.
- Instruction violations: ignored constraints, wrong language, missing steps.
- Format breaks: invalid JSON, schema drift, truncated output.
- Tone misses: off-brand, condescending, or needlessly verbose.
- Safety issues: leaks, refusals that should happen (or shouldn’t).





