Get the 2026 ML Training Cookbook | 52 recipes — GRPO, Flow Matching, World Models, and everything in between Download Now →
Judge Models
Train evaluator models that can assess the quality, safety, and correctness of LLM outputs — replacing human evaluation at scale for data filtering, reward modeling, and automated benchmarking. Judge models are LLMs fine-tuned specifically to evaluate the quality of other model outputs. They take a (prompt, response) pair and produce a score, classification, or critique. Training uses human preference data or outputs from stronger models as reference. Modern judge models can assess multiple dimensions (helpfulness, harmlessness, correctness, style) and provide explainable judgments with reasoning. They are used for automated data filtering, as reward models for RLHF, and for benchmarking.
Papers, code, and datasets
Datasets & ModelsWant to explore this concept?
Whether you're evaluating judge models for your workflow or need help implementing it, we can help.