LLM-as-Judge
Using a strong model to grade another model's output.
/ quick answer
LLM-as-judge automates eval scoring: give the judge the input, the output, and a rubric, and get a score. Cheap alternative to human labeling — with known biases (length, position) to control for. Using a strong model to grade another model's output.
What is LLM-as-Judge?
LLM-as-judge automates eval scoring: give the judge the input, the output, and a rubric, and get a score. Cheap alternative to human labeling — with known biases (length, position) to control for.
What is an example of LLM-as-Judge?
GPT-4 grades 500 summaries against a rubric: relevance 1-5, faithfulness 1-5.
Why does LLM-as-Judge matter for AI and automation?
Using a strong model to grade another model's output. It connects to the workflows, prompts and tool stacks linked on this page, so you can move from definition to execution without leaving Onexial.