Method

LLM-as-Judge

Canonical ID https://id.searchplex.net/llm-as-judge/

A method that uses a large language model to evaluate, label, score, or compare outputs or query-item pairs.

Also called

  • LLM Judge
  • LLM-as-a-Judge

Scope

In retrieval evaluation, LLM-as-judge methods may be used to generate or assist relevance assessments. They are not the same as human relevance judgments, qrels, or the relevance relationship being judged.

Not equivalent to

  • LLM-as-Judge is not Qrels. LLM-as-judge is a judging method; qrels are a structured set of relevance labels.
  • LLM-as-Judge is not Relevance. LLM-as-judge may estimate or label relevance, but it is not the relevance relationship itself.
  • LLM-as-Judge is not Relevance Judgment. LLM-as-judge is a method for producing or assisting judgments; a relevance judgment is the assessment itself.

Learn more

ID