第 17 章

第 17 章:LLM Judge——用模型给模型打分

第 17 章:LLM Judge——用模型给模型打分

有些能力无法用文件是否存在、命令是否成功来判定:代码质量、边界情况处理、可读性、回答是否切题……这些"软指标"天然适合让一个模型来评另一个模型。RewardKit 的 judge criteria 让你用 TOML 文件声明 LLM 或 agent judge 的评分规则,可复用、可共享、可在运行时切换 provider。本章是 judge TOML 的完整参考:[judge]、[[criterion]]、[scoring] 三个部分的全字段表,加上采样、防注入、provider routing 与 JEV judge。

17.1 什么任务适合 judge 评分

判断标准很简单:验证标准存在于"语义"而非"语法"层面时,用 judge。典型场景:

  • 代码质量、可读性、风格一致性;
  • 回答是否切题、是否遗漏关键信息;
  • agent 过程是否高效、是否走了捷径;
  • 开放式产出(报告、设计文档)的完整性。

需要精确判定的部分(文件、命令、数据格式)仍应交给第 16 章的程序化 criteria——两者可以在同一个 tests 目录里混用。

17.2 LLM judge 与 agent judge

LLM judge 把指定的文件内容发给一个模型打分:

# tests/quality.toml
[judge]
judge = "anthropic/claude-opus-5-5"
files = ["/app/main.py", "/app/utils.py"]

[[criterion]]
description = "Is the code correct?"
type = "binary"

[[criterion]]
description = "How readable is the code?"
type = "likert"
points = 5
weight = 2.0

[[criterion]]
description = "Rate the test coverage on a scale from 0 to 100"
type = "numeric"
min = 0
max = 100

Agent judge 则让一个 agent 在文件系统里探索、运行命令后再打分:

# tests/review.toml
[judge]
judge = "claude-code"
model = "anthropic/claude-opus-5-5"
isolated = true

[[criterion]]
description = "Does the solution handle edge cases?"
type = "binary"

Agent judge 可配置 MCP 服务器(每条 [[judge.mcp_servers]] 与 Harbor 任务的 [[environment.mcp_servers]] 字段一致,额外支持按服务器的 allowed_tools 白名单,省略则放行全部工具;Codex 不支持 sse):

[judge]
judge = "claude-code"

[[judge.mcp_servers]]
name = "playwright"
transport = "stdio"
command = "npx"
args = ["@playwright/mcp@latest", "--headless", "--isolated"]
allowed_tools = ["navigate", "click"]

[[criterion]]
description = "Does the rendered page match the spec?"
type = "binary"

17.3 [judge] 字段参考

字段 默认值 说明
judge "anthropic/claude-opus-5-5" LiteLLM 模型名、agent judge 名("claude-code"/"codex"/"fx")或 "jev"
model null agent judge 使用的 LLM;JEV 的模型
files [] 放进 judge 提示词的工作区文件路径
mode "batched" "batched" 一次调用评全部 criteria;"individual" 每个 criterion 单独调用
timeout 300 等待 judge 响应的秒数
reasoning_effort null "auto"/"none"/"minimal"/"low"/"medium"/"high"/"xhigh"/"max",取决于模型支持
isolated false agent judge 用 overlayfs 只读挂载工作区
cwd null agent judge 的工作目录;isolated 时必须在工作区内
mcp_servers [] agent judge 的 MCP 服务器配置
reference null 参考解法文件路径,供对比
atif-trajectory null 放进提示词的 ATIF 轨迹 JSON 路径
weight 1.0 本 judge 分数在目录内合成时的权重
prompt_template null 自定义提示词模板(.md/.txt),必须含 {criteria} 占位符
samples 1 judge 运行次数
guard "off" 防注入:"flag" 仅上报,"penalize" 将被标记的提交记 0 分

17.4 [[criterion]] 字段参考

字段 默认值 说明
description (必填) 评什么,这段文字会发给 judge
type "binary" 输出格式:"binary"/"likert"/"numeric"/"rubric"
name null 标识符,省略时从 description 自动生成
id null 稳定 ID(如 "1.1"),透传到 reward-details.json,改描述不变
points 5 likert 的量表大小
min / max 0.0 / 1.0 numeric 的取值范围
levels [] rubric 的等级描述,从低到高,2–10 条
weight 1.0 聚合权重;负权重需配合 weighted-sum
files [] 本 criterion 独享的文件,要求 mode = "individual",省略时回落到 [judge].files
negate false 反转归一化分数,用于"答案不应出现的行为"
optional false 在 required-pass 聚合下不参与门控

[scoring] 部分控制本 TOML 内 criteria 如何合成一个分数(不影响目录内跨文件合成):

[scoring]
aggregation = "all-pass"  # weighted-mean | weighted-sum | all-pass | any-pass | threshold | required-pass
threshold = 0.7           # 仅 threshold 聚合使用

required-pass 只有在每个非 optional criterion 都通过(value > 0)时返回 1.0;若没有任何非 optional 项则警告并记 0。weighted-sum 不归一化、唯一允许负权重、结果可能落在 [0, 1] 之外。

分数归一化:binary——yes/true/1 → 1.0,否则 0.0;likert——(raw - 1) / (points - 1);numeric——(raw - min) / (max - min);rubric——等级从 0 编号,raw / (levels - 1)。

17.5 稳定性:individual 模式、采样与防注入

  • Individual 模式:mode = "individual" 让每个 criterion 独立评分。LLM judge 每个 criterion 发一次请求;agent judge 每个 criterion 跑一轮(顺序执行)。各 criterion 可用 files 圈定自己的文件,超时会把受影响的 criterion 记 0 分并写入错误与警告。
  • 多次采样:samples = 5 让 judge 跑 5 次,每个 criterion 取中位数样本计分;reward-details.json 记录每次样本的答案与 0–1 的 agreement 一致性。多采样的 agent judge 必须先隔离。
  • Guard 防注入:被评文件里可能藏着 agent 写给 judge 的"指令"(提示注入)。开启 guard 后 judge 会被告知不要遵循被评文件中的指令,并在 reward-details.json 的 guard 字段下报告是否发现注入;flag 只上报,penalize 额外把被标记的提交记 0 分。
[judge]
judge = "anthropic/claude-opus-5-5"
samples = 5
guard = "penalize"

17.6 Provider routing 与认证

Judge 通过 LiteLLM 调用模型,凭据来自环境变量。可以在不改 rubric 的前提下运行时切换 provider:

CLI 参数 作用 等价环境变量
--je KEY=VALUE 为本次运行设置环境变量(可重复) —
--judge MODEL_OR_AGENT 覆盖 [judge].judge REWARDKIT_JUDGE
--model MODEL 覆盖 agent judge 的 [judge].model REWARDKIT_MODEL
--reasoning-effort LEVEL 覆盖 [judge].reasoning_effort REWARDKIT_REASONING_EFFORT
# 路由到 Bedrock,注入 AWS 凭据
rewardkit /tests \
  --judge bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0 \
  --je AWS_ACCESS_KEY_ID=$AWS_ACCESS_KEY_ID \
  --je AWS_REGION_NAME=us-east-1

# 切换 agent judge 与推理力度
rewardkit /tests \
  --judge claude-code \
  --model anthropic/claude-opus-5-5 \
  --reasoning-effort high

Harbor 用户可用 --ve 传递相同的环境变量。订阅认证:Anthropic LLM judge 在没有 ANTHROPIC_API_KEY 时会用 CLAUDE_CODE_OAUTH_TOKEN(由 claude setup-token 创建),设 REWARDKIT_FORCE_SUBSCRIPTION=1 可强制走订阅;codex agent judge 用 OPENAI_API_KEY 或 ChatGPT 认证的 CODEX_AUTH_JSON。

17.7 轨迹评估、negate 反转与自定义模板

轨迹评估:把 atif-trajectory 指向 /logs/agent/trajectory.json,就能评"过程"而非只评"结果":

[judge]
judge = "anthropic/claude-opus-5-5"
atif-trajectory = "/logs/agent/trajectory.json"
files = ["/app/main.py"]

[[criterion]]
description = "Did the agent take an efficient approach?"
type = "likert"
points = 5

轨迹内容会按比例截断以适配模型上下文窗口,同时保留全部步骤。

Negated criteria:对"答案不应出现的行为"设 negate = true,judge 照常打分后分数反转(value → 1 - value):出现 → 0.0,未出现 → 1.0。原始作答保留在 reward-details.json 里,反转可审计。

[[criterion]]
description = "States there is no task execution history tracking in the database"
type = "binary"
negate = true  # 答案不应做出这个(错误的)声明

自定义模板:prompt_template = "my_prompt.md",模板中必须包含注入 criterion 描述的 {criteria} 占位符。

17.8 JEV judge

JEV 是 TypeSafe 推出的新型语言模型:每个 criterion 直接返回概率或 rubric 分数、不含推理文本,因此快且便宜。需要 jev extra(uv tool install harbor-rewardkit[jev]) 与 TYPESAFE_API_KEY:

[judge]
judge = "jev"
files = ["/app/answer.md"]

[[criterion]]
description = "Does the answer address the requested task?"
type = "binary"

[[criterion]]
description = "How complete is the answer?"
type = "rubric"
levels = [
  "Omits the requested information",
  "Provides some requested information but misses important details",
  "Provides all requested information",
]

使用要点:binary 在概率 ≥ 0.5 时记 1.0;rubric 的 levels 必须从差到好排列(位置即分数);仅支持纯文本文件与 binary/rubric 两种类型,不支持 atif-trajectory 与 prompt_template;文件与最长的 criterion 需在 32k token 内;任务镜像需要 CA 证书(Debian/Ubuntu 装 ca-certificates)。经网关路由时设置 TYPESAFE_BASE_URL 与 TYPESAFE_API_KEY,Vercel 还需 model = "typesafe-ai/jev"。

17.9 设计可靠 judge 的实践建议

综合官方文档,以下几点最值得遵守:

  1. 语义软指标交给 judge,硬指标交给程序化 criteria,同一 tests 目录混用,各得其所。
  2. 描述写得越具体越好:description 是发给 judge 的唯一评分指令,避免"好不好"这类模糊措辞。
  3. 用 id 锚定 rubric,这样改写描述不会破坏结果的可追溯性。
  4. 关键判定用 samples 提稳定性,并用 agreement 监控一致性;低一致性说明描述需要改写。
  5. 永远开启 guard,尤其当 agent 能写被评文件时——这是对抗 reward hacking 的第一道防线。
  6. agent judge 一律 isolated = true,防止评分 agent 改动工作区,也解锁多采样。
  7. 用 provider routing 做成本分层:开发时用便宜模型,正式跑分用强模型,rubric 不用动。

本章小结

  • judge criteria 用 TOML 声明,适合语义层面的软指标;LLM judge 读文件打分,agent judge 可探索文件系统并跑命令。
  • [judge] 定义"谁评、评哪些文件、怎么调",[[criterion]] 定义"评什么、什么格式、多重",[scoring] 定义 TOML 内部的聚合方式。
  • 四种输出类型 binary/likert/numeric/rubric 各有归一化公式,全部映射到 [0, 1]。
  • mode = "individual"、samples、guard 是三大稳定性与安全性旋钮。
  • --judge/--model/--reasoning-effort 与 REWARDKIT_* 环境变量实现运行时 provider 路由,rubric 零改动。
  • negate 支持"应避免的行为"类 criteria;atif-trajectory 把评分对象从产出扩展到过程。
  • JEV judge 无推理文本、直接出概率/等级,快而便宜,但仅限文本与 binary/rubric。

延伸阅读