第 17 章:LLM Judge——用模型给模型打分
第 17 章:LLM Judge——用模型给模型打分
有些能力无法用文件是否存在、命令是否成功来判定:代码质量、边界情况处理、可读性、回答是否切题……这些"软指标"天然适合让一个模型来评另一个模型。RewardKit 的 judge criteria 让你用 TOML 文件声明 LLM 或 agent judge 的评分规则,可复用、可共享、可在运行时切换 provider。本章是 judge TOML 的完整参考:
[judge]、[[criterion]]、[scoring]三个部分的全字段表,加上采样、防注入、provider routing 与 JEV judge。
17.1 什么任务适合 judge 评分
判断标准很简单:验证标准存在于"语义"而非"语法"层面时,用 judge。典型场景:
- 代码质量、可读性、风格一致性;
- 回答是否切题、是否遗漏关键信息;
- agent 过程是否高效、是否走了捷径;
- 开放式产出(报告、设计文档)的完整性。
需要精确判定的部分(文件、命令、数据格式)仍应交给第 16 章的程序化 criteria——两者可以在同一个 tests 目录里混用。
17.2 LLM judge 与 agent judge
LLM judge 把指定的文件内容发给一个模型打分:
# tests/quality.toml
[judge]
judge = "anthropic/claude-opus-5-5"
files = ["/app/main.py", "/app/utils.py"]
[[criterion]]
description = "Is the code correct?"
type = "binary"
[[criterion]]
description = "How readable is the code?"
type = "likert"
points = 5
weight = 2.0
[[criterion]]
description = "Rate the test coverage on a scale from 0 to 100"
type = "numeric"
min = 0
max = 100Agent judge 则让一个 agent 在文件系统里探索、运行命令后再打分:
# tests/review.toml
[judge]
judge = "claude-code"
model = "anthropic/claude-opus-5-5"
isolated = true
[[criterion]]
description = "Does the solution handle edge cases?"
type = "binary"Agent judge 可配置 MCP 服务器(每条 [[judge.mcp_servers]] 与 Harbor 任务的 [[environment.mcp_servers]] 字段一致,额外支持按服务器的 allowed_tools 白名单,省略则放行全部工具;Codex 不支持 sse):
[judge]
judge = "claude-code"
[[judge.mcp_servers]]
name = "playwright"
transport = "stdio"
command = "npx"
args = ["@playwright/mcp@latest", "--headless", "--isolated"]
allowed_tools = ["navigate", "click"]
[[criterion]]
description = "Does the rendered page match the spec?"
type = "binary"17.3 [judge] 字段参考
| 字段 | 默认值 | 说明 |
|---|---|---|
judge |
"anthropic/claude-opus-5-5" |
LiteLLM 模型名、agent judge 名("claude-code"/"codex"/"fx")或 "jev" |
model |
null |
agent judge 使用的 LLM;JEV 的模型 |
files |
[] |
放进 judge 提示词的工作区文件路径 |
mode |
"batched" |
"batched" 一次调用评全部 criteria;"individual" 每个 criterion 单独调用 |
timeout |
300 |
等待 judge 响应的秒数 |
reasoning_effort |
null |
"auto"/"none"/"minimal"/"low"/"medium"/"high"/"xhigh"/"max",取决于模型支持 |
isolated |
false |
agent judge 用 overlayfs 只读挂载工作区 |
cwd |
null |
agent judge 的工作目录;isolated 时必须在工作区内 |
mcp_servers |
[] |
agent judge 的 MCP 服务器配置 |
reference |
null |
参考解法文件路径,供对比 |
atif-trajectory |
null |
放进提示词的 ATIF 轨迹 JSON 路径 |
weight |
1.0 |
本 judge 分数在目录内合成时的权重 |
prompt_template |
null |
自定义提示词模板(.md/.txt),必须含 {criteria} 占位符 |
samples |
1 |
judge 运行次数 |
guard |
"off" |
防注入:"flag" 仅上报,"penalize" 将被标记的提交记 0 分 |
17.4 [[criterion]] 字段参考
| 字段 | 默认值 | 说明 |
|---|---|---|
description |
(必填) | 评什么,这段文字会发给 judge |
type |
"binary" |
输出格式:"binary"/"likert"/"numeric"/"rubric" |
name |
null |
标识符,省略时从 description 自动生成 |
id |
null |
稳定 ID(如 "1.1"),透传到 reward-details.json,改描述不变 |
points |
5 |
likert 的量表大小 |
min / max |
0.0 / 1.0 |
numeric 的取值范围 |
levels |
[] |
rubric 的等级描述,从低到高,2–10 条 |
weight |
1.0 |
聚合权重;负权重需配合 weighted-sum |
files |
[] |
本 criterion 独享的文件,要求 mode = "individual",省略时回落到 [judge].files |
negate |
false |
反转归一化分数,用于"答案不应出现的行为" |
optional |
false |
在 required-pass 聚合下不参与门控 |
[scoring] 部分控制本 TOML 内 criteria 如何合成一个分数(不影响目录内跨文件合成):
[scoring]
aggregation = "all-pass" # weighted-mean | weighted-sum | all-pass | any-pass | threshold | required-pass
threshold = 0.7 # 仅 threshold 聚合使用required-pass 只有在每个非 optional criterion 都通过(value > 0)时返回 1.0;若没有任何非 optional 项则警告并记 0。weighted-sum 不归一化、唯一允许负权重、结果可能落在 [0, 1] 之外。
分数归一化:binary——yes/true/1 → 1.0,否则 0.0;likert——(raw - 1) / (points - 1);numeric——(raw - min) / (max - min);rubric——等级从 0 编号,raw / (levels - 1)。
17.5 稳定性:individual 模式、采样与防注入
- Individual 模式:
mode = "individual"让每个 criterion 独立评分。LLM judge 每个 criterion 发一次请求;agent judge 每个 criterion 跑一轮(顺序执行)。各 criterion 可用files圈定自己的文件,超时会把受影响的 criterion 记 0 分并写入错误与警告。 - 多次采样:
samples = 5让 judge 跑 5 次,每个 criterion 取中位数样本计分;reward-details.json记录每次样本的答案与 0–1 的agreement一致性。多采样的 agent judge 必须先隔离。 - Guard 防注入:被评文件里可能藏着 agent 写给 judge 的"指令"(提示注入)。开启
guard后 judge 会被告知不要遵循被评文件中的指令,并在reward-details.json的guard字段下报告是否发现注入;flag只上报,penalize额外把被标记的提交记 0 分。
[judge]
judge = "anthropic/claude-opus-5-5"
samples = 5
guard = "penalize"17.6 Provider routing 与认证
Judge 通过 LiteLLM 调用模型,凭据来自环境变量。可以在不改 rubric 的前提下运行时切换 provider:
| CLI 参数 | 作用 | 等价环境变量 |
|---|---|---|
--je KEY=VALUE |
为本次运行设置环境变量(可重复) | — |
--judge MODEL_OR_AGENT |
覆盖 [judge].judge |
REWARDKIT_JUDGE |
--model MODEL |
覆盖 agent judge 的 [judge].model |
REWARDKIT_MODEL |
--reasoning-effort LEVEL |
覆盖 [judge].reasoning_effort |
REWARDKIT_REASONING_EFFORT |
# 路由到 Bedrock,注入 AWS 凭据
rewardkit /tests \
--judge bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0 \
--je AWS_ACCESS_KEY_ID=$AWS_ACCESS_KEY_ID \
--je AWS_REGION_NAME=us-east-1
# 切换 agent judge 与推理力度
rewardkit /tests \
--judge claude-code \
--model anthropic/claude-opus-5-5 \
--reasoning-effort highHarbor 用户可用 --ve 传递相同的环境变量。订阅认证:Anthropic LLM judge 在没有 ANTHROPIC_API_KEY 时会用 CLAUDE_CODE_OAUTH_TOKEN(由 claude setup-token 创建),设 REWARDKIT_FORCE_SUBSCRIPTION=1 可强制走订阅;codex agent judge 用 OPENAI_API_KEY 或 ChatGPT 认证的 CODEX_AUTH_JSON。
17.7 轨迹评估、negate 反转与自定义模板
轨迹评估:把 atif-trajectory 指向 /logs/agent/trajectory.json,就能评"过程"而非只评"结果":
[judge]
judge = "anthropic/claude-opus-5-5"
atif-trajectory = "/logs/agent/trajectory.json"
files = ["/app/main.py"]
[[criterion]]
description = "Did the agent take an efficient approach?"
type = "likert"
points = 5轨迹内容会按比例截断以适配模型上下文窗口,同时保留全部步骤。
Negated criteria:对"答案不应出现的行为"设 negate = true,judge 照常打分后分数反转(value → 1 - value):出现 → 0.0,未出现 → 1.0。原始作答保留在 reward-details.json 里,反转可审计。
[[criterion]]
description = "States there is no task execution history tracking in the database"
type = "binary"
negate = true # 答案不应做出这个(错误的)声明自定义模板:prompt_template = "my_prompt.md",模板中必须包含注入 criterion 描述的 {criteria} 占位符。
17.8 JEV judge
JEV 是 TypeSafe 推出的新型语言模型:每个 criterion 直接返回概率或 rubric 分数、不含推理文本,因此快且便宜。需要 jev extra(uv tool install harbor-rewardkit[jev]) 与 TYPESAFE_API_KEY:
[judge]
judge = "jev"
files = ["/app/answer.md"]
[[criterion]]
description = "Does the answer address the requested task?"
type = "binary"
[[criterion]]
description = "How complete is the answer?"
type = "rubric"
levels = [
"Omits the requested information",
"Provides some requested information but misses important details",
"Provides all requested information",
]使用要点:binary 在概率 ≥ 0.5 时记 1.0;rubric 的 levels 必须从差到好排列(位置即分数);仅支持纯文本文件与 binary/rubric 两种类型,不支持 atif-trajectory 与 prompt_template;文件与最长的 criterion 需在 32k token 内;任务镜像需要 CA 证书(Debian/Ubuntu 装 ca-certificates)。经网关路由时设置 TYPESAFE_BASE_URL 与 TYPESAFE_API_KEY,Vercel 还需 model = "typesafe-ai/jev"。
17.9 设计可靠 judge 的实践建议
综合官方文档,以下几点最值得遵守:
- 语义软指标交给 judge,硬指标交给程序化 criteria,同一 tests 目录混用,各得其所。
- 描述写得越具体越好:
description是发给 judge 的唯一评分指令,避免"好不好"这类模糊措辞。 - 用
id锚定 rubric,这样改写描述不会破坏结果的可追溯性。 - 关键判定用
samples提稳定性,并用agreement监控一致性;低一致性说明描述需要改写。 - 永远开启
guard,尤其当 agent 能写被评文件时——这是对抗 reward hacking 的第一道防线。 - agent judge 一律
isolated = true,防止评分 agent 改动工作区,也解锁多采样。 - 用 provider routing 做成本分层:开发时用便宜模型,正式跑分用强模型,rubric 不用动。
本章小结
- judge criteria 用 TOML 声明,适合语义层面的软指标;LLM judge 读文件打分,agent judge 可探索文件系统并跑命令。
[judge]定义"谁评、评哪些文件、怎么调",[[criterion]]定义"评什么、什么格式、多重",[scoring]定义 TOML 内部的聚合方式。- 四种输出类型 binary/likert/numeric/rubric 各有归一化公式,全部映射到 [0, 1]。
mode = "individual"、samples、guard是三大稳定性与安全性旋钮。--judge/--model/--reasoning-effort与REWARDKIT_*环境变量实现运行时 provider 路由,rubric 零改动。negate支持"应避免的行为"类 criteria;atif-trajectory把评分对象从产出扩展到过程。- JEV judge 无推理文本、直接出概率/等级,快而便宜,但仅限文本与 binary/rubric。