第 18 章:实战——从零创建一个完整任务
第 18 章:实战——从零创建一个完整任务
前面各章分别讲了任务结构、agent、沙箱与 RewardKit,本章把它们串成一条完整的实战线:跟随官方 create-a-task 教程,用
harbor task init脚手架从零创建一个名为 ssh-key-pair 的任务——写指令、配元数据、搭环境、写解法与测试、本地调试、用 Oracle 校验、再用真实 agent 试跑并检查结果。章末把全书知识浓缩成一张创建任务的检查清单。
18.1 安装与脚手架(Step 0–1)
先按官方安装文档装好 Harbor,然后一条命令生成任务骨架:
harbor task init ssh-key-pair生成的目录结构就是第 3 章讲过的任务四件套加测试:
ssh-key-pair/
├── instruction.md # 任务指令
├── task.toml # 配置与元数据
├── environment/
│ └── Dockerfile # 容器定义
├── solution/
│ └── solve.sh # 解法脚本
└── tests/
├── test_outputs.py # pytest 单元测试
└── test.sh # 测试验证脚本flowchart LR
A[task init
生成骨架] --> B[写 instruction
与 task.toml]
B --> C[搭 environment]
C --> D[手动验证解法]
D --> E[写 solution 与 tests]
E --> F[Oracle 校验]
F --> G[真实 agent 试跑]
G --> H[harbor view
检查结果]18.2 写指令与元数据(Step 2–3)
instruction.md 面向 agent,要求清晰、无歧义、可直接执行:
# SSH Key Pair Generation
Generate an SSH key pair in the files `~/.ssh/id_rsa` and `~/.ssh/id_rsa.pub`.
Don't make them password protected.task.toml 声明元数据与各阶段超时:
version = "1.0"
[metadata]
author_name = "Your Name"
author_email = "your.email@example.com"
difficulty_explanation = "Simple SSH key generation command"
category = "system-administration"
tags = ["ssh", "cryptography", "linux"]
[verifier]
timeout_sec = 120.0
[agent]
timeout_sec = 120.0
[environment]
build_timeout_sec = 600.0需要 Windows 容器时在 [environment] 加 os = "windows"(默认 "linux");任务需要显式资源时加 cpus、memory_mb、storage_mb 或 gpus(详见第 11 章)。
18.3 搭建环境(Step 4)
任务的 Dockerfile 定义 agent 将通过终端交互的环境,任务需要什么依赖就装什么:
FROM ubuntu:24.04
# Create working directory
WORKDIR /app
# Install openssh-client for the task
RUN apt-get update && apt-get install -y openssh-client && rm -rf /var/lib/apt/lists/*18.4 手动验证解法(Step 5)
写自动化脚本之前,先进入容器亲手验证解法可行——这一步能省掉后面大量返工:
harbor task start-env -p ssh-key-pair -e docker -a -i # 也可用 daytona 或 modal在容器内测试解法命令(无需交互输入即可完成):
ssh-keygen -t rsa -f ~/.ssh/id_rsa -N ""
ls -l ~/.ssh/id_rsa*预期看到私钥 600、公钥 644 的权限,然后用 exit 或 Ctrl+D 退出容器。
18.5 写解法与测试(Step 6–7)
把上一步验证过的命令固化进 solution/solve.sh(Oracle agent 靠它证明任务可解),并赋予执行权限:
#!/bin/bash
ssh-keygen -t rsa -f ~/.ssh/id_rsa -N ""chmod +x ssh-key-pair/solution/solve.shtests/test.sh 负责验证 agent 是否完成任务,核心约定是:必须把 reward 写到 /logs/verifier/。本教程用 pytest,也可以换成 RewardKit(第 16–17 章)做多 criterion、加权评分或 LLM judge:
#!/bin/bash
apt-get update
apt-get install -y curl
curl -LsSf https://astral.sh/uv/0.9.5/install.sh | sh
source $HOME/.local/bin/env
# Run pytest tests
uvx \
--python 3.12 \
--with pytest==8.4.1 \
pytest /tests/test_outputs.py
# Check exit code and write reward
if [ $? -eq 0 ]; then
echo 1 > /logs/verifier/reward.txt
else
echo 0 > /logs/verifier/reward.txt
fi对应的 tests/test_outputs.py 检查三件事:文件存在、权限正确、公钥格式合法:
import os
from pathlib import Path
def test_key_files_exist() -> None:
"""Test that both private and public key files exist."""
private_key = Path.home() / ".ssh" / "id_rsa"
public_key = Path.home() / ".ssh" / "id_rsa.pub"
assert private_key.exists(), "Private key file does not exist"
assert public_key.exists(), "Public key file does not exist"
def test_key_file_permissions() -> None:
"""Test that the key files have correct permissions."""
private_key = Path.home() / ".ssh" / "id_rsa"
public_key = Path.home() / ".ssh" / "id_rsa.pub"
private_perms = oct(os.stat(private_key).st_mode)[-3:]
public_perms = oct(os.stat(public_key).st_mode)[-3:]
assert private_perms == "600", (
f"Private key has incorrect permissions: {private_perms}"
)
assert public_perms == "644", (
f"Public key has incorrect permissions: {public_perms}"
)
def test_key_format() -> None:
"""Test that the public key has the correct RSA format."""
public_key = Path.home() / ".ssh" / "id_rsa.pub"
with open(public_key, 'r') as f:
content = f.read()
assert content.startswith("ssh-rsa "), "Public key does not start with 'ssh-rsa'"
assert len(content.split()) >= 2, "Public key format is invalid"18.6 Oracle 校验与真实 agent 试跑(Step 8–9)
Oracle 校验是任务发布前的硬性关卡——解法脚本必须能拿满分:
harbor run -p ssh-key-pair -a oracle成功时输出会显示任务完成且 reward 为 1。若失败,按顺序排查:
- 解法脚本是否有执行权限;
- Dockerfile 是否装齐了依赖;
- test.sh 是否正确写入了
/logs/verifier/reward.txt; - tests 中的路径与 solution 中的路径是否一致。
Oracle 通过后,再用真实 agent 试跑(需配置对应 provider 的 API key):
harbor run \
-p ssh-key-pair \
-a claude-code \
-m anthropic/claude-haiku-4-5这一步用的是第 14 章的 harbor run:-p 指任务、-a 选 agent、-m 选模型。
18.7 检查 reward 与轨迹(Step 10)
用查看器浏览本次运行的轨迹、verifier 日志与产物:
harbor view ./jobs打开最新的 Job,选中 ssh-key-pair 的 Trial:Verifier Logs 标签页可以看到 pytest 输出与最终 reward;Trajectory 标签页逐步回放 agent 的每个动作——当真实 agent 意外失败时,这里是定位原因的第一现场(配合第 15 章的 --stream,还能在运行中实时观察)。
18.8 全书串联:任务创建检查清单
最后,把前十七章的知识浓缩成一张表。每次创建新任务,按此自检:
| # | 检查项 | 对应章节 |
|---|---|---|
| 1 | harbor task init 生成的四件套结构齐全:instruction、task.toml、environment、solution、tests |
第 3 章 |
| 2 | instruction.md 指令清晰无歧义,agent 只凭它就能完成任务 | 第 3 章 |
| 3 | task.toml 的元数据、各阶段 timeout_sec 与资源声明(cpus/memory_mb 等)合理 |
第 3、11 章 |
| 4 | Dockerfile 最小化且装齐依赖;需要网络访问的任务已配网络策略 | 第 11 章 |
| 5 | 解法先在 start-env 容器里手动验证过,再固化进 solve.sh |
第 3、18 章 |
| 6 | solve.sh 有执行权限,Oracle 跑通且 reward = 1 | 第 18 章 |
| 7 | verifier 稳定写出 /logs/verifier/ 下的 reward 文件;多维度评分考虑 RewardKit |
第 16、17 章 |
| 8 | harbor run -p ... -a ... -m ... 用真实 agent 试跑,harbor view 检查轨迹与日志 |
第 14、15 章 |
| 9 | 需要防作弊/防注入时:verifier 加 guard,必要时用隔离与独立 verifier 环境 |
第 16、17 章 |
| 10 | 大规模评测前:配置文件化、设好 -n/-k/-r、先 -l 小样本试跑 |
第 14、15 章 |
本章小结
harbor task init ssh-key-pair生成任务标准骨架:instruction.md、task.toml、environment/、solution/、tests/。- 指令写给 agent 看,元数据与超时写给 Harbor 看,两者都在第 2–3 步定型。
- 动手写自动化之前,先用
harbor task start-env -p ... -e docker -a -i进容器手动验证解法命令。 - solve.sh 是 Oracle agent 的解法,test.sh 的核心契约是把 reward 写进
/logs/verifier/。 harbor run -p ssh-key-pair -a oracle必须拿满 1 分,才算一个可用的任务。harbor view ./jobs的 Verifier Logs 与 Trajectory 标签页是检查评分与调试 agent 行为的入口。- 从任务结构、RewardKit 打分、judge 防注入到 Job 配置与并发控制,本章的检查清单把全书内容收拢为可执行的十步。