第 18 章

第 18 章:实战——从零创建一个完整任务

第 18 章:实战——从零创建一个完整任务

前面各章分别讲了任务结构、agent、沙箱与 RewardKit,本章把它们串成一条完整的实战线:跟随官方 create-a-task 教程,用 harbor task init 脚手架从零创建一个名为 ssh-key-pair 的任务——写指令、配元数据、搭环境、写解法与测试、本地调试、用 Oracle 校验、再用真实 agent 试跑并检查结果。章末把全书知识浓缩成一张创建任务的检查清单。

18.1 安装与脚手架(Step 0–1)

先按官方安装文档装好 Harbor,然后一条命令生成任务骨架:

harbor task init ssh-key-pair

生成的目录结构就是第 3 章讲过的任务四件套加测试:

ssh-key-pair/
├── instruction.md      # 任务指令
├── task.toml           # 配置与元数据
├── environment/
│   └── Dockerfile      # 容器定义
├── solution/
│   └── solve.sh        # 解法脚本
└── tests/
    ├── test_outputs.py # pytest 单元测试
    └── test.sh         # 测试验证脚本
flowchart LR
    A[task init
生成骨架] --> B[写 instruction
与 task.toml] B --> C[搭 environment] C --> D[手动验证解法] D --> E[写 solution 与 tests] E --> F[Oracle 校验] F --> G[真实 agent 试跑] G --> H[harbor view
检查结果]

18.2 写指令与元数据(Step 2–3)

instruction.md 面向 agent,要求清晰、无歧义、可直接执行:

# SSH Key Pair Generation

Generate an SSH key pair in the files `~/.ssh/id_rsa` and `~/.ssh/id_rsa.pub`.

Don't make them password protected.

task.toml 声明元数据与各阶段超时:

version = "1.0"

[metadata]
author_name = "Your Name"
author_email = "your.email@example.com"
difficulty_explanation = "Simple SSH key generation command"
category = "system-administration"
tags = ["ssh", "cryptography", "linux"]

[verifier]
timeout_sec = 120.0

[agent]
timeout_sec = 120.0

[environment]
build_timeout_sec = 600.0

需要 Windows 容器时在 [environment] 加 os = "windows"(默认 "linux");任务需要显式资源时加 cpus、memory_mb、storage_mb 或 gpus(详见第 11 章)。

18.3 搭建环境(Step 4)

任务的 Dockerfile 定义 agent 将通过终端交互的环境,任务需要什么依赖就装什么:

FROM ubuntu:24.04

# Create working directory
WORKDIR /app

# Install openssh-client for the task
RUN apt-get update && apt-get install -y openssh-client && rm -rf /var/lib/apt/lists/*

18.4 手动验证解法(Step 5)

写自动化脚本之前,先进入容器亲手验证解法可行——这一步能省掉后面大量返工:

harbor task start-env -p ssh-key-pair -e docker -a -i  # 也可用 daytona 或 modal

在容器内测试解法命令(无需交互输入即可完成):

ssh-keygen -t rsa -f ~/.ssh/id_rsa -N ""
ls -l ~/.ssh/id_rsa*

预期看到私钥 600、公钥 644 的权限,然后用 exit 或 Ctrl+D 退出容器。

18.5 写解法与测试(Step 6–7)

把上一步验证过的命令固化进 solution/solve.sh(Oracle agent 靠它证明任务可解),并赋予执行权限:

#!/bin/bash

ssh-keygen -t rsa -f ~/.ssh/id_rsa -N ""
chmod +x ssh-key-pair/solution/solve.sh

tests/test.sh 负责验证 agent 是否完成任务,核心约定是:必须把 reward 写到 /logs/verifier/。本教程用 pytest,也可以换成 RewardKit(第 16–17 章)做多 criterion、加权评分或 LLM judge:

#!/bin/bash

apt-get update
apt-get install -y curl

curl -LsSf https://astral.sh/uv/0.9.5/install.sh | sh
source $HOME/.local/bin/env

# Run pytest tests
uvx \
  --python 3.12 \
  --with pytest==8.4.1 \
  pytest /tests/test_outputs.py

# Check exit code and write reward
if [ $? -eq 0 ]; then
  echo 1 > /logs/verifier/reward.txt
else
  echo 0 > /logs/verifier/reward.txt
fi

对应的 tests/test_outputs.py 检查三件事:文件存在、权限正确、公钥格式合法:

import os
from pathlib import Path


def test_key_files_exist() -> None:
    """Test that both private and public key files exist."""
    private_key = Path.home() / ".ssh" / "id_rsa"
    public_key = Path.home() / ".ssh" / "id_rsa.pub"

    assert private_key.exists(), "Private key file does not exist"
    assert public_key.exists(), "Public key file does not exist"


def test_key_file_permissions() -> None:
    """Test that the key files have correct permissions."""
    private_key = Path.home() / ".ssh" / "id_rsa"
    public_key = Path.home() / ".ssh" / "id_rsa.pub"

    private_perms = oct(os.stat(private_key).st_mode)[-3:]
    public_perms = oct(os.stat(public_key).st_mode)[-3:]

    assert private_perms == "600", (
        f"Private key has incorrect permissions: {private_perms}"
    )
    assert public_perms == "644", (
        f"Public key has incorrect permissions: {public_perms}"
    )


def test_key_format() -> None:
    """Test that the public key has the correct RSA format."""
    public_key = Path.home() / ".ssh" / "id_rsa.pub"

    with open(public_key, 'r') as f:
        content = f.read()

    assert content.startswith("ssh-rsa "), "Public key does not start with 'ssh-rsa'"
    assert len(content.split()) >= 2, "Public key format is invalid"

18.6 Oracle 校验与真实 agent 试跑(Step 8–9)

Oracle 校验是任务发布前的硬性关卡——解法脚本必须能拿满分:

harbor run -p ssh-key-pair -a oracle

成功时输出会显示任务完成且 reward 为 1。若失败,按顺序排查:

  • 解法脚本是否有执行权限;
  • Dockerfile 是否装齐了依赖;
  • test.sh 是否正确写入了 /logs/verifier/reward.txt;
  • tests 中的路径与 solution 中的路径是否一致。

Oracle 通过后,再用真实 agent 试跑(需配置对应 provider 的 API key):

harbor run \
  -p ssh-key-pair \
  -a claude-code \
  -m anthropic/claude-haiku-4-5

这一步用的是第 14 章的 harbor run:-p 指任务、-a 选 agent、-m 选模型。

18.7 检查 reward 与轨迹(Step 10)

用查看器浏览本次运行的轨迹、verifier 日志与产物:

harbor view ./jobs

打开最新的 Job,选中 ssh-key-pair 的 Trial:Verifier Logs 标签页可以看到 pytest 输出与最终 reward;Trajectory 标签页逐步回放 agent 的每个动作——当真实 agent 意外失败时,这里是定位原因的第一现场(配合第 15 章的 --stream,还能在运行中实时观察)。

18.8 全书串联:任务创建检查清单

最后,把前十七章的知识浓缩成一张表。每次创建新任务,按此自检:

# 检查项 对应章节
1 harbor task init 生成的四件套结构齐全:instruction、task.toml、environment、solution、tests 第 3 章
2 instruction.md 指令清晰无歧义,agent 只凭它就能完成任务 第 3 章
3 task.toml 的元数据、各阶段 timeout_sec 与资源声明(cpus/memory_mb 等)合理 第 3、11 章
4 Dockerfile 最小化且装齐依赖;需要网络访问的任务已配网络策略 第 11 章
5 解法先在 start-env 容器里手动验证过,再固化进 solve.sh 第 3、18 章
6 solve.sh 有执行权限,Oracle 跑通且 reward = 1 第 18 章
7 verifier 稳定写出 /logs/verifier/ 下的 reward 文件;多维度评分考虑 RewardKit 第 16、17 章
8 harbor run -p ... -a ... -m ... 用真实 agent 试跑,harbor view 检查轨迹与日志 第 14、15 章
9 需要防作弊/防注入时:verifier 加 guard,必要时用隔离与独立 verifier 环境 第 16、17 章
10 大规模评测前:配置文件化、设好 -n/-k/-r、先 -l 小样本试跑 第 14、15 章

本章小结

  • harbor task init ssh-key-pair 生成任务标准骨架:instruction.md、task.toml、environment/、solution/、tests/。
  • 指令写给 agent 看,元数据与超时写给 Harbor 看,两者都在第 2–3 步定型。
  • 动手写自动化之前,先用 harbor task start-env -p ... -e docker -a -i 进容器手动验证解法命令。
  • solve.sh 是 Oracle agent 的解法,test.sh 的核心契约是把 reward 写进 /logs/verifier/。
  • harbor run -p ssh-key-pair -a oracle 必须拿满 1 分,才算一个可用的任务。
  • harbor view ./jobs 的 Verifier Logs 与 Trajectory 标签页是检查评分与调试 agent 行为的入口。
  • 从任务结构、RewardKit 打分、judge 防注入到 Job 配置与并发控制,本章的检查清单把全书内容收拢为可执行的十步。

延伸阅读