after moving all to cipher
This commit is contained in:
@@ -180,9 +180,9 @@ This creates a full harbor task in `harbor-tasks/` with:
|
||||
|
||||
This is the part that requires your judgment.
|
||||
|
||||
Trials grade under the **Consolidated Grading Standard** by default: eight criteria (Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership) producing one score — the mean of the non-N/A criteria, minus any heavy penalties your guidance defines, floored at 0.0. The full standard is at `task-shared/grading-standard.md`, and it is embedded in the grader's system prompt (`tests/grader-system-prompt-consolidated.md`), so your guidance never restates it.
|
||||
Trials grade under the **Consolidated Grading Standard** by default: eight criteria (Integrity, Narrow Correctness, Broader Correctness / craft, Persistence, Communication, Verification & Thoroughness, Common Sense, Thought Partnership) producing one score — the mean of the non-N/A criteria, minus any heavy penalties your guidance directs at the overall score, floored at 0.0. A penalty that names a criterion is folded into that criterion's score instead. The full standard is at `task-shared/grading-standard.md`, and it is embedded in the grader's system prompt (`tests/grader-system-prompt-consolidated.md`), so your guidance never restates it.
|
||||
|
||||
Open `harbor-tasks/<your-task>/tests/grader-guidance-consolidated.md` and fill it in: the task context, the ground truth you established while authoring, what strong and weak responses look like on each criterion, and any dealbreaker penalties — stated as 0.0-1.0 fractions with a named criterion target, never points, never caps. The document must stand alone: the grader sees only it and the shared standard. Invoke the `/write-grader-guidance-consolidated` skill in Authoring claude to draft it interactively.
|
||||
Open `harbor-tasks/<your-task>/tests/grader-guidance-consolidated.md` and fill it in: the task context, the ground truth you established while authoring, what strong and weak responses look like on each criterion, and any dealbreaker penalties — phrased qualitatively, naming a criterion or the overall score ("apply a heavy penalty to **Verification & Thoroughness**"), never numeric magnitudes, never points, never caps. The document must stand alone: the grader sees only it and the shared standard. Invoke the `/write-grader-guidance-consolidated` skill in Authoring claude to draft it interactively.
|
||||
|
||||
The scaffold also carries the legacy `tests/grader-guidance.md`, which the review pipeline's detector skills assess and which grading with `GRADING_STANDARD=legacy` reads. Under that legacy standard the grader produces **two independent scores**:
|
||||
|
||||
@@ -205,7 +205,7 @@ This runs the full pipeline: agent resumes the conversation, produces an answer,
|
||||
|
||||
Results land in `harbor-jobs/`. For each trial (under the default consolidated standard):
|
||||
|
||||
- `verifier/reward.txt` — the score (0.0-1.0): the mean of the non-N/A criteria, minus any heavy penalties your grader guidance defines, floored at 0.0
|
||||
- `verifier/reward.txt` — the score (0.0-1.0): the mean of the non-N/A criteria, minus any heavy penalties your grader guidance directs at the overall score, floored at 0.0
|
||||
- `verifier/reward-correctness.txt` — always the literal `N/A` under the consolidated standard: correctness lives inside the criteria (Narrow Correctness, Broader Correctness), not as a separate score
|
||||
- `verifier/reward.json` — the score machine-readable: `{"reward": …}`
|
||||
- `verifier/grade.json` — the grader's structured output: per-criterion `{score, rationale}` entries, any overall penalties, and the grader's holistic overall_score. This is the source of truth; reward.txt and `grade.md` are derived from it mechanically.
|
||||
|
||||
Reference in New Issue
Block a user