From 95d5787868fd46bd35f82d32b00b67031ad3f4b2 Mon Sep 17 00:00:00 2001 From: Eric Bell Date: Mon, 17 Aug 2026 15:52:25 -0400 Subject: [PATCH] rem pii --- sources/task-instructions.md | 10 +--------- 1 file changed, 1 insertion(+), 9 deletions(-) diff --git a/sources/task-instructions.md b/sources/task-instructions.md index 654804b..92bc38b 100644 --- a/sources/task-instructions.md +++ b/sources/task-instructions.md @@ -903,7 +903,7 @@ themselves. Before you click submit, export your answers using the Import/Export panel on the left-hand side of the task page. The export produces a JSON save-state file, and that file is what lets you rebuild your platform answers when you revise the task later. Export before every submission so your work is saved. If you forgot to export your task, find your previous -submission and export it from [your past responses page](https://app.dataannotation.tech/workers/past_responses). +submission and export it from [your past responses page] On the Submit page, upload the tarball, complete the feedback request field using the four-item format described below, and paste your Slack thread URL into the Slack thread URL field. The page also asks whether the task is complete. @@ -1367,7 +1367,6 @@ each criterion. **Takeaway.** Support grader guidance with the exact code paths that make the failure real. -[Download the full task](https://app.dataannotation.tech/publish/s/1ce2abf5-1c1d-446e-8dee-8ad2833d94a8.zip) ## ✅ **Example 2. A heavy penalty for load-bearing ambiguity** @@ -1405,7 +1404,6 @@ because the resulting test passes. **Takeaway.** Ambiguity can be useful when it is intentional and load-bearing. If the task tests whether the agent asks before assuming, say that directly in the grader guidance. -[Download the full task](https://app.dataannotation.tech/publish/s/2b4e4621-0a36-4801-b5e0-1da3dd27c8b8.zip) ## ✅ **Example 3. A heavy penalty for unsupported UI claims** @@ -1448,7 +1446,6 @@ in proportion. The heavy penalty prevents agents from earning high credit for co **Takeaway.** Make clear in the grader guidance which findings are central to the user's situation and which are secondary. Enforce unverified claims with an appropriately heavy penalty. -[Download the full task](https://app.dataannotation.tech/publish/s/f96c34b1-8abe-46a4-9b6a-1e87313e6758.zip) ## ✅ **Example 4. A scoping task where the agent complied too narrowly** **Context.** A worker built this task by asking the agent to support sorting perk records by the time from creation to @@ -1499,7 +1496,6 @@ done. a wide blast radius or rests on a false premise, the agent should investigate and surface that before declaring success. -[Download the full task](https://app.dataannotation.tech/publish/s/95f7db0f-1246-4c86-905d-75e0f5c3043c.zip) ## ✅ **Example 5. A task that required pushback with options on** ## **regulated data** @@ -1542,7 +1538,6 @@ asking design-shaped questions while missing the regulated-data questions that a from load-bearing risk questions. Naming risks after the file is written is not the same as surfacing them while the user can still choose safeguards. -[Download the full task](https://app.dataannotation.tech/publish/s/2aa365f5-9808-4c4c-b33b-9e23cfac172a.zip) ## ❌ **Example 6. A prompt too ambiguous to write grader guidance for** @@ -1661,6 +1656,3 @@ project Slack channel if it persists. The agent refuses or gets overly cautious on a security task Rephrase the prompt so the legitimate engineering intent is explicit, for example by naming the defensive goal of the review. If refusals persist across trials, post the task in the project Slack channel. -[Code of Conduct](https://docs.google.com/document/u/1/d/e/2PACX-1vQ3pmQsDlC3_aA4zcV1g3Bd9-KYrGPg9Z37FMt0C8UdBObA5qAa9f36uEDAtXtT_NCecyqKAH6xD7UZ/pub) -[Support](https://app.dataannotation.tech/workers/support) -© 2026 DataAnnotation. All rights reserved.