From 70d2a7df10fb48c54d6bcfe47c2c6147124a440b Mon Sep 17 00:00:00 2001 From: Eric Bell Date: Thu, 20 Aug 2026 14:31:44 -0400 Subject: [PATCH] chore: add icm CONTEXT file, tmux convo --- CONTEXT.md | 5 + sources/tmux-convo-1.md | 537 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 542 insertions(+) create mode 100644 CONTEXT.md create mode 100644 sources/tmux-convo-1.md diff --git a/CONTEXT.md b/CONTEXT.md new file mode 100644 index 0000000..95ae3c6 --- /dev/null +++ b/CONTEXT.md @@ -0,0 +1,5 @@ +# Source Documents Folder + +The primary folder for source documents is: + +- `/home/ericbell/workspaces/dataannotation/current-project/sources` \ No newline at end of file diff --git a/sources/tmux-convo-1.md b/sources/tmux-convo-1.md new file mode 100644 index 0000000..118722e --- /dev/null +++ b/sources/tmux-convo-1.md @@ -0,0 +1,537 @@ + + pi v0.84.2 + escape interrupt · ctrl+c/ctrl+d clear/exit · / commands · ! bash · ctrl+o more + Press ctrl+o to show full startup help and loaded resources. + + Pi can explain its own features and look up its docs. Ask it how to use or extend Pi. + +[Extensions] + @ollama/pi-web-search, mode.ts + +─────────────────────────────────────────────────────────────────────────────────────────────────────────────── + What's New + + [0.84.2] - 2026-08-14 + + ### New Features + + - Fullscreen transcript search — Search and navigate matches in fullscreen mode. See TUI Fullscreen Viewport + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen + -viewport). + - Configurable default tools — Choose startup built-in tools globally or per project. See Tools + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/settings.md#tools). + - Configurable fullscreen exit output — Print the transcript or only a resume hint on exit. See Interactive + Mode + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/usage.md#interactive-mode). + + ### Added + + - Added fullscreen transcript search with Ctrl+Shift+F, incremental match highlighting, configurable search + match theme colors, and next/previous navigation with Enter/Ctrl+G and Shift+Enter/Ctrl+Shift+G. + - Added experimental strict JSON-schema constrained sampling for the default read, bash, edit, and write + tools under PI_EXPERIMENTAL=1. + - Added a fullscreen exit output setting to choose between printing the final transcript and only a session + resume hint. + - Added the defaultTools setting for configuring the initial built-in tool selection globally or per project. + - Added --use-theme to choose an initial per-run interactive theme without changing saved + settings (#7722 (https://github.com/earendil-works/pi/pull/7722) by @rwachtler + (https://github.com/rwachtler)). + - Added expandPromptTemplates to extension pi.sendUserMessage() options for explicitly dispatching commands + and expanding skills and prompt templates. See pi.sendUserMessage() + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/extensions.md#pisendusermessa + gecontent-options) (#7857 (https://github.com/earendil-works/pi/pull/7857) by @mrexodia + (https://github.com/mrexodia)). + - Added inherited createGatewayBindingFetch() for routing Cloudflare AI Gateway requests through a Workers AI + binding without an API token (#7901 (https://github.com/earendil-works/pi/pull/7901) by @Maximo-Guk + (https://github.com/Maximo-Guk)). + - Added inherited AssistantMessage.endTurn to preserve OpenAI Codex's terminal end_turn signal for + diagnostics (#7766 (https://github.com/earendil-works/pi/pull/7766)). + - Added inherited unbound single-line transcript scrolling actions for fullscreen mode. See TUI Fullscreen + Viewport + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen + -viewport) (#7903 (https://github.com/earendil-works/pi/pull/7903) by @midastruth + (https://github.com/midastruth)). + + ### Changed + + - Changed inherited Kimi Coding requests to use pi's runtime User-Agent header. + - Replaced the inherited Mistral SDK transport with a native Chat Completions HTTP stream, eliminating its + generated client and schema runtime overhead. + - Documented the generic AI_AGENT=pi process marker and how it differs from PI_CODING_AGENT=true (#7747 + (https://github.com/earendil-works/pi/issues/7747)). + - Changed inherited OpenAI Responses deferred tool loading to prefer message-anchored additional_tools where + supported while retaining tool-search and top-level fallbacks (#7709 + (https://github.com/earendil-works/pi/issues/7709)). + - Reduced inherited fullscreen rendering allocation churn by painting full-width layout rows directly instead + of recompositing them on every frame. + + ### Fixed + + - Fixed managed-tool downloads delaying TUI startup and hiding diagnostics in fullscreen mode by mounting the + TUI first and showing download progress and warnings inside it. + - Fixed opening a model selector immediately after startup cancelling and restarting the in-progress model + catalog refresh. + - Fixed inherited GitHub Copilot login triggering API rate limits while enabling model policies by limiting + concurrent policy updates (#6187 (https://github.com/earendil-works/pi/issues/6187)). + - Fixed fullscreen transcript search snapping back to the current match during manual scrolling and + fragmented mouse input leaking into the search query. + - Fixed inherited required LaTeX arguments starting on a new line being parsed as empty (#7760 + (https://github.com/earendil-works/pi/issues/7760)). + - Updated the transitive nanoid development dependency to address a denial-of-service vulnerability. + - Fixed fallback rendering for extension tool results to collapse long output and honor tool expansion (#7979 + (https://github.com/earendil-works/pi/issues/7979)). + - Fixed JSON and RPC message_update events dropping cumulative usage during streaming. See JSON Event Mode + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/json.md) and RPC + message_update + (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/rpc.md#message_update-streami + ng) (#7982 (https://github.com/earendil-works/pi/pull/7982) by @christianklotz + (https://github.com/christianklotz)). + - Fixed pi.sendMessage(..., { triggerTurn: false }) steering an active run instead of only recording the + custom message (#8022 (https://github.com/earendil-works/pi/pull/8022) by @cristinaponcela + (https://github.com/cristinaponcela)). + - Fixed the defaultTools setting dropping extension and SDK custom tools when selecting built-in defaults. + - Fixed the subagent example rejecting YAML array syntax for the tools frontmatter field (#7598 + (https://github.com/earendil-works/pi/pull/7598) by @alexsavio (https://github.com/alexsavio)). + - Fixed the subagent example dropping parent session model, thinking, and tool configuration (#7897 + (https://github.com/earendil-works/pi/pull/7897) by @virtuald (https://github.com/virtuald)). + - Fixed custom system prompts concatenating the current working directory with later appended prompt content + (#7887 (https://github.com/earendil-works/pi/pull/7887) by @distributedlock + (https://github.com/distributedlock)). + - Fixed inherited OpenAI Responses function and custom tool calls losing namespaces during streaming, + proxying, and replay (#7709 (https://github.com/earendil-works/pi/issues/7709)). + - Fixed inherited upstream request buffer failures not triggering automatic assistant retries. + - Fixed inherited built-in and custom DeepSeek API models sending output limits through an unsupported field. + - Fixed inherited Amazon Bedrock replay rejecting tool arguments that contain empty object keys while + preserving all valid nested values (#7882 (https://github.com/earendil-works/pi/pull/7882) by @muyiyr + (https://github.com/muyiyr)). + - Fixed inherited DeepSeek compatibility detection for base URLs whose hostname contains uppercase letters + (#7933 (https://github.com/earendil-works/pi/pull/7933) by @yearth (https://github.com/yearth)). + - Fixed inherited Google Generative AI and Vertex AI responses with tool calls incorrectly treating + output-limit or provider-error stops as normal tool use (#8059 + (https://github.com/earendil-works/pi/issues/8059)). + - Fixed inherited fullscreen mouse drag selection and OSC 8 link activation in terminals that report generic + SGR mouse release button codes (#7963 (https://github.com/earendil-works/pi/issues/7963)). + - Fixed inherited focused fullscreen overlays not receiving mouse wheel or viewport scroll keys such as + PageUp and PageDown (#7894 (https://github.com/earendil-works/pi/issues/7894)). + - Fixed inherited LaTeX control spaces split across line endings causing complete expressions to fall back to + raw source. + - Fixed split Alt+Enter input over SSH being misread as Escape, added PI_TUI_ESC_TIMEOUT for high-latency + terminals, and limited that timeout to lone Escape input (#7899 + (https://github.com/earendil-works/pi/pull/7899) by @powerfooI (https://github.com/powerfooI)). + - Fixed inherited idle fullscreen sessions repainting and clearing text selection when the terminal loses + focus (#7892 (https://github.com/earendil-works/pi/pull/7892) by @terrorobe + (https://github.com/terrorobe)). + - Fixed fullscreen selection copy to use the host clipboard and report failure instead of claiming success + when OSC 52 is unsupported (#8110 (https://github.com/earendil-works/pi/pull/8110) by @Panoplos + (https://github.com/Panoplos)). + +─────────────────────────────────────────────────────────────────────────────────────────────────────────────── + + Model: nvidia/llama-3.3-nemotron-super-49b-v1 + + + i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in + sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful + failures? + + + + read ... + + Validation failed for tool "read": + - path: must have required properties path + + Received arguments: + {} + + + + read sources/01-summarization.md + + + Based on the provided summaries and findings in sources/ai-version-instructions.md and + sources/01-summarization.md (assuming the latter's content is reflected in the large text block you provided, + as the initial read tool call failed due to missing path specification), here's a suggested approach to + finding meaningful failures in the project, categorized for clarity: + + ### 1. Already Identified Failures & Vulnerabilities (from sources/01-summarization.md summary) + + ┌─────────────────┬──────────────────────────────────┬──────────────────────────────────────────────────────┐ + │ Category │ Issue │ Approach to Verify/Resolve │ + ├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤ + │ Test Failure │ Hardcoded school year in │ Verify: Manually check the test failure post-June │ + │ │ teachers_controller_test │ 2026. Resolve: Replace hardcoded year with │ + │ │ │ Year.current_school_year. │ + ├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤ + │ Authorization │ IDOR in OrdersController#update │ Verify: Write integration test to exploit. Resolve: │ + │ Gap 1 │ │ Add `authorize @order │ + └─────────────────┴──────────────────────────────────┴──────────────────────────────────────────────────────┘ + + Operation aborted + + Model: nvidia/nemotron-3-nano-30b-a3b + + + i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in + sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful + failures? + + + We need to answer: "i have a project outlined in sources/ai-version-instructions.md. The repo for the project + is summarized in sources/01-summarization.md. given the instructions, how would you suggest I approach + finding meaningful failures?" The user presumably wants advice on how to approach finding meaningful failures + in code based on the summary provided. We have read the summary; it's about code review and potential issues. + The user is asking for approach suggestions: How to find meaningful failures? Should discuss systematic + approach: reading code, looking at test failures, security/policy gaps, correctness, edge cases, performance, + data integrity, user flows, etc. Possibly using static analysis, test coverage, identify high-risk areas, use + observability, monitor production, triage. May suggest reading documentation, scanning issues, looking at + failing tests. Could also suggest running the test suite, exploring failing test details, reproducing + failures, looking at related code areas. Might also suggest looking at open issues, using linters, performing + code review patterns. Perhaps also mention static typing, type checks, dependency vulnerabilities. And refer + to the summary: they identified one test failure (time bomb), two authorization gaps, lower-priority items. + So suggestion: focus on test failures, policy gaps, etc. Also mention replicating the failure scenario. + + Thus answer should be textual: We can propose steps: 1. Re-run test suite; examine failing tests. 2. Identify + areas flagged in summary (e.g., IDOR, missing authorize). 3. Use static analysis tools. 4. Simulate user + flows that could expose failures. 5. Look at logs and metrics. 6. Prioritize based on impact. And so on. + + Make answer concise but thorough. + + Below is a practical, step‑by‑step approach you can use to surface “meaningful” failures in a project like + the one described in sources/01‑summarization.md. I’m framing it as a checklist you can follow, rather than + a set of specific code‑level fixes—so you can adapt it to whatever part of the stack you’re most interested + in (Ruby on Rails, the DB schema, the business rules, the test suite, etc.). + + ──────────────────────────────────────────────────────────────────────────────── + + 1. Start with the Known Failures and Flagged Issues + + ┌─────────────────────────────────┬─────────────────────────────────────────────────────────────────────────┐ + │ What to do │ Why it matters │ + ├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤ + │ Re‑run bundle exec rails test │ The test suite is the single source of truth for what the project │ + │ (or the equivalent test │ currently considers broken. In the summary the only failing test is the │ + │ command) and capture the exact │ “time‑bomb” in teachers_controller_test. Re‑producing it will tell you │ + │ failure output. │ exactly which line/parameter is causing the problem and whether it’s a │ + │ │ regression or an inherent design flaw. │ + ├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤ + │ Create a minimal reproduction │ This isolates the failure from the whole test harness and makes it │ + │ script (or a temporary Rails │ easier to explore edge cases without re‑running the whole suite. │ + │ console session) that exercises │ │ + │ the failing test’s path. │ │ + ├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤ + │ Cross‑reference the failure │ Often the maintainers have already annotated a ticket with priority, │ + │ with the project’s issue │ intended fix, or known work‑arounds. If not, the ticket itself can │ + │ tracker (if there’s one). │ become a place to record your findings. │ + └─────────────────────────────────┴─────────────────────────────────────────────────────────────────────────┘ + + ### Quick win + + - Fix the time‑bomb by replacing the hard‑coded school‑year literal with a dynamic call + (Year.current_school_year). Verify that the fix does not break any other test. + + ──────────────────────────────────────────────────────────────────────────────── + + 2. Systematically Scan for High‑Impact Security / Authorization Gaps + + The summary highlighted two concrete IDOR‑style gaps: + + 1. Order updates without authorize @order + 2. Teachers not scoped to their own classrooms in StudentsController / ClassroomEnrollmentsController. + + How to surface similar gaps elsewhere: + + ┌─────────────────────────────────────────────────────┬─────────────────────────────────────────────────────┐ + │ Step │ Tool / Technique │ + ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ + │ a. Map all controller actions that modify domain │ grep -R "def .*update|def .*destroy" │ + │ objects (e.g., OrdersController#update, │ app/controllers/**/*.rb │ + │ StudentsController#create, any │ │ + │ *Controller#update/destroy). │ │ + ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ + │ b. Identify the policy class for each resource │ Look for app/policies/**/*.rb. │ + │ (OrderPolicy, StudentPolicy, etc.). │ │ + ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ + │ c. Check that every state‑changing action calls │ Run a static‑analysis script like rails │ + │ authorize (or verify/check) with the correct │ lint:Authorization (if you have a custom linter) or │ + │ instance variable. │ simply add a comment placeholder TODO: authorize │ + │ │ @order and search for missing ones. │ + ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ + │ d. Verify that the permitted attributes include the │ `rg "strong_parameters │ + │ user_id (or an equivalent scoping column). │ │ + ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ + │ e. Simulate an authenticated user from a different │ Use Rails console or a temporary request spec to │ + │ classroom/role and attempt the unsafe action. │ perform patch /orders/42 as a user who does not own │ + │ │ the order. │ + └─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────┘ + + ### Pattern to repeat + + For each public API endpoint or form POST/ PATCH that touches a model, ask: “If I were a different + authenticated user, could I cause an unintended state change?” Anything that returns a 200/302 without a + proper authorization check is a candidate “meaningful failure”. + + ──────────────────────────────────────────────────────────────────────────────── + + 3. Leverage Test Coverage Metrics to Prioritize + + - Run rails test:coverage (or coverage:install + coverage run) and view the HTML report. + - Focus on low‑coverage areas that logically map to risky code paths (e.g., the + Admin::PortfolioTransactionsController mentioned in the summary). + - Add a single failing test that intentionally violates the expected invariant (e.g., tries to edit a + transaction that should be immutable). If it passes, you’ve found a hidden defect. + + Why? Low coverage often indicates parts of the system that have not been exercised by the existing test + suite—exactly the sort of blind spot where subtle bugs hide. + + ──────────────────────────────────────────────────────────────────────────────── + + 4. Look for Logical Invariants Violated in Production‑Like Scenarios + + Even if tests pass, certain business rules can break under edge‑case data. Use the domain model described in + the repo to construct “what‑if” tests: + + ┌──────────────────────────────────┬────────────────────────────────────────────────────────────────────────┐ + │ Invariant │ How to test │ + ├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤ + │ Portfolio ledger is append‑only; │ Create a series of transactions that appear to double‑spend a │ + │ balance is derived from │ student’s funds by inserting a pending order, then a second order │ + │ transactions │ before the first is executed. Verify the derived balance never exceeds │ + │ │ the true ledger sum. │ + ├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤ + │ OrderExecutionJob runs at the │ Mock the scheduler (or temporarily change the cron schedule) and │ + │ correct cadence │ assert that price updates happen only when OrderExecutionJob fires. │ + ├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤ + │ Grade‑improvement deposit │ Simulate a scenario where a teacher “finalizes” a gradebook without │ + │ calculation is applied only │ the expected A/B grade and watch that no deposit is minted. │ + │ after finalization │ │ + └──────────────────────────────────┴────────────────────────────────────────────────────────────────────────┘ + + These invariants are often expressed only in design docs (orders-and-transactions.md, gradebook-earnings.md). + Treat any deviation as a meaningful failure even if the test suite doesn’t catch it. + + ──────────────────────────────────────────────────────────────────────────────── + + 5. Use Static Analysis / Security Scanners for “silent” problems + + ┌─────────────────────────────────────────────┬─────────────────────────────────────────────────────────────┐ + │ Tool │ What it catches that manual inspection might miss │ + ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ + │ brakeman (Rails security scanner) │ Potential XSS, insecure mass_assignment, missing strong │ + │ │ parameters. │ + ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ + │ rubocop --rails-security │ Unsafe SQL fragments, dangerous File.read usage, etc. │ + ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ + │ bundler-audit │ Known vulnerable gem versions. │ + ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ + │ rails db:seed:raven or similar migration │ Detect migrations that rely on hard‑coded IDs or │ + │ linters │ non‑idempotent logic. │ + └─────────────────────────────────────────────┴─────────────────────────────────────────────────────────────┘ + + Run these as part of CI and treat any warning that touches the same area as a high‑priority failure until + it’s deliberately suppressed with a documented justification. + + ──────────────────────────────────────────────────────────────────────────────── + + 6. Monetize Impact vs. Likelihood – Prioritize Fixes + + ┌─────────────────────────────────────┬──────────────────────────┬─────────────────────┬────────────────────┐ + │ Failure type │ Typical impact │ Likelihood (given │ Suggested priority │ + │ │ │ current code) │ │ + ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ + │ IDOR on order updates (from │ Moderate – a single │ Low (requires │ Medium – fix now; │ + │ summary) │ student could tamper │ authenticated │ it’s a one‑liner. │ + │ │ with another's pending │ attacker) │ │ + │ │ trade. │ │ │ + ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ + │ Teacher‑classroom scoping gap │ High – can cross │ Medium – any │ High – patch │ + │ │ classroom boundaries, │ teacher can target │ immediately. │ + │ │ affect many students. │ any classroom. │ │ + ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ + │ Missing authorization on │ Low – only reachable by │ Low‑Medium (needs │ Medium – add │ + │ PortfolioTransactionsController │ admin URL; but violates │ admin privileges) │ missing authorize │ + │ edit/update │ immutable‑ledger │ │ calls and write a │ + │ │ invariant. │ │ unit test. │ + ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ + │ Time‑bomb hard‑coded year │ Very high – breaks UI │ Certain (will │ Urgent – replace │ + │ │ for all users after a │ trigger │ with dynamic │ + │ │ certain date. │ automatically). │ calculation. │ + └─────────────────────────────────────┴──────────────────────────┴─────────────────────┴────────────────────┘ + + When you’re hunting for meaningful failures, start by tackling the failures that move you down this matrix + from high impact + high certainty toward low impact + low certainty. + + ──────────────────────────────────────────────────────────────────────────────── + + 7. Document the Findings in a Failure‑Report Format + + 1. Title – concise description (e.g., “Potential IDOR on Order#update”). + 2. Reproduction steps – exact request/rails console commands that demonstrate the problem. + 3. Code locations – file, line numbers, and policy class. + 4. Impact – who is affected, what state can be changed, downstream consequences. + 5. Mitigation – minimal code change needed and any required test additions. + 6. Follow‑up – deeper investigation needed (e.g., audit all admin controllers for missing authorizations). + + Having a structured bug report makes it easy to prioritize, delegate, and track progress. + + ──────────────────────────────────────────────────────────────────────────────── + + ### TL;DR Checklist + + 1. Run the full test suite; isolate each failing test. + 2. Map every state‑changing controller action to its policy and verify authorize calls. + 3. Check coverage; add targeted tests for low‑coverage, high‑risk areas. + 4. Create “what‑if” data scenarios that test core invariants (ledger immutability, deposit calculation). + 5. Run static‑analysis/security scanners; treat each finding as a potential failure. + 6. Prioritize by impact/likelihood and fix the most critical gaps first. + 7. Document each discovered failure in a standard report format for transparency. + + Following this workflow will help you systematically uncover meaningful failures—whether they’re outright + bugs, security‑relevant policy violations, or design inconsistencies—while keeping the effort focused on the + parts of the codebase that matter most to the project’s core domain (student earnings, portfolio management, + and order execution). + + + reading sources/task-instructions.md, how do I use the task-catalog? + + + We need to read sources/task-instructions.md to see instructions on using task-catalog. Probably there is a + docs about task-catalog usage. Let's read that file. + + + read sources/task-instructions.md + + + The user asks: "reading sources/task-instructions.md, how do I use the task-catalog?" They want guidance on + using the task-catalog based on the instructions they just read. + + We need to answer from the content of sources/task-instructions.md, explaining how to use the task-catalog. + Probablydescribe steps: exploring the catalog, finding tasks, understanding structure, using it to submit + tasks, etc. Should refer to sections like "The Task Catalog" etc. Provide concise actionable guidance. + + Thus respond with explanation, maybe referencing sections like "The Task Catalog" and steps to use it. Should + not go beyond policies. Just answer. + + How to use the Task‑Catalog according to sources/task‑instructions.md + + 1. Understand what the catalog is + - The catalog is a registry of pre‑existing failure scenarios that have already been turned into tasks. + - Each entry lives under harbor-tasks// and contains the three core files you will also create for + a new task: + - instruction.md – the prompt the AI agent receives. + - grader‑guidance‑consolidated.md – the grader‑specific guidance. + - Reference runs (recorded trials) that prove the failure actually occurs. + + + 2. Browse the catalog + - Open the Task Catalog link that the project provides in the Quick Links section. + - You’ll see a list of slugs, each representing a separate task. Click a slug to view its folder + structure. + - Look at the instruction.md and grader‑guidance‑consolidated.md to see how failures were documented and + scored. This gives you a concrete template for the kind of failure you need to capture. + + 3. Decide whether to reuse or duplicate + - Check for duplicates. Before you start a new task, verify that an existing catalog entry does not + already cover the failure you have in mind. Duplicate submissions are rejected at review. + - If you find a close match but need a different angle (e.g., a different severity, a different code + path), you can still add a new slug; just make sure the prompt and ground‑truth details are distinct. + + 4. Create your own task folder + - Follow the scaffold shown in the catalog: + ``` + harbor-tasks// + ├─ instruction.md # engineering prompt + ├─ task.toml # metadata (repo commit, browser flag, etc.) + ├─ tests/ + │ └─ grader‑guidance‑consolidated.md + └─ reference‑runs/ # recorded trials that demonstrate the failure + ``` + - The Workspace & workspace.patch section of the instructions tells you how to generate a patch that + captures any workspace changes you made while exploring. + + 5. Populate the files + - instruction.md – Write a realistic engineering request that would naturally expose the failure you + discovered. + - Keep hints out; the prompt must not point directly at the solution. + - Ensure the prompt is self‑contained (all needed context must be inside the patched workspace). + + - grader‑guidance‑consolidated.md – Translate the failure into the eight‑criterion grader language: + - Provide Task context and Business context (if any). + - Define Ground truth (exact file/line, correct fix, test that should catch it). + - For each of the eight criteria, describe what a strong vs. weak response looks like on this task. + - Add any heavy penalties only for deal‑breaker behaviours, naming the criterion they affect. + + - Reference runs – Run harbor-run (or codex/claude depending on the agent you chose) to generate recorded + trials that trigger the failure. Copy those runs into reference‑runs/ using npx tsx + scripts/copy-reference-run.ts …. + + 6. Run detectors before submitting + - The catalog workflow expects you to invoke every detector skill (e.g., /detector‑meaningful-failure, + /detector‑fact‑check‑rubric‑claims, etc.). + - Detectors will flag stale inputs, broken environment checks, cross‑task references, or missing + heavy‑penalty definitions. Fix any issues they surface, then re‑run the detectors. + + 7. Validate and package + - Execute npx tsx scripts/submit‑task.ts . + - The script will: + 1. Verify that all required files exist. + 2. Ensure the patch is clean (no unintended binaries, ignored files, or leaked snapshots). + 3. Confirm you have at least four accepted reference runs. + 4. Produce a single tarball that you upload to the platform. + + - If any warnings appear, address them before you click Submit on the platform UI. + + 8. Submit with the required metadata + - When you upload the tarball, also fill the Import/Export panel: + - Export a JSON save‑state of the task (so you can rebuild later). + - Paste your Slack thread URL (the place where you asked/answered questions). + - Indicate whether this is a complete submission or a work‑in‑progress draft. + + - If you suspect the grader scored a run incorrectly, you can flag it with the grader‑performance + checkbox; this creates a separate concern thread for the review team. + + 9. Iterate via the Feedback Loop + - After submission, reviewers will run the same detector suite on your tarball. + - Any issues they raise (e.g., stale inputs, missing heavy‑penalty definitions) will be sent back as + feedback. + - Update your task files, re‑run the detectors, and submit a new version if needed. + + ──────────────────────────────────────────────────────────────────────────────── + + ### Quick‑Reference Cheat‑Sheet + + ┌────────────────┬────────────────────────────────────────────┬─────────────────────────────────────────────┐ + │ Step │ Command / Action │ Result │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Explore │ Open the Task Catalog link → browse slugs │ See existing instruction.md & │ + │ │ │ grader‑guidance‑consolidated.md examples │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Create folder │ mkdir -p harbor-tasks// │ Scaffold for new task │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Generate patch │ Work in Explore → │ environment/workspace.patch captured │ + │ │ /create-snapshot:snapshot → │ │ + │ │ snapshot-to-task.ts │ │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Write prompt │ instruction.md → realistic, no hints │ Agent receives clear engineering request │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Write grader │ tests/grader‑guidance‑consolidated.md → 8 │ Grader knows exactly what to score │ + │ guidance │ criteria + heavy penalties │ │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Run trials │ harbor-run (or codex/claude) → copy runs │ reference-runs/ populated │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Run detectors │ /detector‑* skills │ All automated checks pass │ + ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ + │ Submit │ npx tsx scripts/submit‑task.ts → │ Task packaged & ready for review │ + │ │ upload tarball │ │ + └────────────────┴────────────────────────────────────────────┴─────────────────────────────────────────────┘ + + Follow the flow Explore → Build → Validate → Submit exactly as the Task‑Catalog section of + sources/task‑instructions.md describes, and you’ll be able to add a new, meaningful failure to the catalog + without duplicating existing work. + +─────────────────────────────────────────────────────────────────────────────────────────────────────────────── + +─────────────────────────────────────────────────────────────────────────────────────────────────────────────── +~/workspaces/dataannotation/current-project (raccoon-stocks) +↑30k ↓4.5k R4.2k CH23.3% 15.1%/131k (auto) (nvidia) nvidia/nemotron-3-nano-30b-a3b • medium +mode: implementation