pi v0.84.2 escape interrupt · ctrl+c/ctrl+d clear/exit · / commands · ! bash · ctrl+o more Press ctrl+o to show full startup help and loaded resources. Pi can explain its own features and look up its docs. Ask it how to use or extend Pi. [Extensions] @ollama/pi-web-search, mode.ts ─────────────────────────────────────────────────────────────────────────────────────────────────────────────── What's New [0.84.2] - 2026-08-14 ### New Features - Fullscreen transcript search — Search and navigate matches in fullscreen mode. See TUI Fullscreen Viewport (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen -viewport). - Configurable default tools — Choose startup built-in tools globally or per project. See Tools (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/settings.md#tools). - Configurable fullscreen exit output — Print the transcript or only a resume hint on exit. See Interactive Mode (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/usage.md#interactive-mode). ### Added - Added fullscreen transcript search with Ctrl+Shift+F, incremental match highlighting, configurable search match theme colors, and next/previous navigation with Enter/Ctrl+G and Shift+Enter/Ctrl+Shift+G. - Added experimental strict JSON-schema constrained sampling for the default read, bash, edit, and write tools under PI_EXPERIMENTAL=1. - Added a fullscreen exit output setting to choose between printing the final transcript and only a session resume hint. - Added the defaultTools setting for configuring the initial built-in tool selection globally or per project. - Added --use-theme to choose an initial per-run interactive theme without changing saved settings (#7722 (https://github.com/earendil-works/pi/pull/7722) by @rwachtler (https://github.com/rwachtler)). - Added expandPromptTemplates to extension pi.sendUserMessage() options for explicitly dispatching commands and expanding skills and prompt templates. See pi.sendUserMessage() (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/extensions.md#pisendusermessa gecontent-options) (#7857 (https://github.com/earendil-works/pi/pull/7857) by @mrexodia (https://github.com/mrexodia)). - Added inherited createGatewayBindingFetch() for routing Cloudflare AI Gateway requests through a Workers AI binding without an API token (#7901 (https://github.com/earendil-works/pi/pull/7901) by @Maximo-Guk (https://github.com/Maximo-Guk)). - Added inherited AssistantMessage.endTurn to preserve OpenAI Codex's terminal end_turn signal for diagnostics (#7766 (https://github.com/earendil-works/pi/pull/7766)). - Added inherited unbound single-line transcript scrolling actions for fullscreen mode. See TUI Fullscreen Viewport (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen -viewport) (#7903 (https://github.com/earendil-works/pi/pull/7903) by @midastruth (https://github.com/midastruth)). ### Changed - Changed inherited Kimi Coding requests to use pi's runtime User-Agent header. - Replaced the inherited Mistral SDK transport with a native Chat Completions HTTP stream, eliminating its generated client and schema runtime overhead. - Documented the generic AI_AGENT=pi process marker and how it differs from PI_CODING_AGENT=true (#7747 (https://github.com/earendil-works/pi/issues/7747)). - Changed inherited OpenAI Responses deferred tool loading to prefer message-anchored additional_tools where supported while retaining tool-search and top-level fallbacks (#7709 (https://github.com/earendil-works/pi/issues/7709)). - Reduced inherited fullscreen rendering allocation churn by painting full-width layout rows directly instead of recompositing them on every frame. ### Fixed - Fixed managed-tool downloads delaying TUI startup and hiding diagnostics in fullscreen mode by mounting the TUI first and showing download progress and warnings inside it. - Fixed opening a model selector immediately after startup cancelling and restarting the in-progress model catalog refresh. - Fixed inherited GitHub Copilot login triggering API rate limits while enabling model policies by limiting concurrent policy updates (#6187 (https://github.com/earendil-works/pi/issues/6187)). - Fixed fullscreen transcript search snapping back to the current match during manual scrolling and fragmented mouse input leaking into the search query. - Fixed inherited required LaTeX arguments starting on a new line being parsed as empty (#7760 (https://github.com/earendil-works/pi/issues/7760)). - Updated the transitive nanoid development dependency to address a denial-of-service vulnerability. - Fixed fallback rendering for extension tool results to collapse long output and honor tool expansion (#7979 (https://github.com/earendil-works/pi/issues/7979)). - Fixed JSON and RPC message_update events dropping cumulative usage during streaming. See JSON Event Mode (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/json.md) and RPC message_update (https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/rpc.md#message_update-streami ng) (#7982 (https://github.com/earendil-works/pi/pull/7982) by @christianklotz (https://github.com/christianklotz)). - Fixed pi.sendMessage(..., { triggerTurn: false }) steering an active run instead of only recording the custom message (#8022 (https://github.com/earendil-works/pi/pull/8022) by @cristinaponcela (https://github.com/cristinaponcela)). - Fixed the defaultTools setting dropping extension and SDK custom tools when selecting built-in defaults. - Fixed the subagent example rejecting YAML array syntax for the tools frontmatter field (#7598 (https://github.com/earendil-works/pi/pull/7598) by @alexsavio (https://github.com/alexsavio)). - Fixed the subagent example dropping parent session model, thinking, and tool configuration (#7897 (https://github.com/earendil-works/pi/pull/7897) by @virtuald (https://github.com/virtuald)). - Fixed custom system prompts concatenating the current working directory with later appended prompt content (#7887 (https://github.com/earendil-works/pi/pull/7887) by @distributedlock (https://github.com/distributedlock)). - Fixed inherited OpenAI Responses function and custom tool calls losing namespaces during streaming, proxying, and replay (#7709 (https://github.com/earendil-works/pi/issues/7709)). - Fixed inherited upstream request buffer failures not triggering automatic assistant retries. - Fixed inherited built-in and custom DeepSeek API models sending output limits through an unsupported field. - Fixed inherited Amazon Bedrock replay rejecting tool arguments that contain empty object keys while preserving all valid nested values (#7882 (https://github.com/earendil-works/pi/pull/7882) by @muyiyr (https://github.com/muyiyr)). - Fixed inherited DeepSeek compatibility detection for base URLs whose hostname contains uppercase letters (#7933 (https://github.com/earendil-works/pi/pull/7933) by @yearth (https://github.com/yearth)). - Fixed inherited Google Generative AI and Vertex AI responses with tool calls incorrectly treating output-limit or provider-error stops as normal tool use (#8059 (https://github.com/earendil-works/pi/issues/8059)). - Fixed inherited fullscreen mouse drag selection and OSC 8 link activation in terminals that report generic SGR mouse release button codes (#7963 (https://github.com/earendil-works/pi/issues/7963)). - Fixed inherited focused fullscreen overlays not receiving mouse wheel or viewport scroll keys such as PageUp and PageDown (#7894 (https://github.com/earendil-works/pi/issues/7894)). - Fixed inherited LaTeX control spaces split across line endings causing complete expressions to fall back to raw source. - Fixed split Alt+Enter input over SSH being misread as Escape, added PI_TUI_ESC_TIMEOUT for high-latency terminals, and limited that timeout to lone Escape input (#7899 (https://github.com/earendil-works/pi/pull/7899) by @powerfooI (https://github.com/powerfooI)). - Fixed inherited idle fullscreen sessions repainting and clearing text selection when the terminal loses focus (#7892 (https://github.com/earendil-works/pi/pull/7892) by @terrorobe (https://github.com/terrorobe)). - Fixed fullscreen selection copy to use the host clipboard and report failure instead of claiming success when OSC 52 is unsupported (#8110 (https://github.com/earendil-works/pi/pull/8110) by @Panoplos (https://github.com/Panoplos)). ─────────────────────────────────────────────────────────────────────────────────────────────────────────────── Model: nvidia/llama-3.3-nemotron-super-49b-v1 i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful failures? read ... Validation failed for tool "read": - path: must have required properties path Received arguments: {} read sources/01-summarization.md Based on the provided summaries and findings in sources/ai-version-instructions.md and sources/01-summarization.md (assuming the latter's content is reflected in the large text block you provided, as the initial read tool call failed due to missing path specification), here's a suggested approach to finding meaningful failures in the project, categorized for clarity: ### 1. Already Identified Failures & Vulnerabilities (from sources/01-summarization.md summary) ┌─────────────────┬──────────────────────────────────┬──────────────────────────────────────────────────────┐ │ Category │ Issue │ Approach to Verify/Resolve │ ├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤ │ Test Failure │ Hardcoded school year in │ Verify: Manually check the test failure post-June │ │ │ teachers_controller_test │ 2026. Resolve: Replace hardcoded year with │ │ │ │ Year.current_school_year. │ ├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤ │ Authorization │ IDOR in OrdersController#update │ Verify: Write integration test to exploit. Resolve: │ │ Gap 1 │ │ Add `authorize @order │ └─────────────────┴──────────────────────────────────┴──────────────────────────────────────────────────────┘ Operation aborted Model: nvidia/nemotron-3-nano-30b-a3b i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful failures? We need to answer: "i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful failures?" The user presumably wants advice on how to approach finding meaningful failures in code based on the summary provided. We have read the summary; it's about code review and potential issues. The user is asking for approach suggestions: How to find meaningful failures? Should discuss systematic approach: reading code, looking at test failures, security/policy gaps, correctness, edge cases, performance, data integrity, user flows, etc. Possibly using static analysis, test coverage, identify high-risk areas, use observability, monitor production, triage. May suggest reading documentation, scanning issues, looking at failing tests. Could also suggest running the test suite, exploring failing test details, reproducing failures, looking at related code areas. Might also suggest looking at open issues, using linters, performing code review patterns. Perhaps also mention static typing, type checks, dependency vulnerabilities. And refer to the summary: they identified one test failure (time bomb), two authorization gaps, lower-priority items. So suggestion: focus on test failures, policy gaps, etc. Also mention replicating the failure scenario. Thus answer should be textual: We can propose steps: 1. Re-run test suite; examine failing tests. 2. Identify areas flagged in summary (e.g., IDOR, missing authorize). 3. Use static analysis tools. 4. Simulate user flows that could expose failures. 5. Look at logs and metrics. 6. Prioritize based on impact. And so on. Make answer concise but thorough. Below is a practical, step‑by‑step approach you can use to surface “meaningful” failures in a project like the one described in sources/01‑summarization.md. I’m framing it as a checklist you can follow, rather than a set of specific code‑level fixes—so you can adapt it to whatever part of the stack you’re most interested in (Ruby on Rails, the DB schema, the business rules, the test suite, etc.). ──────────────────────────────────────────────────────────────────────────────── 1. Start with the Known Failures and Flagged Issues ┌─────────────────────────────────┬─────────────────────────────────────────────────────────────────────────┐ │ What to do │ Why it matters │ ├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤ │ Re‑run bundle exec rails test │ The test suite is the single source of truth for what the project │ │ (or the equivalent test │ currently considers broken. In the summary the only failing test is the │ │ command) and capture the exact │ “time‑bomb” in teachers_controller_test. Re‑producing it will tell you │ │ failure output. │ exactly which line/parameter is causing the problem and whether it’s a │ │ │ regression or an inherent design flaw. │ ├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤ │ Create a minimal reproduction │ This isolates the failure from the whole test harness and makes it │ │ script (or a temporary Rails │ easier to explore edge cases without re‑running the whole suite. │ │ console session) that exercises │ │ │ the failing test’s path. │ │ ├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤ │ Cross‑reference the failure │ Often the maintainers have already annotated a ticket with priority, │ │ with the project’s issue │ intended fix, or known work‑arounds. If not, the ticket itself can │ │ tracker (if there’s one). │ become a place to record your findings. │ └─────────────────────────────────┴─────────────────────────────────────────────────────────────────────────┘ ### Quick win - Fix the time‑bomb by replacing the hard‑coded school‑year literal with a dynamic call (Year.current_school_year). Verify that the fix does not break any other test. ──────────────────────────────────────────────────────────────────────────────── 2. Systematically Scan for High‑Impact Security / Authorization Gaps The summary highlighted two concrete IDOR‑style gaps: 1. Order updates without authorize @order 2. Teachers not scoped to their own classrooms in StudentsController / ClassroomEnrollmentsController. How to surface similar gaps elsewhere: ┌─────────────────────────────────────────────────────┬─────────────────────────────────────────────────────┐ │ Step │ Tool / Technique │ ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ │ a. Map all controller actions that modify domain │ grep -R "def .*update|def .*destroy" │ │ objects (e.g., OrdersController#update, │ app/controllers/**/*.rb │ │ StudentsController#create, any │ │ │ *Controller#update/destroy). │ │ ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ │ b. Identify the policy class for each resource │ Look for app/policies/**/*.rb. │ │ (OrderPolicy, StudentPolicy, etc.). │ │ ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ │ c. Check that every state‑changing action calls │ Run a static‑analysis script like rails │ │ authorize (or verify/check) with the correct │ lint:Authorization (if you have a custom linter) or │ │ instance variable. │ simply add a comment placeholder TODO: authorize │ │ │ @order and search for missing ones. │ ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ │ d. Verify that the permitted attributes include the │ `rg "strong_parameters │ │ user_id (or an equivalent scoping column). │ │ ├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤ │ e. Simulate an authenticated user from a different │ Use Rails console or a temporary request spec to │ │ classroom/role and attempt the unsafe action. │ perform patch /orders/42 as a user who does not own │ │ │ the order. │ └─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────┘ ### Pattern to repeat For each public API endpoint or form POST/ PATCH that touches a model, ask: “If I were a different authenticated user, could I cause an unintended state change?” Anything that returns a 200/302 without a proper authorization check is a candidate “meaningful failure”. ──────────────────────────────────────────────────────────────────────────────── 3. Leverage Test Coverage Metrics to Prioritize - Run rails test:coverage (or coverage:install + coverage run) and view the HTML report. - Focus on low‑coverage areas that logically map to risky code paths (e.g., the Admin::PortfolioTransactionsController mentioned in the summary). - Add a single failing test that intentionally violates the expected invariant (e.g., tries to edit a transaction that should be immutable). If it passes, you’ve found a hidden defect. Why? Low coverage often indicates parts of the system that have not been exercised by the existing test suite—exactly the sort of blind spot where subtle bugs hide. ──────────────────────────────────────────────────────────────────────────────── 4. Look for Logical Invariants Violated in Production‑Like Scenarios Even if tests pass, certain business rules can break under edge‑case data. Use the domain model described in the repo to construct “what‑if” tests: ┌──────────────────────────────────┬────────────────────────────────────────────────────────────────────────┐ │ Invariant │ How to test │ ├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤ │ Portfolio ledger is append‑only; │ Create a series of transactions that appear to double‑spend a │ │ balance is derived from │ student’s funds by inserting a pending order, then a second order │ │ transactions │ before the first is executed. Verify the derived balance never exceeds │ │ │ the true ledger sum. │ ├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤ │ OrderExecutionJob runs at the │ Mock the scheduler (or temporarily change the cron schedule) and │ │ correct cadence │ assert that price updates happen only when OrderExecutionJob fires. │ ├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤ │ Grade‑improvement deposit │ Simulate a scenario where a teacher “finalizes” a gradebook without │ │ calculation is applied only │ the expected A/B grade and watch that no deposit is minted. │ │ after finalization │ │ └──────────────────────────────────┴────────────────────────────────────────────────────────────────────────┘ These invariants are often expressed only in design docs (orders-and-transactions.md, gradebook-earnings.md). Treat any deviation as a meaningful failure even if the test suite doesn’t catch it. ──────────────────────────────────────────────────────────────────────────────── 5. Use Static Analysis / Security Scanners for “silent” problems ┌─────────────────────────────────────────────┬─────────────────────────────────────────────────────────────┐ │ Tool │ What it catches that manual inspection might miss │ ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ │ brakeman (Rails security scanner) │ Potential XSS, insecure mass_assignment, missing strong │ │ │ parameters. │ ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ │ rubocop --rails-security │ Unsafe SQL fragments, dangerous File.read usage, etc. │ ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ │ bundler-audit │ Known vulnerable gem versions. │ ├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤ │ rails db:seed:raven or similar migration │ Detect migrations that rely on hard‑coded IDs or │ │ linters │ non‑idempotent logic. │ └─────────────────────────────────────────────┴─────────────────────────────────────────────────────────────┘ Run these as part of CI and treat any warning that touches the same area as a high‑priority failure until it’s deliberately suppressed with a documented justification. ──────────────────────────────────────────────────────────────────────────────── 6. Monetize Impact vs. Likelihood – Prioritize Fixes ┌─────────────────────────────────────┬──────────────────────────┬─────────────────────┬────────────────────┐ │ Failure type │ Typical impact │ Likelihood (given │ Suggested priority │ │ │ │ current code) │ │ ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ │ IDOR on order updates (from │ Moderate – a single │ Low (requires │ Medium – fix now; │ │ summary) │ student could tamper │ authenticated │ it’s a one‑liner. │ │ │ with another's pending │ attacker) │ │ │ │ trade. │ │ │ ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ │ Teacher‑classroom scoping gap │ High – can cross │ Medium – any │ High – patch │ │ │ classroom boundaries, │ teacher can target │ immediately. │ │ │ affect many students. │ any classroom. │ │ ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ │ Missing authorization on │ Low – only reachable by │ Low‑Medium (needs │ Medium – add │ │ PortfolioTransactionsController │ admin URL; but violates │ admin privileges) │ missing authorize │ │ edit/update │ immutable‑ledger │ │ calls and write a │ │ │ invariant. │ │ unit test. │ ├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤ │ Time‑bomb hard‑coded year │ Very high – breaks UI │ Certain (will │ Urgent – replace │ │ │ for all users after a │ trigger │ with dynamic │ │ │ certain date. │ automatically). │ calculation. │ └─────────────────────────────────────┴──────────────────────────┴─────────────────────┴────────────────────┘ When you’re hunting for meaningful failures, start by tackling the failures that move you down this matrix from high impact + high certainty toward low impact + low certainty. ──────────────────────────────────────────────────────────────────────────────── 7. Document the Findings in a Failure‑Report Format 1. Title – concise description (e.g., “Potential IDOR on Order#update”). 2. Reproduction steps – exact request/rails console commands that demonstrate the problem. 3. Code locations – file, line numbers, and policy class. 4. Impact – who is affected, what state can be changed, downstream consequences. 5. Mitigation – minimal code change needed and any required test additions. 6. Follow‑up – deeper investigation needed (e.g., audit all admin controllers for missing authorizations). Having a structured bug report makes it easy to prioritize, delegate, and track progress. ──────────────────────────────────────────────────────────────────────────────── ### TL;DR Checklist 1. Run the full test suite; isolate each failing test. 2. Map every state‑changing controller action to its policy and verify authorize calls. 3. Check coverage; add targeted tests for low‑coverage, high‑risk areas. 4. Create “what‑if” data scenarios that test core invariants (ledger immutability, deposit calculation). 5. Run static‑analysis/security scanners; treat each finding as a potential failure. 6. Prioritize by impact/likelihood and fix the most critical gaps first. 7. Document each discovered failure in a standard report format for transparency. Following this workflow will help you systematically uncover meaningful failures—whether they’re outright bugs, security‑relevant policy violations, or design inconsistencies—while keeping the effort focused on the parts of the codebase that matter most to the project’s core domain (student earnings, portfolio management, and order execution). reading sources/task-instructions.md, how do I use the task-catalog? We need to read sources/task-instructions.md to see instructions on using task-catalog. Probably there is a docs about task-catalog usage. Let's read that file. read sources/task-instructions.md The user asks: "reading sources/task-instructions.md, how do I use the task-catalog?" They want guidance on using the task-catalog based on the instructions they just read. We need to answer from the content of sources/task-instructions.md, explaining how to use the task-catalog. Probablydescribe steps: exploring the catalog, finding tasks, understanding structure, using it to submit tasks, etc. Should refer to sections like "The Task Catalog" etc. Provide concise actionable guidance. Thus respond with explanation, maybe referencing sections like "The Task Catalog" and steps to use it. Should not go beyond policies. Just answer. How to use the Task‑Catalog according to sources/task‑instructions.md 1. Understand what the catalog is - The catalog is a registry of pre‑existing failure scenarios that have already been turned into tasks. - Each entry lives under harbor-tasks// and contains the three core files you will also create for a new task: - instruction.md – the prompt the AI agent receives. - grader‑guidance‑consolidated.md – the grader‑specific guidance. - Reference runs (recorded trials) that prove the failure actually occurs. 2. Browse the catalog - Open the Task Catalog link that the project provides in the Quick Links section. - You’ll see a list of slugs, each representing a separate task. Click a slug to view its folder structure. - Look at the instruction.md and grader‑guidance‑consolidated.md to see how failures were documented and scored. This gives you a concrete template for the kind of failure you need to capture. 3. Decide whether to reuse or duplicate - Check for duplicates. Before you start a new task, verify that an existing catalog entry does not already cover the failure you have in mind. Duplicate submissions are rejected at review. - If you find a close match but need a different angle (e.g., a different severity, a different code path), you can still add a new slug; just make sure the prompt and ground‑truth details are distinct. 4. Create your own task folder - Follow the scaffold shown in the catalog: ``` harbor-tasks// ├─ instruction.md # engineering prompt ├─ task.toml # metadata (repo commit, browser flag, etc.) ├─ tests/ │ └─ grader‑guidance‑consolidated.md └─ reference‑runs/ # recorded trials that demonstrate the failure ``` - The Workspace & workspace.patch section of the instructions tells you how to generate a patch that captures any workspace changes you made while exploring. 5. Populate the files - instruction.md – Write a realistic engineering request that would naturally expose the failure you discovered. - Keep hints out; the prompt must not point directly at the solution. - Ensure the prompt is self‑contained (all needed context must be inside the patched workspace). - grader‑guidance‑consolidated.md – Translate the failure into the eight‑criterion grader language: - Provide Task context and Business context (if any). - Define Ground truth (exact file/line, correct fix, test that should catch it). - For each of the eight criteria, describe what a strong vs. weak response looks like on this task. - Add any heavy penalties only for deal‑breaker behaviours, naming the criterion they affect. - Reference runs – Run harbor-run (or codex/claude depending on the agent you chose) to generate recorded trials that trigger the failure. Copy those runs into reference‑runs/ using npx tsx scripts/copy-reference-run.ts …. 6. Run detectors before submitting - The catalog workflow expects you to invoke every detector skill (e.g., /detector‑meaningful-failure, /detector‑fact‑check‑rubric‑claims, etc.). - Detectors will flag stale inputs, broken environment checks, cross‑task references, or missing heavy‑penalty definitions. Fix any issues they surface, then re‑run the detectors. 7. Validate and package - Execute npx tsx scripts/submit‑task.ts . - The script will: 1. Verify that all required files exist. 2. Ensure the patch is clean (no unintended binaries, ignored files, or leaked snapshots). 3. Confirm you have at least four accepted reference runs. 4. Produce a single tarball that you upload to the platform. - If any warnings appear, address them before you click Submit on the platform UI. 8. Submit with the required metadata - When you upload the tarball, also fill the Import/Export panel: - Export a JSON save‑state of the task (so you can rebuild later). - Paste your Slack thread URL (the place where you asked/answered questions). - Indicate whether this is a complete submission or a work‑in‑progress draft. - If you suspect the grader scored a run incorrectly, you can flag it with the grader‑performance checkbox; this creates a separate concern thread for the review team. 9. Iterate via the Feedback Loop - After submission, reviewers will run the same detector suite on your tarball. - Any issues they raise (e.g., stale inputs, missing heavy‑penalty definitions) will be sent back as feedback. - Update your task files, re‑run the detectors, and submit a new version if needed. ──────────────────────────────────────────────────────────────────────────────── ### Quick‑Reference Cheat‑Sheet ┌────────────────┬────────────────────────────────────────────┬─────────────────────────────────────────────┐ │ Step │ Command / Action │ Result │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Explore │ Open the Task Catalog link → browse slugs │ See existing instruction.md & │ │ │ │ grader‑guidance‑consolidated.md examples │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Create folder │ mkdir -p harbor-tasks// │ Scaffold for new task │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Generate patch │ Work in Explore → │ environment/workspace.patch captured │ │ │ /create-snapshot:snapshot → │ │ │ │ snapshot-to-task.ts │ │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Write prompt │ instruction.md → realistic, no hints │ Agent receives clear engineering request │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Write grader │ tests/grader‑guidance‑consolidated.md → 8 │ Grader knows exactly what to score │ │ guidance │ criteria + heavy penalties │ │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Run trials │ harbor-run (or codex/claude) → copy runs │ reference-runs/ populated │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Run detectors │ /detector‑* skills │ All automated checks pass │ ├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤ │ Submit │ npx tsx scripts/submit‑task.ts → │ Task packaged & ready for review │ │ │ upload tarball │ │ └────────────────┴────────────────────────────────────────────┴─────────────────────────────────────────────┘ Follow the flow Explore → Build → Validate → Submit exactly as the Task‑Catalog section of sources/task‑instructions.md describes, and you’ll be able to add a new, meaningful failure to the catalog without duplicating existing work. ─────────────────────────────────────────────────────────────────────────────────────────────────────────────── ─────────────────────────────────────────────────────────────────────────────────────────────────────────────── ~/workspaces/dataannotation/current-project (raccoon-stocks) ↑30k ↓4.5k R4.2k CH23.3% 15.1%/131k (auto) (nvidia) nvidia/nemotron-3-nano-30b-a3b • medium mode: implementation