538 lines
49 KiB
Markdown
538 lines
49 KiB
Markdown
|
||
pi v0.84.2
|
||
escape interrupt · ctrl+c/ctrl+d clear/exit · / commands · ! bash · ctrl+o more
|
||
Press ctrl+o to show full startup help and loaded resources.
|
||
|
||
Pi can explain its own features and look up its docs. Ask it how to use or extend Pi.
|
||
|
||
[Extensions]
|
||
@ollama/pi-web-search, mode.ts
|
||
|
||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||
What's New
|
||
|
||
[0.84.2] - 2026-08-14
|
||
|
||
### New Features
|
||
|
||
- Fullscreen transcript search — Search and navigate matches in fullscreen mode. See TUI Fullscreen Viewport
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen
|
||
-viewport).
|
||
- Configurable default tools — Choose startup built-in tools globally or per project. See Tools
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/settings.md#tools).
|
||
- Configurable fullscreen exit output — Print the transcript or only a resume hint on exit. See Interactive
|
||
Mode
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/usage.md#interactive-mode).
|
||
|
||
### Added
|
||
|
||
- Added fullscreen transcript search with Ctrl+Shift+F, incremental match highlighting, configurable search
|
||
match theme colors, and next/previous navigation with Enter/Ctrl+G and Shift+Enter/Ctrl+Shift+G.
|
||
- Added experimental strict JSON-schema constrained sampling for the default read, bash, edit, and write
|
||
tools under PI_EXPERIMENTAL=1.
|
||
- Added a fullscreen exit output setting to choose between printing the final transcript and only a session
|
||
resume hint.
|
||
- Added the defaultTools setting for configuring the initial built-in tool selection globally or per project.
|
||
- Added --use-theme <name[/name]> to choose an initial per-run interactive theme without changing saved
|
||
settings (#7722 (https://github.com/earendil-works/pi/pull/7722) by @rwachtler
|
||
(https://github.com/rwachtler)).
|
||
- Added expandPromptTemplates to extension pi.sendUserMessage() options for explicitly dispatching commands
|
||
and expanding skills and prompt templates. See pi.sendUserMessage()
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/extensions.md#pisendusermessa
|
||
gecontent-options) (#7857 (https://github.com/earendil-works/pi/pull/7857) by @mrexodia
|
||
(https://github.com/mrexodia)).
|
||
- Added inherited createGatewayBindingFetch() for routing Cloudflare AI Gateway requests through a Workers AI
|
||
binding without an API token (#7901 (https://github.com/earendil-works/pi/pull/7901) by @Maximo-Guk
|
||
(https://github.com/Maximo-Guk)).
|
||
- Added inherited AssistantMessage.endTurn to preserve OpenAI Codex's terminal end_turn signal for
|
||
diagnostics (#7766 (https://github.com/earendil-works/pi/pull/7766)).
|
||
- Added inherited unbound single-line transcript scrolling actions for fullscreen mode. See TUI Fullscreen
|
||
Viewport
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen
|
||
-viewport) (#7903 (https://github.com/earendil-works/pi/pull/7903) by @midastruth
|
||
(https://github.com/midastruth)).
|
||
|
||
### Changed
|
||
|
||
- Changed inherited Kimi Coding requests to use pi's runtime User-Agent header.
|
||
- Replaced the inherited Mistral SDK transport with a native Chat Completions HTTP stream, eliminating its
|
||
generated client and schema runtime overhead.
|
||
- Documented the generic AI_AGENT=pi process marker and how it differs from PI_CODING_AGENT=true (#7747
|
||
(https://github.com/earendil-works/pi/issues/7747)).
|
||
- Changed inherited OpenAI Responses deferred tool loading to prefer message-anchored additional_tools where
|
||
supported while retaining tool-search and top-level fallbacks (#7709
|
||
(https://github.com/earendil-works/pi/issues/7709)).
|
||
- Reduced inherited fullscreen rendering allocation churn by painting full-width layout rows directly instead
|
||
of recompositing them on every frame.
|
||
|
||
### Fixed
|
||
|
||
- Fixed managed-tool downloads delaying TUI startup and hiding diagnostics in fullscreen mode by mounting the
|
||
TUI first and showing download progress and warnings inside it.
|
||
- Fixed opening a model selector immediately after startup cancelling and restarting the in-progress model
|
||
catalog refresh.
|
||
- Fixed inherited GitHub Copilot login triggering API rate limits while enabling model policies by limiting
|
||
concurrent policy updates (#6187 (https://github.com/earendil-works/pi/issues/6187)).
|
||
- Fixed fullscreen transcript search snapping back to the current match during manual scrolling and
|
||
fragmented mouse input leaking into the search query.
|
||
- Fixed inherited required LaTeX arguments starting on a new line being parsed as empty (#7760
|
||
(https://github.com/earendil-works/pi/issues/7760)).
|
||
- Updated the transitive nanoid development dependency to address a denial-of-service vulnerability.
|
||
- Fixed fallback rendering for extension tool results to collapse long output and honor tool expansion (#7979
|
||
(https://github.com/earendil-works/pi/issues/7979)).
|
||
- Fixed JSON and RPC message_update events dropping cumulative usage during streaming. See JSON Event Mode
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/json.md) and RPC
|
||
message_update
|
||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/rpc.md#message_update-streami
|
||
ng) (#7982 (https://github.com/earendil-works/pi/pull/7982) by @christianklotz
|
||
(https://github.com/christianklotz)).
|
||
- Fixed pi.sendMessage(..., { triggerTurn: false }) steering an active run instead of only recording the
|
||
custom message (#8022 (https://github.com/earendil-works/pi/pull/8022) by @cristinaponcela
|
||
(https://github.com/cristinaponcela)).
|
||
- Fixed the defaultTools setting dropping extension and SDK custom tools when selecting built-in defaults.
|
||
- Fixed the subagent example rejecting YAML array syntax for the tools frontmatter field (#7598
|
||
(https://github.com/earendil-works/pi/pull/7598) by @alexsavio (https://github.com/alexsavio)).
|
||
- Fixed the subagent example dropping parent session model, thinking, and tool configuration (#7897
|
||
(https://github.com/earendil-works/pi/pull/7897) by @virtuald (https://github.com/virtuald)).
|
||
- Fixed custom system prompts concatenating the current working directory with later appended prompt content
|
||
(#7887 (https://github.com/earendil-works/pi/pull/7887) by @distributedlock
|
||
(https://github.com/distributedlock)).
|
||
- Fixed inherited OpenAI Responses function and custom tool calls losing namespaces during streaming,
|
||
proxying, and replay (#7709 (https://github.com/earendil-works/pi/issues/7709)).
|
||
- Fixed inherited upstream request buffer failures not triggering automatic assistant retries.
|
||
- Fixed inherited built-in and custom DeepSeek API models sending output limits through an unsupported field.
|
||
- Fixed inherited Amazon Bedrock replay rejecting tool arguments that contain empty object keys while
|
||
preserving all valid nested values (#7882 (https://github.com/earendil-works/pi/pull/7882) by @muyiyr
|
||
(https://github.com/muyiyr)).
|
||
- Fixed inherited DeepSeek compatibility detection for base URLs whose hostname contains uppercase letters
|
||
(#7933 (https://github.com/earendil-works/pi/pull/7933) by @yearth (https://github.com/yearth)).
|
||
- Fixed inherited Google Generative AI and Vertex AI responses with tool calls incorrectly treating
|
||
output-limit or provider-error stops as normal tool use (#8059
|
||
(https://github.com/earendil-works/pi/issues/8059)).
|
||
- Fixed inherited fullscreen mouse drag selection and OSC 8 link activation in terminals that report generic
|
||
SGR mouse release button codes (#7963 (https://github.com/earendil-works/pi/issues/7963)).
|
||
- Fixed inherited focused fullscreen overlays not receiving mouse wheel or viewport scroll keys such as
|
||
PageUp and PageDown (#7894 (https://github.com/earendil-works/pi/issues/7894)).
|
||
- Fixed inherited LaTeX control spaces split across line endings causing complete expressions to fall back to
|
||
raw source.
|
||
- Fixed split Alt+Enter input over SSH being misread as Escape, added PI_TUI_ESC_TIMEOUT for high-latency
|
||
terminals, and limited that timeout to lone Escape input (#7899
|
||
(https://github.com/earendil-works/pi/pull/7899) by @powerfooI (https://github.com/powerfooI)).
|
||
- Fixed inherited idle fullscreen sessions repainting and clearing text selection when the terminal loses
|
||
focus (#7892 (https://github.com/earendil-works/pi/pull/7892) by @terrorobe
|
||
(https://github.com/terrorobe)).
|
||
- Fixed fullscreen selection copy to use the host clipboard and report failure instead of claiming success
|
||
when OSC 52 is unsupported (#8110 (https://github.com/earendil-works/pi/pull/8110) by @Panoplos
|
||
(https://github.com/Panoplos)).
|
||
|
||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||
|
||
Model: nvidia/llama-3.3-nemotron-super-49b-v1
|
||
|
||
|
||
i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in
|
||
sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful
|
||
failures?
|
||
|
||
|
||
|
||
read ...
|
||
|
||
Validation failed for tool "read":
|
||
- path: must have required properties path
|
||
|
||
Received arguments:
|
||
{}
|
||
|
||
|
||
|
||
read sources/01-summarization.md
|
||
|
||
|
||
Based on the provided summaries and findings in sources/ai-version-instructions.md and
|
||
sources/01-summarization.md (assuming the latter's content is reflected in the large text block you provided,
|
||
as the initial read tool call failed due to missing path specification), here's a suggested approach to
|
||
finding meaningful failures in the project, categorized for clarity:
|
||
|
||
### 1. Already Identified Failures & Vulnerabilities (from sources/01-summarization.md summary)
|
||
|
||
┌─────────────────┬──────────────────────────────────┬──────────────────────────────────────────────────────┐
|
||
│ Category │ Issue │ Approach to Verify/Resolve │
|
||
├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤
|
||
│ Test Failure │ Hardcoded school year in │ Verify: Manually check the test failure post-June │
|
||
│ │ teachers_controller_test │ 2026. Resolve: Replace hardcoded year with │
|
||
│ │ │ Year.current_school_year. │
|
||
├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤
|
||
│ Authorization │ IDOR in OrdersController#update │ Verify: Write integration test to exploit. Resolve: │
|
||
│ Gap 1 │ │ Add `authorize @order │
|
||
└─────────────────┴──────────────────────────────────┴──────────────────────────────────────────────────────┘
|
||
|
||
Operation aborted
|
||
|
||
Model: nvidia/nemotron-3-nano-30b-a3b
|
||
|
||
|
||
i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in
|
||
sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful
|
||
failures?
|
||
|
||
|
||
We need to answer: "i have a project outlined in sources/ai-version-instructions.md. The repo for the project
|
||
is summarized in sources/01-summarization.md. given the instructions, how would you suggest I approach
|
||
finding meaningful failures?" The user presumably wants advice on how to approach finding meaningful failures
|
||
in code based on the summary provided. We have read the summary; it's about code review and potential issues.
|
||
The user is asking for approach suggestions: How to find meaningful failures? Should discuss systematic
|
||
approach: reading code, looking at test failures, security/policy gaps, correctness, edge cases, performance,
|
||
data integrity, user flows, etc. Possibly using static analysis, test coverage, identify high-risk areas, use
|
||
observability, monitor production, triage. May suggest reading documentation, scanning issues, looking at
|
||
failing tests. Could also suggest running the test suite, exploring failing test details, reproducing
|
||
failures, looking at related code areas. Might also suggest looking at open issues, using linters, performing
|
||
code review patterns. Perhaps also mention static typing, type checks, dependency vulnerabilities. And refer
|
||
to the summary: they identified one test failure (time bomb), two authorization gaps, lower-priority items.
|
||
So suggestion: focus on test failures, policy gaps, etc. Also mention replicating the failure scenario.
|
||
|
||
Thus answer should be textual: We can propose steps: 1. Re-run test suite; examine failing tests. 2. Identify
|
||
areas flagged in summary (e.g., IDOR, missing authorize). 3. Use static analysis tools. 4. Simulate user
|
||
flows that could expose failures. 5. Look at logs and metrics. 6. Prioritize based on impact. And so on.
|
||
|
||
Make answer concise but thorough.
|
||
|
||
Below is a practical, step‑by‑step approach you can use to surface “meaningful” failures in a project like
|
||
the one described in sources/01‑summarization.md. I’m framing it as a checklist you can follow, rather than
|
||
a set of specific code‑level fixes—so you can adapt it to whatever part of the stack you’re most interested
|
||
in (Ruby on Rails, the DB schema, the business rules, the test suite, etc.).
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
1. Start with the Known Failures and Flagged Issues
|
||
|
||
┌─────────────────────────────────┬─────────────────────────────────────────────────────────────────────────┐
|
||
│ What to do │ Why it matters │
|
||
├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤
|
||
│ Re‑run bundle exec rails test │ The test suite is the single source of truth for what the project │
|
||
│ (or the equivalent test │ currently considers broken. In the summary the only failing test is the │
|
||
│ command) and capture the exact │ “time‑bomb” in teachers_controller_test. Re‑producing it will tell you │
|
||
│ failure output. │ exactly which line/parameter is causing the problem and whether it’s a │
|
||
│ │ regression or an inherent design flaw. │
|
||
├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤
|
||
│ Create a minimal reproduction │ This isolates the failure from the whole test harness and makes it │
|
||
│ script (or a temporary Rails │ easier to explore edge cases without re‑running the whole suite. │
|
||
│ console session) that exercises │ │
|
||
│ the failing test’s path. │ │
|
||
├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤
|
||
│ Cross‑reference the failure │ Often the maintainers have already annotated a ticket with priority, │
|
||
│ with the project’s issue │ intended fix, or known work‑arounds. If not, the ticket itself can │
|
||
│ tracker (if there’s one). │ become a place to record your findings. │
|
||
└─────────────────────────────────┴─────────────────────────────────────────────────────────────────────────┘
|
||
|
||
### Quick win
|
||
|
||
- Fix the time‑bomb by replacing the hard‑coded school‑year literal with a dynamic call
|
||
(Year.current_school_year). Verify that the fix does not break any other test.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
2. Systematically Scan for High‑Impact Security / Authorization Gaps
|
||
|
||
The summary highlighted two concrete IDOR‑style gaps:
|
||
|
||
1. Order updates without authorize @order
|
||
2. Teachers not scoped to their own classrooms in StudentsController / ClassroomEnrollmentsController.
|
||
|
||
How to surface similar gaps elsewhere:
|
||
|
||
┌─────────────────────────────────────────────────────┬─────────────────────────────────────────────────────┐
|
||
│ Step │ Tool / Technique │
|
||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||
│ a. Map all controller actions that modify domain │ grep -R "def .*update|def .*destroy" │
|
||
│ objects (e.g., OrdersController#update, │ app/controllers/**/*.rb │
|
||
│ StudentsController#create, any │ │
|
||
│ *Controller#update/destroy). │ │
|
||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||
│ b. Identify the policy class for each resource │ Look for app/policies/**/*.rb. │
|
||
│ (OrderPolicy, StudentPolicy, etc.). │ │
|
||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||
│ c. Check that every state‑changing action calls │ Run a static‑analysis script like rails │
|
||
│ authorize (or verify/check) with the correct │ lint:Authorization (if you have a custom linter) or │
|
||
│ instance variable. │ simply add a comment placeholder TODO: authorize │
|
||
│ │ @order and search for missing ones. │
|
||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||
│ d. Verify that the permitted attributes include the │ `rg "strong_parameters │
|
||
│ user_id (or an equivalent scoping column). │ │
|
||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||
│ e. Simulate an authenticated user from a different │ Use Rails console or a temporary request spec to │
|
||
│ classroom/role and attempt the unsafe action. │ perform patch /orders/42 as a user who does not own │
|
||
│ │ the order. │
|
||
└─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────┘
|
||
|
||
### Pattern to repeat
|
||
|
||
For each public API endpoint or form POST/ PATCH that touches a model, ask: “If I were a different
|
||
authenticated user, could I cause an unintended state change?” Anything that returns a 200/302 without a
|
||
proper authorization check is a candidate “meaningful failure”.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
3. Leverage Test Coverage Metrics to Prioritize
|
||
|
||
- Run rails test:coverage (or coverage:install + coverage run) and view the HTML report.
|
||
- Focus on low‑coverage areas that logically map to risky code paths (e.g., the
|
||
Admin::PortfolioTransactionsController mentioned in the summary).
|
||
- Add a single failing test that intentionally violates the expected invariant (e.g., tries to edit a
|
||
transaction that should be immutable). If it passes, you’ve found a hidden defect.
|
||
|
||
Why? Low coverage often indicates parts of the system that have not been exercised by the existing test
|
||
suite—exactly the sort of blind spot where subtle bugs hide.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
4. Look for Logical Invariants Violated in Production‑Like Scenarios
|
||
|
||
Even if tests pass, certain business rules can break under edge‑case data. Use the domain model described in
|
||
the repo to construct “what‑if” tests:
|
||
|
||
┌──────────────────────────────────┬────────────────────────────────────────────────────────────────────────┐
|
||
│ Invariant │ How to test │
|
||
├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤
|
||
│ Portfolio ledger is append‑only; │ Create a series of transactions that appear to double‑spend a │
|
||
│ balance is derived from │ student’s funds by inserting a pending order, then a second order │
|
||
│ transactions │ before the first is executed. Verify the derived balance never exceeds │
|
||
│ │ the true ledger sum. │
|
||
├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤
|
||
│ OrderExecutionJob runs at the │ Mock the scheduler (or temporarily change the cron schedule) and │
|
||
│ correct cadence │ assert that price updates happen only when OrderExecutionJob fires. │
|
||
├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤
|
||
│ Grade‑improvement deposit │ Simulate a scenario where a teacher “finalizes” a gradebook without │
|
||
│ calculation is applied only │ the expected A/B grade and watch that no deposit is minted. │
|
||
│ after finalization │ │
|
||
└──────────────────────────────────┴────────────────────────────────────────────────────────────────────────┘
|
||
|
||
These invariants are often expressed only in design docs (orders-and-transactions.md, gradebook-earnings.md).
|
||
Treat any deviation as a meaningful failure even if the test suite doesn’t catch it.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
5. Use Static Analysis / Security Scanners for “silent” problems
|
||
|
||
┌─────────────────────────────────────────────┬─────────────────────────────────────────────────────────────┐
|
||
│ Tool │ What it catches that manual inspection might miss │
|
||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||
│ brakeman (Rails security scanner) │ Potential XSS, insecure mass_assignment, missing strong │
|
||
│ │ parameters. │
|
||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||
│ rubocop --rails-security │ Unsafe SQL fragments, dangerous File.read usage, etc. │
|
||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||
│ bundler-audit │ Known vulnerable gem versions. │
|
||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||
│ rails db:seed:raven or similar migration │ Detect migrations that rely on hard‑coded IDs or │
|
||
│ linters │ non‑idempotent logic. │
|
||
└─────────────────────────────────────────────┴─────────────────────────────────────────────────────────────┘
|
||
|
||
Run these as part of CI and treat any warning that touches the same area as a high‑priority failure until
|
||
it’s deliberately suppressed with a documented justification.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
6. Monetize Impact vs. Likelihood – Prioritize Fixes
|
||
|
||
┌─────────────────────────────────────┬──────────────────────────┬─────────────────────┬────────────────────┐
|
||
│ Failure type │ Typical impact │ Likelihood (given │ Suggested priority │
|
||
│ │ │ current code) │ │
|
||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||
│ IDOR on order updates (from │ Moderate – a single │ Low (requires │ Medium – fix now; │
|
||
│ summary) │ student could tamper │ authenticated │ it’s a one‑liner. │
|
||
│ │ with another's pending │ attacker) │ │
|
||
│ │ trade. │ │ │
|
||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||
│ Teacher‑classroom scoping gap │ High – can cross │ Medium – any │ High – patch │
|
||
│ │ classroom boundaries, │ teacher can target │ immediately. │
|
||
│ │ affect many students. │ any classroom. │ │
|
||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||
│ Missing authorization on │ Low – only reachable by │ Low‑Medium (needs │ Medium – add │
|
||
│ PortfolioTransactionsController │ admin URL; but violates │ admin privileges) │ missing authorize │
|
||
│ edit/update │ immutable‑ledger │ │ calls and write a │
|
||
│ │ invariant. │ │ unit test. │
|
||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||
│ Time‑bomb hard‑coded year │ Very high – breaks UI │ Certain (will │ Urgent – replace │
|
||
│ │ for all users after a │ trigger │ with dynamic │
|
||
│ │ certain date. │ automatically). │ calculation. │
|
||
└─────────────────────────────────────┴──────────────────────────┴─────────────────────┴────────────────────┘
|
||
|
||
When you’re hunting for meaningful failures, start by tackling the failures that move you down this matrix
|
||
from high impact + high certainty toward low impact + low certainty.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
7. Document the Findings in a Failure‑Report Format
|
||
|
||
1. Title – concise description (e.g., “Potential IDOR on Order#update”).
|
||
2. Reproduction steps – exact request/rails console commands that demonstrate the problem.
|
||
3. Code locations – file, line numbers, and policy class.
|
||
4. Impact – who is affected, what state can be changed, downstream consequences.
|
||
5. Mitigation – minimal code change needed and any required test additions.
|
||
6. Follow‑up – deeper investigation needed (e.g., audit all admin controllers for missing authorizations).
|
||
|
||
Having a structured bug report makes it easy to prioritize, delegate, and track progress.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
### TL;DR Checklist
|
||
|
||
1. Run the full test suite; isolate each failing test.
|
||
2. Map every state‑changing controller action to its policy and verify authorize calls.
|
||
3. Check coverage; add targeted tests for low‑coverage, high‑risk areas.
|
||
4. Create “what‑if” data scenarios that test core invariants (ledger immutability, deposit calculation).
|
||
5. Run static‑analysis/security scanners; treat each finding as a potential failure.
|
||
6. Prioritize by impact/likelihood and fix the most critical gaps first.
|
||
7. Document each discovered failure in a standard report format for transparency.
|
||
|
||
Following this workflow will help you systematically uncover meaningful failures—whether they’re outright
|
||
bugs, security‑relevant policy violations, or design inconsistencies—while keeping the effort focused on the
|
||
parts of the codebase that matter most to the project’s core domain (student earnings, portfolio management,
|
||
and order execution).
|
||
|
||
|
||
reading sources/task-instructions.md, how do I use the task-catalog?
|
||
|
||
|
||
We need to read sources/task-instructions.md to see instructions on using task-catalog. Probably there is a
|
||
docs about task-catalog usage. Let's read that file.
|
||
|
||
|
||
read sources/task-instructions.md
|
||
|
||
|
||
The user asks: "reading sources/task-instructions.md, how do I use the task-catalog?" They want guidance on
|
||
using the task-catalog based on the instructions they just read.
|
||
|
||
We need to answer from the content of sources/task-instructions.md, explaining how to use the task-catalog.
|
||
Probablydescribe steps: exploring the catalog, finding tasks, understanding structure, using it to submit
|
||
tasks, etc. Should refer to sections like "The Task Catalog" etc. Provide concise actionable guidance.
|
||
|
||
Thus respond with explanation, maybe referencing sections like "The Task Catalog" and steps to use it. Should
|
||
not go beyond policies. Just answer.
|
||
|
||
How to use the Task‑Catalog according to sources/task‑instructions.md
|
||
|
||
1. Understand what the catalog is
|
||
- The catalog is a registry of pre‑existing failure scenarios that have already been turned into tasks.
|
||
- Each entry lives under harbor-tasks/<slug>/ and contains the three core files you will also create for
|
||
a new task:
|
||
- instruction.md – the prompt the AI agent receives.
|
||
- grader‑guidance‑consolidated.md – the grader‑specific guidance.
|
||
- Reference runs (recorded trials) that prove the failure actually occurs.
|
||
|
||
|
||
2. Browse the catalog
|
||
- Open the Task Catalog link that the project provides in the Quick Links section.
|
||
- You’ll see a list of slugs, each representing a separate task. Click a slug to view its folder
|
||
structure.
|
||
- Look at the instruction.md and grader‑guidance‑consolidated.md to see how failures were documented and
|
||
scored. This gives you a concrete template for the kind of failure you need to capture.
|
||
|
||
3. Decide whether to reuse or duplicate
|
||
- Check for duplicates. Before you start a new task, verify that an existing catalog entry does not
|
||
already cover the failure you have in mind. Duplicate submissions are rejected at review.
|
||
- If you find a close match but need a different angle (e.g., a different severity, a different code
|
||
path), you can still add a new slug; just make sure the prompt and ground‑truth details are distinct.
|
||
|
||
4. Create your own task folder
|
||
- Follow the scaffold shown in the catalog:
|
||
```
|
||
harbor-tasks/<your‑slug>/
|
||
├─ instruction.md # engineering prompt
|
||
├─ task.toml # metadata (repo commit, browser flag, etc.)
|
||
├─ tests/
|
||
│ └─ grader‑guidance‑consolidated.md
|
||
└─ reference‑runs/ # recorded trials that demonstrate the failure
|
||
```
|
||
- The Workspace & workspace.patch section of the instructions tells you how to generate a patch that
|
||
captures any workspace changes you made while exploring.
|
||
|
||
5. Populate the files
|
||
- instruction.md – Write a realistic engineering request that would naturally expose the failure you
|
||
discovered.
|
||
- Keep hints out; the prompt must not point directly at the solution.
|
||
- Ensure the prompt is self‑contained (all needed context must be inside the patched workspace).
|
||
|
||
- grader‑guidance‑consolidated.md – Translate the failure into the eight‑criterion grader language:
|
||
- Provide Task context and Business context (if any).
|
||
- Define Ground truth (exact file/line, correct fix, test that should catch it).
|
||
- For each of the eight criteria, describe what a strong vs. weak response looks like on this task.
|
||
- Add any heavy penalties only for deal‑breaker behaviours, naming the criterion they affect.
|
||
|
||
- Reference runs – Run harbor-run (or codex/claude depending on the agent you chose) to generate recorded
|
||
trials that trigger the failure. Copy those runs into reference‑runs/ using npx tsx
|
||
scripts/copy-reference-run.ts ….
|
||
|
||
6. Run detectors before submitting
|
||
- The catalog workflow expects you to invoke every detector skill (e.g., /detector‑meaningful-failure,
|
||
/detector‑fact‑check‑rubric‑claims, etc.).
|
||
- Detectors will flag stale inputs, broken environment checks, cross‑task references, or missing
|
||
heavy‑penalty definitions. Fix any issues they surface, then re‑run the detectors.
|
||
|
||
7. Validate and package
|
||
- Execute npx tsx scripts/submit‑task.ts <your‑slug> .
|
||
- The script will:
|
||
1. Verify that all required files exist.
|
||
2. Ensure the patch is clean (no unintended binaries, ignored files, or leaked snapshots).
|
||
3. Confirm you have at least four accepted reference runs.
|
||
4. Produce a single tarball that you upload to the platform.
|
||
|
||
- If any warnings appear, address them before you click Submit on the platform UI.
|
||
|
||
8. Submit with the required metadata
|
||
- When you upload the tarball, also fill the Import/Export panel:
|
||
- Export a JSON save‑state of the task (so you can rebuild later).
|
||
- Paste your Slack thread URL (the place where you asked/answered questions).
|
||
- Indicate whether this is a complete submission or a work‑in‑progress draft.
|
||
|
||
- If you suspect the grader scored a run incorrectly, you can flag it with the grader‑performance
|
||
checkbox; this creates a separate concern thread for the review team.
|
||
|
||
9. Iterate via the Feedback Loop
|
||
- After submission, reviewers will run the same detector suite on your tarball.
|
||
- Any issues they raise (e.g., stale inputs, missing heavy‑penalty definitions) will be sent back as
|
||
feedback.
|
||
- Update your task files, re‑run the detectors, and submit a new version if needed.
|
||
|
||
────────────────────────────────────────────────────────────────────────────────
|
||
|
||
### Quick‑Reference Cheat‑Sheet
|
||
|
||
┌────────────────┬────────────────────────────────────────────┬─────────────────────────────────────────────┐
|
||
│ Step │ Command / Action │ Result │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Explore │ Open the Task Catalog link → browse slugs │ See existing instruction.md & │
|
||
│ │ │ grader‑guidance‑consolidated.md examples │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Create folder │ mkdir -p harbor-tasks/<slug>/ │ Scaffold for new task │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Generate patch │ Work in Explore → │ environment/workspace.patch captured │
|
||
│ │ /create-snapshot:snapshot → │ │
|
||
│ │ snapshot-to-task.ts │ │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Write prompt │ instruction.md → realistic, no hints │ Agent receives clear engineering request │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Write grader │ tests/grader‑guidance‑consolidated.md → 8 │ Grader knows exactly what to score │
|
||
│ guidance │ criteria + heavy penalties │ │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Run trials │ harbor-run (or codex/claude) → copy runs │ reference-runs/ populated │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Run detectors │ /detector‑* skills │ All automated checks pass │
|
||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||
│ Submit │ npx tsx scripts/submit‑task.ts <slug> → │ Task packaged & ready for review │
|
||
│ │ upload tarball │ │
|
||
└────────────────┴────────────────────────────────────────────┴─────────────────────────────────────────────┘
|
||
|
||
Follow the flow Explore → Build → Validate → Submit exactly as the Task‑Catalog section of
|
||
sources/task‑instructions.md describes, and you’ll be able to add a new, meaningful failure to the catalog
|
||
without duplicating existing work.
|
||
|
||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||
|
||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||
~/workspaces/dataannotation/current-project (raccoon-stocks)
|
||
↑30k ↓4.5k R4.2k CH23.3% 15.1%/131k (auto) (nvidia) nvidia/nemotron-3-nano-30b-a3b • medium
|
||
mode: implementation
|