chore: add icm CONTEXT file, tmux convo
This commit is contained in:
537
sources/tmux-convo-1.md
Normal file
537
sources/tmux-convo-1.md
Normal file
@@ -0,0 +1,537 @@
|
||||
|
||||
pi v0.84.2
|
||||
escape interrupt · ctrl+c/ctrl+d clear/exit · / commands · ! bash · ctrl+o more
|
||||
Press ctrl+o to show full startup help and loaded resources.
|
||||
|
||||
Pi can explain its own features and look up its docs. Ask it how to use or extend Pi.
|
||||
|
||||
[Extensions]
|
||||
@ollama/pi-web-search, mode.ts
|
||||
|
||||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||||
What's New
|
||||
|
||||
[0.84.2] - 2026-08-14
|
||||
|
||||
### New Features
|
||||
|
||||
- Fullscreen transcript search — Search and navigate matches in fullscreen mode. See TUI Fullscreen Viewport
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen
|
||||
-viewport).
|
||||
- Configurable default tools — Choose startup built-in tools globally or per project. See Tools
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/settings.md#tools).
|
||||
- Configurable fullscreen exit output — Print the transcript or only a resume hint on exit. See Interactive
|
||||
Mode
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/usage.md#interactive-mode).
|
||||
|
||||
### Added
|
||||
|
||||
- Added fullscreen transcript search with Ctrl+Shift+F, incremental match highlighting, configurable search
|
||||
match theme colors, and next/previous navigation with Enter/Ctrl+G and Shift+Enter/Ctrl+Shift+G.
|
||||
- Added experimental strict JSON-schema constrained sampling for the default read, bash, edit, and write
|
||||
tools under PI_EXPERIMENTAL=1.
|
||||
- Added a fullscreen exit output setting to choose between printing the final transcript and only a session
|
||||
resume hint.
|
||||
- Added the defaultTools setting for configuring the initial built-in tool selection globally or per project.
|
||||
- Added --use-theme <name[/name]> to choose an initial per-run interactive theme without changing saved
|
||||
settings (#7722 (https://github.com/earendil-works/pi/pull/7722) by @rwachtler
|
||||
(https://github.com/rwachtler)).
|
||||
- Added expandPromptTemplates to extension pi.sendUserMessage() options for explicitly dispatching commands
|
||||
and expanding skills and prompt templates. See pi.sendUserMessage()
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/extensions.md#pisendusermessa
|
||||
gecontent-options) (#7857 (https://github.com/earendil-works/pi/pull/7857) by @mrexodia
|
||||
(https://github.com/mrexodia)).
|
||||
- Added inherited createGatewayBindingFetch() for routing Cloudflare AI Gateway requests through a Workers AI
|
||||
binding without an API token (#7901 (https://github.com/earendil-works/pi/pull/7901) by @Maximo-Guk
|
||||
(https://github.com/Maximo-Guk)).
|
||||
- Added inherited AssistantMessage.endTurn to preserve OpenAI Codex's terminal end_turn signal for
|
||||
diagnostics (#7766 (https://github.com/earendil-works/pi/pull/7766)).
|
||||
- Added inherited unbound single-line transcript scrolling actions for fullscreen mode. See TUI Fullscreen
|
||||
Viewport
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/keybindings.md#tui-fullscreen
|
||||
-viewport) (#7903 (https://github.com/earendil-works/pi/pull/7903) by @midastruth
|
||||
(https://github.com/midastruth)).
|
||||
|
||||
### Changed
|
||||
|
||||
- Changed inherited Kimi Coding requests to use pi's runtime User-Agent header.
|
||||
- Replaced the inherited Mistral SDK transport with a native Chat Completions HTTP stream, eliminating its
|
||||
generated client and schema runtime overhead.
|
||||
- Documented the generic AI_AGENT=pi process marker and how it differs from PI_CODING_AGENT=true (#7747
|
||||
(https://github.com/earendil-works/pi/issues/7747)).
|
||||
- Changed inherited OpenAI Responses deferred tool loading to prefer message-anchored additional_tools where
|
||||
supported while retaining tool-search and top-level fallbacks (#7709
|
||||
(https://github.com/earendil-works/pi/issues/7709)).
|
||||
- Reduced inherited fullscreen rendering allocation churn by painting full-width layout rows directly instead
|
||||
of recompositing them on every frame.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed managed-tool downloads delaying TUI startup and hiding diagnostics in fullscreen mode by mounting the
|
||||
TUI first and showing download progress and warnings inside it.
|
||||
- Fixed opening a model selector immediately after startup cancelling and restarting the in-progress model
|
||||
catalog refresh.
|
||||
- Fixed inherited GitHub Copilot login triggering API rate limits while enabling model policies by limiting
|
||||
concurrent policy updates (#6187 (https://github.com/earendil-works/pi/issues/6187)).
|
||||
- Fixed fullscreen transcript search snapping back to the current match during manual scrolling and
|
||||
fragmented mouse input leaking into the search query.
|
||||
- Fixed inherited required LaTeX arguments starting on a new line being parsed as empty (#7760
|
||||
(https://github.com/earendil-works/pi/issues/7760)).
|
||||
- Updated the transitive nanoid development dependency to address a denial-of-service vulnerability.
|
||||
- Fixed fallback rendering for extension tool results to collapse long output and honor tool expansion (#7979
|
||||
(https://github.com/earendil-works/pi/issues/7979)).
|
||||
- Fixed JSON and RPC message_update events dropping cumulative usage during streaming. See JSON Event Mode
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/json.md) and RPC
|
||||
message_update
|
||||
(https://github.com/earendil-works/pi/blob/v0.84.2/packages/coding-agent/docs/rpc.md#message_update-streami
|
||||
ng) (#7982 (https://github.com/earendil-works/pi/pull/7982) by @christianklotz
|
||||
(https://github.com/christianklotz)).
|
||||
- Fixed pi.sendMessage(..., { triggerTurn: false }) steering an active run instead of only recording the
|
||||
custom message (#8022 (https://github.com/earendil-works/pi/pull/8022) by @cristinaponcela
|
||||
(https://github.com/cristinaponcela)).
|
||||
- Fixed the defaultTools setting dropping extension and SDK custom tools when selecting built-in defaults.
|
||||
- Fixed the subagent example rejecting YAML array syntax for the tools frontmatter field (#7598
|
||||
(https://github.com/earendil-works/pi/pull/7598) by @alexsavio (https://github.com/alexsavio)).
|
||||
- Fixed the subagent example dropping parent session model, thinking, and tool configuration (#7897
|
||||
(https://github.com/earendil-works/pi/pull/7897) by @virtuald (https://github.com/virtuald)).
|
||||
- Fixed custom system prompts concatenating the current working directory with later appended prompt content
|
||||
(#7887 (https://github.com/earendil-works/pi/pull/7887) by @distributedlock
|
||||
(https://github.com/distributedlock)).
|
||||
- Fixed inherited OpenAI Responses function and custom tool calls losing namespaces during streaming,
|
||||
proxying, and replay (#7709 (https://github.com/earendil-works/pi/issues/7709)).
|
||||
- Fixed inherited upstream request buffer failures not triggering automatic assistant retries.
|
||||
- Fixed inherited built-in and custom DeepSeek API models sending output limits through an unsupported field.
|
||||
- Fixed inherited Amazon Bedrock replay rejecting tool arguments that contain empty object keys while
|
||||
preserving all valid nested values (#7882 (https://github.com/earendil-works/pi/pull/7882) by @muyiyr
|
||||
(https://github.com/muyiyr)).
|
||||
- Fixed inherited DeepSeek compatibility detection for base URLs whose hostname contains uppercase letters
|
||||
(#7933 (https://github.com/earendil-works/pi/pull/7933) by @yearth (https://github.com/yearth)).
|
||||
- Fixed inherited Google Generative AI and Vertex AI responses with tool calls incorrectly treating
|
||||
output-limit or provider-error stops as normal tool use (#8059
|
||||
(https://github.com/earendil-works/pi/issues/8059)).
|
||||
- Fixed inherited fullscreen mouse drag selection and OSC 8 link activation in terminals that report generic
|
||||
SGR mouse release button codes (#7963 (https://github.com/earendil-works/pi/issues/7963)).
|
||||
- Fixed inherited focused fullscreen overlays not receiving mouse wheel or viewport scroll keys such as
|
||||
PageUp and PageDown (#7894 (https://github.com/earendil-works/pi/issues/7894)).
|
||||
- Fixed inherited LaTeX control spaces split across line endings causing complete expressions to fall back to
|
||||
raw source.
|
||||
- Fixed split Alt+Enter input over SSH being misread as Escape, added PI_TUI_ESC_TIMEOUT for high-latency
|
||||
terminals, and limited that timeout to lone Escape input (#7899
|
||||
(https://github.com/earendil-works/pi/pull/7899) by @powerfooI (https://github.com/powerfooI)).
|
||||
- Fixed inherited idle fullscreen sessions repainting and clearing text selection when the terminal loses
|
||||
focus (#7892 (https://github.com/earendil-works/pi/pull/7892) by @terrorobe
|
||||
(https://github.com/terrorobe)).
|
||||
- Fixed fullscreen selection copy to use the host clipboard and report failure instead of claiming success
|
||||
when OSC 52 is unsupported (#8110 (https://github.com/earendil-works/pi/pull/8110) by @Panoplos
|
||||
(https://github.com/Panoplos)).
|
||||
|
||||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
Model: nvidia/llama-3.3-nemotron-super-49b-v1
|
||||
|
||||
|
||||
i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in
|
||||
sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful
|
||||
failures?
|
||||
|
||||
|
||||
|
||||
read ...
|
||||
|
||||
Validation failed for tool "read":
|
||||
- path: must have required properties path
|
||||
|
||||
Received arguments:
|
||||
{}
|
||||
|
||||
|
||||
|
||||
read sources/01-summarization.md
|
||||
|
||||
|
||||
Based on the provided summaries and findings in sources/ai-version-instructions.md and
|
||||
sources/01-summarization.md (assuming the latter's content is reflected in the large text block you provided,
|
||||
as the initial read tool call failed due to missing path specification), here's a suggested approach to
|
||||
finding meaningful failures in the project, categorized for clarity:
|
||||
|
||||
### 1. Already Identified Failures & Vulnerabilities (from sources/01-summarization.md summary)
|
||||
|
||||
┌─────────────────┬──────────────────────────────────┬──────────────────────────────────────────────────────┐
|
||||
│ Category │ Issue │ Approach to Verify/Resolve │
|
||||
├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤
|
||||
│ Test Failure │ Hardcoded school year in │ Verify: Manually check the test failure post-June │
|
||||
│ │ teachers_controller_test │ 2026. Resolve: Replace hardcoded year with │
|
||||
│ │ │ Year.current_school_year. │
|
||||
├─────────────────┼──────────────────────────────────┼──────────────────────────────────────────────────────┤
|
||||
│ Authorization │ IDOR in OrdersController#update │ Verify: Write integration test to exploit. Resolve: │
|
||||
│ Gap 1 │ │ Add `authorize @order │
|
||||
└─────────────────┴──────────────────────────────────┴──────────────────────────────────────────────────────┘
|
||||
|
||||
Operation aborted
|
||||
|
||||
Model: nvidia/nemotron-3-nano-30b-a3b
|
||||
|
||||
|
||||
i have a project outlined in sources/ai-version-instructions.md. The repo for the project is summarized in
|
||||
sources/01-summarization.md. given the instructions, how would you suggest I approach finding meaningful
|
||||
failures?
|
||||
|
||||
|
||||
We need to answer: "i have a project outlined in sources/ai-version-instructions.md. The repo for the project
|
||||
is summarized in sources/01-summarization.md. given the instructions, how would you suggest I approach
|
||||
finding meaningful failures?" The user presumably wants advice on how to approach finding meaningful failures
|
||||
in code based on the summary provided. We have read the summary; it's about code review and potential issues.
|
||||
The user is asking for approach suggestions: How to find meaningful failures? Should discuss systematic
|
||||
approach: reading code, looking at test failures, security/policy gaps, correctness, edge cases, performance,
|
||||
data integrity, user flows, etc. Possibly using static analysis, test coverage, identify high-risk areas, use
|
||||
observability, monitor production, triage. May suggest reading documentation, scanning issues, looking at
|
||||
failing tests. Could also suggest running the test suite, exploring failing test details, reproducing
|
||||
failures, looking at related code areas. Might also suggest looking at open issues, using linters, performing
|
||||
code review patterns. Perhaps also mention static typing, type checks, dependency vulnerabilities. And refer
|
||||
to the summary: they identified one test failure (time bomb), two authorization gaps, lower-priority items.
|
||||
So suggestion: focus on test failures, policy gaps, etc. Also mention replicating the failure scenario.
|
||||
|
||||
Thus answer should be textual: We can propose steps: 1. Re-run test suite; examine failing tests. 2. Identify
|
||||
areas flagged in summary (e.g., IDOR, missing authorize). 3. Use static analysis tools. 4. Simulate user
|
||||
flows that could expose failures. 5. Look at logs and metrics. 6. Prioritize based on impact. And so on.
|
||||
|
||||
Make answer concise but thorough.
|
||||
|
||||
Below is a practical, step‑by‑step approach you can use to surface “meaningful” failures in a project like
|
||||
the one described in sources/01‑summarization.md. I’m framing it as a checklist you can follow, rather than
|
||||
a set of specific code‑level fixes—so you can adapt it to whatever part of the stack you’re most interested
|
||||
in (Ruby on Rails, the DB schema, the business rules, the test suite, etc.).
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
1. Start with the Known Failures and Flagged Issues
|
||||
|
||||
┌─────────────────────────────────┬─────────────────────────────────────────────────────────────────────────┐
|
||||
│ What to do │ Why it matters │
|
||||
├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤
|
||||
│ Re‑run bundle exec rails test │ The test suite is the single source of truth for what the project │
|
||||
│ (or the equivalent test │ currently considers broken. In the summary the only failing test is the │
|
||||
│ command) and capture the exact │ “time‑bomb” in teachers_controller_test. Re‑producing it will tell you │
|
||||
│ failure output. │ exactly which line/parameter is causing the problem and whether it’s a │
|
||||
│ │ regression or an inherent design flaw. │
|
||||
├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤
|
||||
│ Create a minimal reproduction │ This isolates the failure from the whole test harness and makes it │
|
||||
│ script (or a temporary Rails │ easier to explore edge cases without re‑running the whole suite. │
|
||||
│ console session) that exercises │ │
|
||||
│ the failing test’s path. │ │
|
||||
├─────────────────────────────────┼─────────────────────────────────────────────────────────────────────────┤
|
||||
│ Cross‑reference the failure │ Often the maintainers have already annotated a ticket with priority, │
|
||||
│ with the project’s issue │ intended fix, or known work‑arounds. If not, the ticket itself can │
|
||||
│ tracker (if there’s one). │ become a place to record your findings. │
|
||||
└─────────────────────────────────┴─────────────────────────────────────────────────────────────────────────┘
|
||||
|
||||
### Quick win
|
||||
|
||||
- Fix the time‑bomb by replacing the hard‑coded school‑year literal with a dynamic call
|
||||
(Year.current_school_year). Verify that the fix does not break any other test.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
2. Systematically Scan for High‑Impact Security / Authorization Gaps
|
||||
|
||||
The summary highlighted two concrete IDOR‑style gaps:
|
||||
|
||||
1. Order updates without authorize @order
|
||||
2. Teachers not scoped to their own classrooms in StudentsController / ClassroomEnrollmentsController.
|
||||
|
||||
How to surface similar gaps elsewhere:
|
||||
|
||||
┌─────────────────────────────────────────────────────┬─────────────────────────────────────────────────────┐
|
||||
│ Step │ Tool / Technique │
|
||||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||||
│ a. Map all controller actions that modify domain │ grep -R "def .*update|def .*destroy" │
|
||||
│ objects (e.g., OrdersController#update, │ app/controllers/**/*.rb │
|
||||
│ StudentsController#create, any │ │
|
||||
│ *Controller#update/destroy). │ │
|
||||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||||
│ b. Identify the policy class for each resource │ Look for app/policies/**/*.rb. │
|
||||
│ (OrderPolicy, StudentPolicy, etc.). │ │
|
||||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||||
│ c. Check that every state‑changing action calls │ Run a static‑analysis script like rails │
|
||||
│ authorize (or verify/check) with the correct │ lint:Authorization (if you have a custom linter) or │
|
||||
│ instance variable. │ simply add a comment placeholder TODO: authorize │
|
||||
│ │ @order and search for missing ones. │
|
||||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||||
│ d. Verify that the permitted attributes include the │ `rg "strong_parameters │
|
||||
│ user_id (or an equivalent scoping column). │ │
|
||||
├─────────────────────────────────────────────────────┼─────────────────────────────────────────────────────┤
|
||||
│ e. Simulate an authenticated user from a different │ Use Rails console or a temporary request spec to │
|
||||
│ classroom/role and attempt the unsafe action. │ perform patch /orders/42 as a user who does not own │
|
||||
│ │ the order. │
|
||||
└─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────┘
|
||||
|
||||
### Pattern to repeat
|
||||
|
||||
For each public API endpoint or form POST/ PATCH that touches a model, ask: “If I were a different
|
||||
authenticated user, could I cause an unintended state change?” Anything that returns a 200/302 without a
|
||||
proper authorization check is a candidate “meaningful failure”.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
3. Leverage Test Coverage Metrics to Prioritize
|
||||
|
||||
- Run rails test:coverage (or coverage:install + coverage run) and view the HTML report.
|
||||
- Focus on low‑coverage areas that logically map to risky code paths (e.g., the
|
||||
Admin::PortfolioTransactionsController mentioned in the summary).
|
||||
- Add a single failing test that intentionally violates the expected invariant (e.g., tries to edit a
|
||||
transaction that should be immutable). If it passes, you’ve found a hidden defect.
|
||||
|
||||
Why? Low coverage often indicates parts of the system that have not been exercised by the existing test
|
||||
suite—exactly the sort of blind spot where subtle bugs hide.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
4. Look for Logical Invariants Violated in Production‑Like Scenarios
|
||||
|
||||
Even if tests pass, certain business rules can break under edge‑case data. Use the domain model described in
|
||||
the repo to construct “what‑if” tests:
|
||||
|
||||
┌──────────────────────────────────┬────────────────────────────────────────────────────────────────────────┐
|
||||
│ Invariant │ How to test │
|
||||
├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤
|
||||
│ Portfolio ledger is append‑only; │ Create a series of transactions that appear to double‑spend a │
|
||||
│ balance is derived from │ student’s funds by inserting a pending order, then a second order │
|
||||
│ transactions │ before the first is executed. Verify the derived balance never exceeds │
|
||||
│ │ the true ledger sum. │
|
||||
├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤
|
||||
│ OrderExecutionJob runs at the │ Mock the scheduler (or temporarily change the cron schedule) and │
|
||||
│ correct cadence │ assert that price updates happen only when OrderExecutionJob fires. │
|
||||
├──────────────────────────────────┼────────────────────────────────────────────────────────────────────────┤
|
||||
│ Grade‑improvement deposit │ Simulate a scenario where a teacher “finalizes” a gradebook without │
|
||||
│ calculation is applied only │ the expected A/B grade and watch that no deposit is minted. │
|
||||
│ after finalization │ │
|
||||
└──────────────────────────────────┴────────────────────────────────────────────────────────────────────────┘
|
||||
|
||||
These invariants are often expressed only in design docs (orders-and-transactions.md, gradebook-earnings.md).
|
||||
Treat any deviation as a meaningful failure even if the test suite doesn’t catch it.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
5. Use Static Analysis / Security Scanners for “silent” problems
|
||||
|
||||
┌─────────────────────────────────────────────┬─────────────────────────────────────────────────────────────┐
|
||||
│ Tool │ What it catches that manual inspection might miss │
|
||||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||||
│ brakeman (Rails security scanner) │ Potential XSS, insecure mass_assignment, missing strong │
|
||||
│ │ parameters. │
|
||||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||||
│ rubocop --rails-security │ Unsafe SQL fragments, dangerous File.read usage, etc. │
|
||||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||||
│ bundler-audit │ Known vulnerable gem versions. │
|
||||
├─────────────────────────────────────────────┼─────────────────────────────────────────────────────────────┤
|
||||
│ rails db:seed:raven or similar migration │ Detect migrations that rely on hard‑coded IDs or │
|
||||
│ linters │ non‑idempotent logic. │
|
||||
└─────────────────────────────────────────────┴─────────────────────────────────────────────────────────────┘
|
||||
|
||||
Run these as part of CI and treat any warning that touches the same area as a high‑priority failure until
|
||||
it’s deliberately suppressed with a documented justification.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
6. Monetize Impact vs. Likelihood – Prioritize Fixes
|
||||
|
||||
┌─────────────────────────────────────┬──────────────────────────┬─────────────────────┬────────────────────┐
|
||||
│ Failure type │ Typical impact │ Likelihood (given │ Suggested priority │
|
||||
│ │ │ current code) │ │
|
||||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||||
│ IDOR on order updates (from │ Moderate – a single │ Low (requires │ Medium – fix now; │
|
||||
│ summary) │ student could tamper │ authenticated │ it’s a one‑liner. │
|
||||
│ │ with another's pending │ attacker) │ │
|
||||
│ │ trade. │ │ │
|
||||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||||
│ Teacher‑classroom scoping gap │ High – can cross │ Medium – any │ High – patch │
|
||||
│ │ classroom boundaries, │ teacher can target │ immediately. │
|
||||
│ │ affect many students. │ any classroom. │ │
|
||||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||||
│ Missing authorization on │ Low – only reachable by │ Low‑Medium (needs │ Medium – add │
|
||||
│ PortfolioTransactionsController │ admin URL; but violates │ admin privileges) │ missing authorize │
|
||||
│ edit/update │ immutable‑ledger │ │ calls and write a │
|
||||
│ │ invariant. │ │ unit test. │
|
||||
├─────────────────────────────────────┼──────────────────────────┼─────────────────────┼────────────────────┤
|
||||
│ Time‑bomb hard‑coded year │ Very high – breaks UI │ Certain (will │ Urgent – replace │
|
||||
│ │ for all users after a │ trigger │ with dynamic │
|
||||
│ │ certain date. │ automatically). │ calculation. │
|
||||
└─────────────────────────────────────┴──────────────────────────┴─────────────────────┴────────────────────┘
|
||||
|
||||
When you’re hunting for meaningful failures, start by tackling the failures that move you down this matrix
|
||||
from high impact + high certainty toward low impact + low certainty.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
7. Document the Findings in a Failure‑Report Format
|
||||
|
||||
1. Title – concise description (e.g., “Potential IDOR on Order#update”).
|
||||
2. Reproduction steps – exact request/rails console commands that demonstrate the problem.
|
||||
3. Code locations – file, line numbers, and policy class.
|
||||
4. Impact – who is affected, what state can be changed, downstream consequences.
|
||||
5. Mitigation – minimal code change needed and any required test additions.
|
||||
6. Follow‑up – deeper investigation needed (e.g., audit all admin controllers for missing authorizations).
|
||||
|
||||
Having a structured bug report makes it easy to prioritize, delegate, and track progress.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
### TL;DR Checklist
|
||||
|
||||
1. Run the full test suite; isolate each failing test.
|
||||
2. Map every state‑changing controller action to its policy and verify authorize calls.
|
||||
3. Check coverage; add targeted tests for low‑coverage, high‑risk areas.
|
||||
4. Create “what‑if” data scenarios that test core invariants (ledger immutability, deposit calculation).
|
||||
5. Run static‑analysis/security scanners; treat each finding as a potential failure.
|
||||
6. Prioritize by impact/likelihood and fix the most critical gaps first.
|
||||
7. Document each discovered failure in a standard report format for transparency.
|
||||
|
||||
Following this workflow will help you systematically uncover meaningful failures—whether they’re outright
|
||||
bugs, security‑relevant policy violations, or design inconsistencies—while keeping the effort focused on the
|
||||
parts of the codebase that matter most to the project’s core domain (student earnings, portfolio management,
|
||||
and order execution).
|
||||
|
||||
|
||||
reading sources/task-instructions.md, how do I use the task-catalog?
|
||||
|
||||
|
||||
We need to read sources/task-instructions.md to see instructions on using task-catalog. Probably there is a
|
||||
docs about task-catalog usage. Let's read that file.
|
||||
|
||||
|
||||
read sources/task-instructions.md
|
||||
|
||||
|
||||
The user asks: "reading sources/task-instructions.md, how do I use the task-catalog?" They want guidance on
|
||||
using the task-catalog based on the instructions they just read.
|
||||
|
||||
We need to answer from the content of sources/task-instructions.md, explaining how to use the task-catalog.
|
||||
Probablydescribe steps: exploring the catalog, finding tasks, understanding structure, using it to submit
|
||||
tasks, etc. Should refer to sections like "The Task Catalog" etc. Provide concise actionable guidance.
|
||||
|
||||
Thus respond with explanation, maybe referencing sections like "The Task Catalog" and steps to use it. Should
|
||||
not go beyond policies. Just answer.
|
||||
|
||||
How to use the Task‑Catalog according to sources/task‑instructions.md
|
||||
|
||||
1. Understand what the catalog is
|
||||
- The catalog is a registry of pre‑existing failure scenarios that have already been turned into tasks.
|
||||
- Each entry lives under harbor-tasks/<slug>/ and contains the three core files you will also create for
|
||||
a new task:
|
||||
- instruction.md – the prompt the AI agent receives.
|
||||
- grader‑guidance‑consolidated.md – the grader‑specific guidance.
|
||||
- Reference runs (recorded trials) that prove the failure actually occurs.
|
||||
|
||||
|
||||
2. Browse the catalog
|
||||
- Open the Task Catalog link that the project provides in the Quick Links section.
|
||||
- You’ll see a list of slugs, each representing a separate task. Click a slug to view its folder
|
||||
structure.
|
||||
- Look at the instruction.md and grader‑guidance‑consolidated.md to see how failures were documented and
|
||||
scored. This gives you a concrete template for the kind of failure you need to capture.
|
||||
|
||||
3. Decide whether to reuse or duplicate
|
||||
- Check for duplicates. Before you start a new task, verify that an existing catalog entry does not
|
||||
already cover the failure you have in mind. Duplicate submissions are rejected at review.
|
||||
- If you find a close match but need a different angle (e.g., a different severity, a different code
|
||||
path), you can still add a new slug; just make sure the prompt and ground‑truth details are distinct.
|
||||
|
||||
4. Create your own task folder
|
||||
- Follow the scaffold shown in the catalog:
|
||||
```
|
||||
harbor-tasks/<your‑slug>/
|
||||
├─ instruction.md # engineering prompt
|
||||
├─ task.toml # metadata (repo commit, browser flag, etc.)
|
||||
├─ tests/
|
||||
│ └─ grader‑guidance‑consolidated.md
|
||||
└─ reference‑runs/ # recorded trials that demonstrate the failure
|
||||
```
|
||||
- The Workspace & workspace.patch section of the instructions tells you how to generate a patch that
|
||||
captures any workspace changes you made while exploring.
|
||||
|
||||
5. Populate the files
|
||||
- instruction.md – Write a realistic engineering request that would naturally expose the failure you
|
||||
discovered.
|
||||
- Keep hints out; the prompt must not point directly at the solution.
|
||||
- Ensure the prompt is self‑contained (all needed context must be inside the patched workspace).
|
||||
|
||||
- grader‑guidance‑consolidated.md – Translate the failure into the eight‑criterion grader language:
|
||||
- Provide Task context and Business context (if any).
|
||||
- Define Ground truth (exact file/line, correct fix, test that should catch it).
|
||||
- For each of the eight criteria, describe what a strong vs. weak response looks like on this task.
|
||||
- Add any heavy penalties only for deal‑breaker behaviours, naming the criterion they affect.
|
||||
|
||||
- Reference runs – Run harbor-run (or codex/claude depending on the agent you chose) to generate recorded
|
||||
trials that trigger the failure. Copy those runs into reference‑runs/ using npx tsx
|
||||
scripts/copy-reference-run.ts ….
|
||||
|
||||
6. Run detectors before submitting
|
||||
- The catalog workflow expects you to invoke every detector skill (e.g., /detector‑meaningful-failure,
|
||||
/detector‑fact‑check‑rubric‑claims, etc.).
|
||||
- Detectors will flag stale inputs, broken environment checks, cross‑task references, or missing
|
||||
heavy‑penalty definitions. Fix any issues they surface, then re‑run the detectors.
|
||||
|
||||
7. Validate and package
|
||||
- Execute npx tsx scripts/submit‑task.ts <your‑slug> .
|
||||
- The script will:
|
||||
1. Verify that all required files exist.
|
||||
2. Ensure the patch is clean (no unintended binaries, ignored files, or leaked snapshots).
|
||||
3. Confirm you have at least four accepted reference runs.
|
||||
4. Produce a single tarball that you upload to the platform.
|
||||
|
||||
- If any warnings appear, address them before you click Submit on the platform UI.
|
||||
|
||||
8. Submit with the required metadata
|
||||
- When you upload the tarball, also fill the Import/Export panel:
|
||||
- Export a JSON save‑state of the task (so you can rebuild later).
|
||||
- Paste your Slack thread URL (the place where you asked/answered questions).
|
||||
- Indicate whether this is a complete submission or a work‑in‑progress draft.
|
||||
|
||||
- If you suspect the grader scored a run incorrectly, you can flag it with the grader‑performance
|
||||
checkbox; this creates a separate concern thread for the review team.
|
||||
|
||||
9. Iterate via the Feedback Loop
|
||||
- After submission, reviewers will run the same detector suite on your tarball.
|
||||
- Any issues they raise (e.g., stale inputs, missing heavy‑penalty definitions) will be sent back as
|
||||
feedback.
|
||||
- Update your task files, re‑run the detectors, and submit a new version if needed.
|
||||
|
||||
────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
### Quick‑Reference Cheat‑Sheet
|
||||
|
||||
┌────────────────┬────────────────────────────────────────────┬─────────────────────────────────────────────┐
|
||||
│ Step │ Command / Action │ Result │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Explore │ Open the Task Catalog link → browse slugs │ See existing instruction.md & │
|
||||
│ │ │ grader‑guidance‑consolidated.md examples │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Create folder │ mkdir -p harbor-tasks/<slug>/ │ Scaffold for new task │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Generate patch │ Work in Explore → │ environment/workspace.patch captured │
|
||||
│ │ /create-snapshot:snapshot → │ │
|
||||
│ │ snapshot-to-task.ts │ │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Write prompt │ instruction.md → realistic, no hints │ Agent receives clear engineering request │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Write grader │ tests/grader‑guidance‑consolidated.md → 8 │ Grader knows exactly what to score │
|
||||
│ guidance │ criteria + heavy penalties │ │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Run trials │ harbor-run (or codex/claude) → copy runs │ reference-runs/ populated │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Run detectors │ /detector‑* skills │ All automated checks pass │
|
||||
├────────────────┼────────────────────────────────────────────┼─────────────────────────────────────────────┤
|
||||
│ Submit │ npx tsx scripts/submit‑task.ts <slug> → │ Task packaged & ready for review │
|
||||
│ │ upload tarball │ │
|
||||
└────────────────┴────────────────────────────────────────────┴─────────────────────────────────────────────┘
|
||||
|
||||
Follow the flow Explore → Build → Validate → Submit exactly as the Task‑Catalog section of
|
||||
sources/task‑instructions.md describes, and you’ll be able to add a new, meaningful failure to the catalog
|
||||
without duplicating existing work.
|
||||
|
||||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
───────────────────────────────────────────────────────────────────────────────────────────────────────────────
|
||||
~/workspaces/dataannotation/current-project (raccoon-stocks)
|
||||
↑30k ↓4.5k R4.2k CH23.3% 15.1%/131k (auto) (nvidia) nvidia/nemotron-3-nano-30b-a3b • medium
|
||||
mode: implementation
|
||||
Reference in New Issue
Block a user