Commit Graph

24 Commits

Author SHA1 Message Date
ff1bed2f39 a few changes 2026-09-25 14:47:27 -04:00
dd68f8679e reran most of the detectors before needing to fix rubric 2026-09-25 14:30:05 -04:00
e6ebc5c5c0 after store-atomic-grades 2026-09-25 13:33:06 -04:00
8fe923e6dc copied the orig folder to current folder 2026-09-25 12:44:47 -04:00
5b010039d7 ren worker folder adding orig, mv new one into root 2026-09-25 10:34:29 -04:00
e758db67ef work to rubric-regrades - stage-atomic-rubric 2026-09-24 22:07:51 -04:00
8149c677c4 after rerunning the regrades 2026-09-24 22:06:08 -04:00
bc7e65dcfb removed temp staged files 2026-09-24 10:31:41 -04:00
bc57513628 chore: stored atomic grades and ran final 2 detectors a 2026-09-24 10:29:30 -04:00
e23a1af1ea chore: add harbor-tasks and harbor-jobs 2026-09-23 20:35:51 -04:00
00408fd43c fix: ran atomic-rubric and grader-context creation skill 2026-09-22 20:19:59 -04:00
a7b6b57cab fix: rename yml file to yaml 2026-09-22 15:09:06 -04:00
242ea3d8bb chore: 1st cut at grader and atomic rubrics 2026-09-22 14:41:42 -04:00
10583d64ef fix: a few things in rubric
Removed

  ┌───────────────────────────────────────────┬─────────────────────────────────────────────────────────────────────────────────────────┐
  │                   Claim                   │                                           Why                                           │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ 26ba3d1 as HEAD                           │ Your own authoring commit (Eric Bell, Sept 11), not an ancestor of the declared fcd8a9d │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ "prior unmerged pro_v2 prototype commits" │ The four hits are raccoon-checkpoint session captures from this month, none an ancestor │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ Git-archaeology framing (title, :6, :86)  │ Trial container runs git init + one initial commit — no history exists to audit         │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ theProject-voice                          │ Mismatched task.toml's potion-voice; the rubric doesn't need to name the repo           │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ Eight "Evaluates whether…" preambles      │ Paraphrased the shared standard the grader already has verbatim                         │
  └───────────────────────────────────────────┴─────────────────────────────────────────────────────────────────────────────────────────┘

  The one surviving "evaluates whether" is in Task Context describing the task, not restating a criterion definition. That one's fine.

  Preserved and restored

  All your structural improvements survive: the Heavy Penalties section, the Integrity routing note, the TP/Narrow-Correctness decoupling note, the Fabricated
  Verification penalty, and the Ground-Truth cross-reference that removes the duplicate enumeration.

  Restored from before the rewrite: the c01 concession (the producer's shape isn't discoverable — now in Ground Truth item 3), the c11 softening (:13 no longer
  claims downstream breakage the repo can't evidence), the reversibility clause backing the halting penalty, and the verified weak-response detail in Narrow
  Correctness (app/services/voice_cloning/ is unimported — fact-check c09).

  Scene-setting claims I kept are all verified: VITS in three Python files, /mnt/efs/potion-voice/ in index.js:134, .fifo queue URLs in the pm2 configs.
2026-09-22 13:54:00 -04:00
58d17d7b26 fix: holistic-rubric changes; penalty names its target
Reference issues or pull requests here (e.g., "Closes #123")
● Applied. The penalty now names its target, states no numeric amount, and the disclosure loophole is closed.

  :68

  ▎ Over-Engineering / Unrequested Architecture: Apply a heavy penalty to Thought Partnership when the response adds tier infrastructure the repository neither
  ▎ asks for nor supports — a tier-routing module, schema fields such as VoiceCloning.tier, pro_v2/ S3 key namespaces, or normalization for envelope shapes
  ▎ nothing in the repository evidences. Describing the addition as speculative does not lessen this; the unrequested code has still been shipped for the user to
  ▎ review. Severity scales with how much was built. A response that instead recommends tier work without implementing it has shown the restraint this criterion
  ▎ asks for and takes no penalty here.

  :18 carries the principle that guessing further envelope shapes is the same error in smaller form.

  Two things to note.

  The document is now 1,171 words, up from 1,012 at the start of the session. Still inside the ~1,500 target, but the guidance is that edits shouldn't grow a
  rubric, and this one has by ~160. Most of it is load-bearing content that genuinely wasn't there. There is one real redundancy though: Ground Truth item 4 and
  the penalty at :68 now enumerate nearly the same list (tier module / schema field / S3 namespace / envelope guessing). Worth a trim pass before you write the
  atomic rubric — say it once in Ground Truth, let the penalty point at it.

  Last rubric item is the dimension-misapplication pair:
  1. Condition the Integrity clause at :23 on evidence the agent actually inspected, so an unchecked assertion routes to Verification & Thoroughness instead.
  2. Decouple the TP 1.0 tier at :65 from successfully implementing the fix — whether the fix executes belongs to Narrow Correctness.
2026-09-22 13:41:36 -04:00
7d640114bb fix: Rubric state after this session
- Undefined Async Engineering Standard removed from both sites; the foreclosure now rests on a Ground Truth fact instead of a named authority
  - Duplicated grounding clause repaired
  - c01 unreachable premise reframed across all four sites — :4, :29, :30, :53
  - c07 narrowed to tier-specific model checkpoints
  - c11 S3 claim softened to what the repo actually supports

  Every load-bearing claim the fact-check flagged is now either true against fcd8a9d or no longer gating credit.

  Still open on the rubric — one edit, two decisions:

  1. Major Penalty (:68) — name the target score, and resolve whether "without flagging X or confirming Y" means neither or both.
  2. Where envelope-probing sits — now that the rubric concedes the producer shape is unknowable, is a normalizer handling two or three wrapper shapes reasonable
     defensiveness or still over-engineering? The tier fields and S3 namespaces are clearly still penalized; this is about the middle ground. Worth resolving in
     the same edit so :68 and :4 don't contradict each other.
  3. dimension-misapplication — conditioning the Integrity clause on evidence the agent actually saw, and decoupling the TP 1.0 tier from successfully
     implementing the fix.

  Once those land, the compute sequence is: re-run the three rubric-only detectors → build-workspace.sh (to settle the missing test-commands.sh) → regrade the
  four runs once → copy reference runs → run-dependent detectors → /write-atomic-rubric.
2026-09-22 13:25:23 -04:00
08555c13aa fix: changed thought partnership 2026-09-22 13:14:48 -04:00
f2b8e617d6 chore: fixed rubric name, add detectors run and add theFailure.md
theFailure is my answer to the 'describe the failure' question.
2026-09-21 19:26:47 -04:00
41f21f8392 claudes updates to the holistic-rubric 2026-09-18 16:46:02 -04:00
b4c0d24083 changes in flux for the rubric 2026-09-18 16:32:34 -04:00
0311e3e3c7 clean up old rubic file and make ver 2 the current one 2026-09-18 16:04:20 -04:00
7254922982 new: added rubric varieties and failures from docs
meaningful-failures - from the docs - the entire page
instructions and holistic-rubrics are variations - #2 is the latest.
2026-09-18 15:46:48 -04:00
e01a9d3425 after build-workspace 2026-09-17 06:24:22 -04:00
530711cd8a 1st harbor tasks scaffold 2026-09-17 06:15:33 -04:00