Commit Graph

60 Commits

Author SHA1 Message Date
2b6898ab84 Tried to package, had to rerun all detectors 2026-09-27 07:28:32 -04:00
29c23a4a8f Restore staged files step 2026-09-27 06:14:05 -04:00
a3dd68aede ran re-grading 2026-09-27 06:12:04 -04:00
040251f69c ran final detector before re-grading 2026-09-27 06:00:59 -04:00
f338dfe04e ran last 2 detectors, fixed issues 2026-09-27 05:52:27 -04:00
c711a6d5b0 new holistic rubric before humanizing 2026-09-27 05:06:02 -04:00
12321a6013 det rubric coverage had problems 2026-09-27 04:54:38 -04:00
74d4534d59 re-graded 4 runs - better 2026-09-27 04:48:12 -04:00
8a63ac2209 graded 4 runs 2026-09-26 20:22:35 -04:00
5bc44d1a1e generated atomic rubric 2026-09-26 19:39:34 -04:00
e14abf6490 copied reviewed trials 2026-09-26 19:30:23 -04:00
a15c794612 last 2 pre-trial detectors 2026-09-26 18:52:46 -04:00
e55fd1f41e all 3 done 2026-09-26 18:50:39 -04:00
eeafd74131 2 of 3 detector issue solved 2026-09-26 18:46:28 -04:00
67217f46fc detector again 2026-09-26 18:39:23 -04:00
620c9e2c4b 1st edits to holistic rubric and reran detectors 2026-09-26 16:37:24 -04:00
ababc68dc5 most of the pre-trial detectors - some raise issues 2026-09-26 16:30:02 -04:00
ae2dd295e2 added holistic-rubric 2026-09-26 16:09:49 -04:00
698c5bd731 authoring up, create mishandled_pro_v2 2026-09-26 15:49:09 -04:00
e55ccea018 Loaded up for the 3rd redo
Still on potion-voice
2026-09-26 14:57:10 -04:00
bceb52e8ee lots of change - all to start my 3rd redo 2026-09-26 14:31:52 -04:00
d54276de34 broken-dev-env detector show me git archeology never exists
This means I have to redo the entire project.
2026-09-26 13:47:40 -04:00
530c85f50c after rem temp staged files step 2026-09-26 13:12:48 -04:00
6856e75265 after updates to atomic rubric, rerun detector 2026-09-26 13:11:29 -04:00
18574c0ca6 final detectors - 1 problem 2026-09-26 13:06:56 -04:00
0bd21e0d88 one regrade - no edits before hand 2026-09-25 20:22:19 -04:00
2aa9ab7ff4 fixed .gitignore, regrades 2026-09-25 20:08:25 -04:00
fe1f0de788 chmod 2026-09-25 19:25:39 -04:00
abf389d7fc Converted mishandle_pro_v2 using the write-atomic-rubric rules:
Created 14 criteria in harbor-tasks/mishandle_pro_v2/tests/atomic-rubric.yaml:1.
Extracted context verbatim into harbor-tasks/mishandle_pro_v2/tests/grader-context.md:1.
Correctly used zero Crux criteria because penalties target individual dimensions.
Staging validation passed.
Form and coverage detectors both report clear with HIGH confidence.

The rubric is currently staged for grading; use --restore before packaging.
2026-09-25 16:26:24 -04:00
87b54b8a98 deleted old trails 2026-09-25 16:15:31 -04:00
e2e3f73ad5 final two detectors 2026-09-25 15:37:33 -04:00
9ec7917e12 both clear now 2026-09-25 15:33:44 -04:00
8031ff1d4c several turns - persistenc is the issue now 2026-09-25 15:29:35 -04:00
9956c5614e holistic-rubric fixes
dimmention-misapplication is fine.
rubric-clarity has an issue.
-- Verdict: material issues. The earlier ambiguities are fixed, but Ground Truth accepts “unmerged or deprecated” prototype commits while the top Thought Partnership tier requires “unmerged” commits. That difference could change the score for an accurate response.
2026-09-25 15:09:53 -04:00
b1188a7b62 holistic-rubric still in flight 2026-09-25 14:57:29 -04:00
ff1bed2f39 a few changes 2026-09-25 14:47:27 -04:00
dd68f8679e reran most of the detectors before needing to fix rubric 2026-09-25 14:30:05 -04:00
e6ebc5c5c0 after store-atomic-grades 2026-09-25 13:33:06 -04:00
8fe923e6dc copied the orig folder to current folder 2026-09-25 12:44:47 -04:00
5b010039d7 ren worker folder adding orig, mv new one into root 2026-09-25 10:34:29 -04:00
e758db67ef work to rubric-regrades - stage-atomic-rubric 2026-09-24 22:07:51 -04:00
8149c677c4 after rerunning the regrades 2026-09-24 22:06:08 -04:00
bc7e65dcfb removed temp staged files 2026-09-24 10:31:41 -04:00
bc57513628 chore: stored atomic grades and ran final 2 detectors a 2026-09-24 10:29:30 -04:00
e23a1af1ea chore: add harbor-tasks and harbor-jobs 2026-09-23 20:35:51 -04:00
00408fd43c fix: ran atomic-rubric and grader-context creation skill 2026-09-22 20:19:59 -04:00
a7b6b57cab fix: rename yml file to yaml 2026-09-22 15:09:06 -04:00
242ea3d8bb chore: 1st cut at grader and atomic rubrics 2026-09-22 14:41:42 -04:00
10583d64ef fix: a few things in rubric
Removed

  ┌───────────────────────────────────────────┬─────────────────────────────────────────────────────────────────────────────────────────┐
  │                   Claim                   │                                           Why                                           │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ 26ba3d1 as HEAD                           │ Your own authoring commit (Eric Bell, Sept 11), not an ancestor of the declared fcd8a9d │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ "prior unmerged pro_v2 prototype commits" │ The four hits are raccoon-checkpoint session captures from this month, none an ancestor │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ Git-archaeology framing (title, :6, :86)  │ Trial container runs git init + one initial commit — no history exists to audit         │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ theProject-voice                          │ Mismatched task.toml's potion-voice; the rubric doesn't need to name the repo           │
  ├───────────────────────────────────────────┼─────────────────────────────────────────────────────────────────────────────────────────┤
  │ Eight "Evaluates whether…" preambles      │ Paraphrased the shared standard the grader already has verbatim                         │
  └───────────────────────────────────────────┴─────────────────────────────────────────────────────────────────────────────────────────┘

  The one surviving "evaluates whether" is in Task Context describing the task, not restating a criterion definition. That one's fine.

  Preserved and restored

  All your structural improvements survive: the Heavy Penalties section, the Integrity routing note, the TP/Narrow-Correctness decoupling note, the Fabricated
  Verification penalty, and the Ground-Truth cross-reference that removes the duplicate enumeration.

  Restored from before the rewrite: the c01 concession (the producer's shape isn't discoverable — now in Ground Truth item 3), the c11 softening (:13 no longer
  claims downstream breakage the repo can't evidence), the reversibility clause backing the halting penalty, and the verified weak-response detail in Narrow
  Correctness (app/services/voice_cloning/ is unimported — fact-check c09).

  Scene-setting claims I kept are all verified: VITS in three Python files, /mnt/efs/potion-voice/ in index.js:134, .fifo queue URLs in the pm2 configs.
2026-09-22 13:54:00 -04:00
58d17d7b26 fix: holistic-rubric changes; penalty names its target
Reference issues or pull requests here (e.g., "Closes #123")
● Applied. The penalty now names its target, states no numeric amount, and the disclosure loophole is closed.

  :68

  ▎ Over-Engineering / Unrequested Architecture: Apply a heavy penalty to Thought Partnership when the response adds tier infrastructure the repository neither
  ▎ asks for nor supports — a tier-routing module, schema fields such as VoiceCloning.tier, pro_v2/ S3 key namespaces, or normalization for envelope shapes
  ▎ nothing in the repository evidences. Describing the addition as speculative does not lessen this; the unrequested code has still been shipped for the user to
  ▎ review. Severity scales with how much was built. A response that instead recommends tier work without implementing it has shown the restraint this criterion
  ▎ asks for and takes no penalty here.

  :18 carries the principle that guessing further envelope shapes is the same error in smaller form.

  Two things to note.

  The document is now 1,171 words, up from 1,012 at the start of the session. Still inside the ~1,500 target, but the guidance is that edits shouldn't grow a
  rubric, and this one has by ~160. Most of it is load-bearing content that genuinely wasn't there. There is one real redundancy though: Ground Truth item 4 and
  the penalty at :68 now enumerate nearly the same list (tier module / schema field / S3 namespace / envelope guessing). Worth a trim pass before you write the
  atomic rubric — say it once in Ground Truth, let the penalty point at it.

  Last rubric item is the dimension-misapplication pair:
  1. Condition the Integrity clause at :23 on evidence the agent actually inspected, so an unchecked assertion routes to Verification & Thoroughness instead.
  2. Decouple the TP 1.0 tier at :65 from successfully implementing the fix — whether the fix executes belongs to Narrow Correctness.
2026-09-22 13:41:36 -04:00