Hiro development journal

Correcting false rejection and false approval in the continuous queue

Implemented, adversarially tested, committed, and active in the live queue Machine-readable JSON

Executive summary

Hiro's approval system now distinguishes an idea verdict from failure to manufacture a valid patch and test artifact. Exhausted repairable construction attempts enter artifact_blocked instead of rejected, remain visible and ranked, and are retried once after the builder or approval harness revision changes.

Untouched-baseline contrast now records per-test phase, exception class, assertion identity, and contract marker. A baseline improvement is attributable only when every recorded failure is a call-phase AssertionError carrying a non-empty hiro_contract identifier. Runtime exceptions such as TypeError are inconclusive even when pytest reports them as JUnit failures.

Canary probes no longer rerun the candidate-authored targeted test. They run fixed mechanism-specific harness tests and, for captured interaction incidents, execute the originating code-owned replay inside the candidate worktree. The replay explicitly records that no candidate-authored test or production endpoint was used.

The invalid writing-assistance canary was quarantined before its final checkpoint. A read-only-to-append-only migration moved 136 clearly non-merit historical failures out of rejected. After restart and one automatic revision-triggered retry, the live queue reports 136 artifact blocked, 37 rejected, 18 waiting, 3 retrying, 1 active candidate, and 3 implemented.

The final code passed 647 repository tests. Focused approval, state-separation, dashboard, candidate worktree, and canary diagnostics passed 59 tests, including real subprocess reproductions of the exact unsupported-keyword TypeError defect.

Work completed

Assertion-provenance contrast

Completed
  • Added a harness-owned pytest plugin that records failed test node id, execution phase, exception type, AssertionError classification, hiro_contract marker, and stable assertion identifier into a bounded JSON receipt.
  • The builder and evaluator copy that trusted plugin into each temporary untouched-baseline worktree rather than importing candidate-controlled instrumentation.
  • Candidate construction prompts now require @pytest.mark.hiro_contract with a stable contract identifier and explicitly state that unmarked assertions, TypeError, and other runtime exceptions are invalid evidence.
  • The shared targeted-contrast classifier requires a non-empty provenance report, at least one marked call-phase assertion, no unexpected failures, matching JUnit failure counts, executed tests, zero JUnit errors, and pytest's ordinary test-failure return code.
  • Repair context retains the assertion-provenance receipt and reports the exception types responsible for inconclusive evidence.

Idea and artifact state separation

Completed
  • Added artifact_blocked as a closed but non-merit queue state with its own lifecycle timestamp, count, list, outcome category, and append-only events.
  • Construction failures, candidate regression failures, inconclusive or already-passing targeted contrasts, validation failures, patch failures, and otherwise unattributable candidate failures no longer become idea rejection after retries are exhausted.
  • Valid negative evaluation evidence such as score delta, category regression, latency ratio, confidence overlap, or statistical-separation failure remains rejected. Safety failures remain rejected immediately.
  • Artifact-blocked outcomes record idea_merit_evaluated false, retriable_after_builder_change true, and the builder revision that blocked them.
  • Each continuous tick may reactivate the highest-ranked blocked artifact once when its recorded builder revision differs from the active revision. It resets candidate artifacts and attempts and re-enters the ordinary ranked queue.

Independent canary evidence

Completed
  • Candidate-authored tests are explicitly excluded from the default canary probe.
  • Each supported mechanism maps to fixed repository-owned regression tests selected independently of the generated patch and test.
  • Captured interaction candidates carry their private originating replay fixture and code-owned task contract into the canary record.
  • A candidate-worktree subprocess runs the originating replay through that candidate's actual response boundary and returns a structured receipt without using the production endpoint.
  • Canary fails closed when independent tests are absent, overlap the candidate-authored tests, fail, or when an available harness replay fails.

Historical queue correction and quarantine

Completed
  • Stopped only Hiro's verified main.py process tree before changing approval behavior; the Qwen model server remained loaded.
  • Quarantined incident-34d35a2be0b229ecb9ae8139 with event candidate_quarantined because its untouched-baseline run raised TypeError before reaching any behavioral assertion.
  • Reclassified 136 clearly non-merit historical records using append-only historical_artifact_failure_reclassified events. The migration covered explicit construction failures, generic legacy candidate_failed construction outcomes, candidate regression failures, invalid targeted contrasts, patch and validation failures, and infrastructure blocking.
  • Did not reclassify score, latency, category, or confidence-based evaluation rejections, the deliberate negative acceptance seed, or six legacy records without enough evidence to infer a different status.
  • After restart the invalid canary remained artifact_blocked, the stale pre-upgrade concept-explanation candidate was requeued, and a different blocked idea re-entered candidate construction on the new commit.

Benchmark-page semantics

Completed
  • Added Artifact blocked as a separate pipeline statistic, ranked-idea filter, section, lifecycle milestone, and count.
  • Changed Rejected copy to describe valid negative, safety, and legacy evidence rather than construction failures.
  • Removed construction-failure totals from the rejected summary and added legacy-unclassified reporting.
  • Artifact-blocked cards state that idea merit was not evaluated and do not expose normal queued priority controls until reactivated by a builder revision.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Exact TypeError adversarial reproduction passed A real pytest subprocess called a baseline-compatible chat fixture with the unsupported keyword message. JUnit reported one failure and zero errors, matching the original defect, while the provenance plugin recorded builtins.TypeError and the classifier returned inconclusive.
Positive assertion control passed A real pytest subprocess with a marked call-phase AssertionError produced builtins.AssertionError, retained its stable contract identifier, and was classified improvement_demonstrated. An otherwise identical unmarked assertion was inconclusive.
Focused approval and queue diagnostics passed 59 tests passed across targeted provenance, candidate builder, candidate evaluator, continuous queue, independent canary, and benchmark-page filtering. Coverage includes exhausted construction becoming artifact_blocked, valid latency rejection remaining rejected, historical migration, revision-triggered retry, and candidate-test exclusion.
Full Hiro regression suite passed 647 tests passed in 188.41 seconds on the exact final committed tree. An earlier pre-dashboard finalization run also passed 647 tests in 203.68 seconds.
Independent harness canary replay passed A writing-assistance fixture ran through hiro.improvement.canary_probe in a child process. The receipt passed its code-owned contract, recorded candidate_authored_test_used false, and recorded production_endpoint_used false.
Historical state migration passed 136 historical records were reclassified from rejected to artifact_blocked through append-only events. Together with the quarantined canary, the pre-restart state contained 137 artifact-blocked records and 37 remaining rejections.
Windows child-process restart passed Hiro restarted through scripts/start_hiro.py, which canonicalizes the process-scoped Path while preserving the inherited runtime environment. Ports 8000, 8001, and 8765 are owned by the verified Hiro process, and the benchmark endpoint returns HTTP 200.
Live queue behavior on commit 908cb83 passed The live queue requeued the stale pre-fix candidate, reactivated one different artifact-blocked idea, and started it with base 908cb83b273dffad89256fdd413470ad67527107. The invalid writing candidate remains artifact_blocked with candidate_quarantined as its latest event.
Live model and dashboard passed The local model endpoint reports qwen/qwen3.8-27b. The benchmark page returns HTTP 200 and contains the Artifact blocked filter and count.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 142 journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps