Hiro development journal

Approval-system audit finds idea/artifact conflation and invalid contrast evidence

Diagnostic completed; live system observed but not modified Machine-readable JSON

Executive summary

A forensic review of the live continuous-improvement queue confirms that the extreme rejection rate is substantially caused by approval-system design, not evidence that nearly every underlying idea is harmful.

The durable queue currently contains 173 rejected and 3 implemented ideas, a 98.3 percent rejection share among terminal rejected-or-implemented records. Since recovery event 1395, the latest classified outcomes for 34 ideas comprise 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass.

The engine currently converts a repairable candidate-construction failure into a terminal idea rejection after the bounded attempt budget is exhausted. This conflates failure to manufacture an evaluable patch/test artifact with evidence that the idea itself lacks merit.

The single recent candidate-gate pass exposed the opposite defect. Its untouched-baseline run failed with TypeError because the candidate-authored test called the baseline chat function with a keyword its baseline signature did not accept. Pytest represented this call-phase exception as a JUnit failure, and Hiro's count-only parser classified any such failure as an attributable assertion contrast.

The live canary cannot correct that false positive because its default probe simply reruns the same single candidate-authored test at synthetic checkpoints. The approval stages therefore reuse correlated evidence instead of adding an independent behavioral check.

Work completed

Live rejection-funnel audit

Completed
  • Read the durable continuous-improvement SQLite state and append-only event history without mutating the queue.
  • Observed 173 rejected, 18 superseded, 18 queued, 3 implemented, 1 candidate, and 1 canary record.
  • Among terminal rejected-or-implemented records, 173 of 176 are rejected, or 98.3 percent.
  • Classified the latest relevant outcomes for 34 ideas since recovery event 1395: 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass.
  • The distribution shows that the dominant bottleneck is artifact construction and test manufacture, not the global evaluation policy deciding that otherwise valid changes are negative.

Terminal-state semantics

Defect confirmed
  • Traced continuous_engine.py candidate handling: non-safety candidate failures are retried only while attempts remain, then the record is unconditionally transitioned to state rejected with event candidate_rejected.
  • This terminal path applies even when the candidate result identifies the failure as repairable and retryable.
  • Common terminal reasons include empty or invalid patch plans, ambiguous edit anchors, candidate validation failure, and candidate-local regression failure.
  • Those observations establish that rejected currently means either the proposed improvement failed or the builder failed to create a valid experiment. The dashboard and ranking system cannot distinguish the two.

Targeted-contrast evidence audit

Defect confirmed
  • Inspected the passing writing-assistance candidate sandbox-20260817043030-a2622db6-01 and its frozen recommendation packet.
  • Its targeted baseline contrast returned pytest code 1 with one JUnit failure and no JUnit errors, which the current targeted_contrast_status function accepts as improvement_demonstrated.
  • The captured traceback shows that baseline execution stopped at TypeError: chat() got an unexpected keyword argument 'message'. No behavioral assertion established that the baseline failed the writing-assistance contract.
  • Pytest's JUnit implementation writes ordinary call-phase test failures through the failure element regardless of whether the exception is AssertionError or another exception. Hiro currently retains only aggregate tests, failures, errors, and skipped counts, discarding exception type and assertion provenance.
  • The earlier assertion-aware gate therefore eliminated collection and setup errors but did not actually distinguish assertion failures from call-phase runtime exceptions.

Canary independence audit

Weakness confirmed
  • The active writing-assistance candidate entered canary after the invalid baseline contrast and had completed synthetic checkpoints at 0, 5, and 15 minutes at audit time.
  • Each default canary probe invokes the candidate's own tests tuple. For this candidate that is one generated test, tests/autonomous/test_continuous_17afe3b06f2e.py.
  • The three recorded probes therefore repeat the same test that helped admit the candidate; they are temporal repetitions, not independent measurements.
  • The final 60-minute checkpoint remained pending during the audit. No live interaction, separately authored oracle, or category-peer replay had yet supplied independent canary evidence.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Durable queue state passed Read-only SQLite queries returned 173 rejected, 18 superseded, 18 queued, 3 implemented, 1 candidate, and 1 canary record. Terminal rejected-or-implemented rejection share is 98.3 percent.
Recent funnel classification passed Read-only event classification since event 1395 found latest relevant outcomes for 34 ideas: 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass.
False-negative control-flow trace passed Source inspection confirmed that after the bounded candidate-attempt limit, continuous_engine.py transitions candidate_failed results to rejected even when they are non-safety, repairable construction failures.
False-positive evidence trace passed The frozen recommendation packet for sandbox-20260817043030-a2622db6-01 records a baseline TypeError caused by an unsupported keyword argument, while the aggregate JUnit counts satisfied the current improvement_demonstrated rule.
Canary independence trace passed Source and durable canary state confirm that default probes rerun the candidate tests; the active candidate's 0, 5, and 15 minute checkpoints each reran the same one-test file and were marked synthetic.
Hiro code changes not run No Hiro source, queue state, service process, candidate, or canary was modified during this diagnostic session.
Journal test and production build passed npm run test:hiro passed. The first build attempt overlapped with a still-finishing generator process and encountered a transient Windows ENOTEMPTY race in the generated public/hiro directory; after confirming no process remained, npm run build generated and validated 141 journal entries and completed TypeScript and Vite production compilation successfully.

Current state

Next steps