Hiro development journal

Overnight rejections traced to candidate-construction and repair-loop friction

Diagnosed; corrective implementation pending Machine-readable JSON

Executive summary

The overnight run produced no promotions, but the thirty-six terminal outcomes were candidate-workflow rejections rather than Stage 6 promotion decisions. No candidate reached the canary or stable promotion governor.

Thirty-one of the thirty-six workflows failed during isolated candidate construction. Every one used all three construction attempts, and the queue preserved only a generic candidate_failed reason instead of the detailed build evidence.

Five candidates reached independent public and held-out evaluation and were legitimately rejected under the current policy because they did not demonstrate the required improvement, and some also regressed latency or a capability category.

The dominant problem is excessive engineering friction in the candidate builder and its repair loop, not evidence that every overnight idea lacked merit. The live process needs repaired retry semantics and complete failure propagation before its rejection rate is meaningful.

Work completed

Queue and promotion reconciliation

Completed
  • The production queue was reconciled from the durable SQLite event ledger against every referenced frozen candidate packet.
  • Thirty-six ideas entered the overnight candidate path and all thirty-six ended in the rejected queue state after three queue-level attempts.
  • The ledger contained no canary, governor, promotion, or implementation event for these ideas. Describing them as failed promotions would therefore be inaccurate.
  • The source mix included locally detected interaction incidents and prompt-safe leads from Moltbook, arXiv, GitHub, and technology-news sources.

Candidate-construction failure analysis

Completed
  • Thirty-one frozen candidate packets had candidate_failed status after all three internal build or repair attempts.
  • Seven candidates initially failed only the Git whitespace check while their syntax and test collection checks passed. Their subsequent repair attempts commonly tried to recreate an already-created test file and were rejected by the patch safety layer.
  • Across the failed packets, repair attempts frequently could not locate an unambiguous edit target or attempted to create a target that already existed. Other candidates had real syntax, import, collection, scope, or test failures.
  • The queue flattened the detailed frozen-packet evidence into candidate_failed for thirty-one outcomes, making the Observatory much less diagnostic than the underlying artifacts.

Independent evaluation outcomes

Completed
  • Five candidates reached candidate_ready and completed public and held-out evaluation.
  • All five failed the policy's minimum score-improvement and confidence-separation requirements.
  • Several also exceeded the latency budget or regressed an evaluated category, so these five rejections should not be bypassed or relabeled as pipeline errors.
  • None of the five was eligible for automatic Stage 5 integration, so no eight-hour canary began.

Event-time interpretation

Diagnosed
  • Queue transition timestamps reuse the scheduler cycle's captured time even when candidate work takes several minutes.
  • This makes investigation, candidate construction, and rejection appear nearly simultaneous in the queue event stream even though the evaluation ledger and frozen packet times show the work continuing normally.
  • The timestamp behavior is an observability defect, but it is not evidence of an asynchronous evaluation race.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Production queue ledger reconciliation passed All thirty-six overnight candidate_rejected events referenced an existing frozen candidate packet; thirty-one packets were candidate_failed and five were candidate_ready but ineligible after evaluation.
Stage 6 event audit passed No overnight canary, governor, fast-forward, promotion, or implementation event was present for the thirty-six rejected workflows.
Construction-attempt audit passed Each of the thirty-one construction failures contained three recorded build or repair attempts with frozen validation evidence.
Journal tests passed npm run test:hiro passed the timestamped-entry unit tests.
Journal production build passed npm run build generated and validated 120 Hiro pages, then completed the TypeScript and Vite production build.

Current state

Next steps