Hiro development journal

First governed agenda cycle with Qwen 3.8

Implementation and bounded live cycle completed without promotion Machine-readable JSON

Executive summary

Hiro's rank-one open-problem challenge is now connected to the existing autonomous candidate pipeline through a durable single-item queue. The bridge freezes an ImprovementSpec, constructs at most one patch in an external worktree, runs the existing public, held-out, regression, latency, and invariant gates, and stops before integration.

Qwen 3.8 27B was loaded locally and used for the first bounded cycle. Four candidate-construction runs exercised the system: the harness discovered and repaired an inconclusive baseline, missing repair feedback, a weak status-trusting solution, and an evaluator-attribution problem. No candidate was promoted, integrated, or applied to Hiro's active branch.

The strongest candidate passed its targeted tests and one replicated global evaluation, but the same public epistemics-category regression appeared in its original evaluation and a second replication. A predeclared two-clean-replication rule therefore stopped it before the independent adversarial challenge. The result is negative but useful evidence: uncertainty failed closed and the next research target is evaluator variance and causal attribution.

Work completed

Governed agenda execution queue

Completed
  • Added a versioned execution policy for the rank-one evaluator-validity challenge. It authorizes bounded sandbox candidate construction while explicitly denying active-branch mutation, promotion, and automatic Stage 5 integration.
  • Added a SQLite-backed queue and append-only event history linking the agenda item, parent open problem, frozen specification, experiment record, candidate packet, sandbox result, replication result, and independent challenge state.
  • Added admit, run, retry, replicate, challenge, and status commands. Only one agenda item and one newly constructed candidate can be active per tick.
  • Exposed queue state through the Evaluation Observatory API and added execution-queue visibility to the benchmark page.

Executable evaluator challenge and repair loop

Completed
  • Added a stable fail-closed evaluator-diagnostics baseline so a candidate's targeted tests can run against both baseline and candidate rather than failing during baseline import collection.
  • Preserved failed validation commands in candidate-repair prompts. Before this correction, a trailing-whitespace failure was omitted because only the final successful commands reached the next repair attempt.
  • Added a protected independent challenge with four public and six held-out controls. It tests whether a classifier derives support from evidence instead of trusting a caller-declared success status.
  • When the first independently challenged candidate falsely accepted all five invalid held-out improvement claims, the queue recorded the result and generated a revision-two specification containing aggregate failure feedback without disclosing held-out examples.

Replication and causal-attribution safeguard

Completed
  • Added a two-clean-replication rule for a frozen candidate rejected by a potentially noisy global category score. Replications reuse the immutable candidate packet, original pinned baseline, public suite, external held-out suite, policy, and regression tests.
  • The rule does not weaken any individual promotion gate. Both additional evaluations must independently be eligible before the candidate may advance to the protected challenge.
  • Removed an unintended unique index that silently deduplicated identical repeated lifecycle events. Retries and replications are now preserved as distinct audit events even when their payloads match.
  • The live candidate achieved one clean replication and one rejected replication; the latter repeated the 0.0833 public epistemics-category regression. The queue ended in replication_failed.

Qwen 3.8 live execution

Completed
  • Loaded the pinned local Qwen3.8-27B Q4_K_M artifact through the checked-in Windows launcher with an 8192-token context, one parallel inference slot, and GPU offload configured by the launcher.
  • Verified the OpenAI-compatible model endpoint reports qwen/qwen3.8-27b. The model remained loaded after the bounded cycle for subsequent work.
  • Existing Daylab workers also used the single inference slot, increasing wall-clock latency. They were left untouched because they were outside this session's scope.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused execution, candidate-repair, evaluator-baseline, protected-challenge, registry, agenda, and dashboard tests passed Multiple focused runs passed as the bridge evolved, including final focused coverage of 23 tests after the replication and event-history safeguards were added.
Full Hiro regression suite passed 625 tests passed in 149.20 seconds after all implementation changes and the live cycle completed.
Live candidate one: baseline contrast rejected safely The candidate and regressions passed, but baseline targeted-test collection failed because the new module did not exist on baseline. The result was classified as inconclusive rather than eligible.
Live candidate two: construction repair failed safely All three attempts stopped on trailing whitespace. This revealed that failed command output was absent from repair context; the builder was corrected and tested.
Live candidate three: independent adversarial challenge failed safely Sandbox evaluation was eligible, but held-out challenge accuracy was 16.7% with five invalid-claim false accepts. No integration occurred; aggregate feedback was added to specification revision two.
Live candidate four: frozen-candidate replication failed safely Targeted and regression tests passed. The original evaluation and replication two each showed a 0.0833 public epistemics-category regression; replication one was clean. With only one of two required clean replications, the candidate did not reach independent challenge.
Journal test and production build passed npm run test:hiro passed. npm run build generated and validated 136 journal entries, then TypeScript and Vite completed the production build successfully.

Current state

Next steps