Hiro development journal

Capability-driven improvement fallback: implementation and qualification

Candidate rejected by final regression gate; qualified extension remains active Machine-readable JSON

Executive summary

Implemented a capability-driven fallback inside the existing ranked queue, reusing the queue database and canonical construction, testing and promotion controller. The central model remains frozen.

A real isolated run established exhaustion, measured a response-preservation weakness, admitted a ranked candidate and invoked local Qwen through the normal builder. Its first construction attempt failed; no autonomous gain or promotion is claimed.

The qualified extension is now active in production at 387a4b0. The restarted scheduler independently established exhaustion, measured the capability and queued a new candidate. No model-authored improvement has been promoted yet.

The production model-authored candidate now passes all 18 unseen validation cases, up from three at baseline, with all six anchors preserved. It has entered canonical canary monitoring; final holdout and promotion remain pending.

Final private holdout also passed18/18 versus baseline3/18, with six anchors preserved. All scheduled canary checkpoints passed; the canonical governor is evaluating the prepared promotion transaction.

The candidate improved validation and holdout but failed one existing brevity regression in the final repository gate. Promotion was rejected and the candidate retired. The capability extension remains active; successful autonomous implementation has not yet been demonstrated.

Work completed

Architecture and provenance

Implemented
  • Reconstructed scheduler, source discovery, queue eligibility, builder, evaluator, governor, activation and probation from source, history and runtime evidence. Existing dashboard changes were preserved byte for byte in a separate local commit.
  • Capability Mode requires recent successful discovery, no viable reservoir entry, resolved prerequisites, a qualified clean revision and no active promotion or infrastructure blockage.
  • Durable capability records and immutable receipts use existing SQLite runtime state and events. Only TRAIN examples reach candidate authoring. Validation and one final holdout remain host verified.
  • Version one abstracts historical response failures into domain-independent preservation of valid answers. It measures the production response boundary, not general reasoning ability.

Benchmark and controller integration

Implemented and tested
  • Generated 18 TRAIN, 18 VALIDATION and 18 HOLDOUT cases with separated domains and task identities, six dimensions, a reproducible seed and provenance. Six fixed human-controlled anchors supplement existing regression and external heldout suites.
  • Baseline measurement precedes candidate admission. Improvement requires perfect unseen preservation, strict gain over baseline and unchanged anchor performance. Existing independent tests, canary, promotion and probation remain authoritative.
  • Fixed a real Windows command-length failure in isolated input staging while preserving UTF-8 bytes and inherited runtime context.
  • Corrected an existing action oracle that classified an explicit no-action statement as an action claim, with positive and negative controls. Old observations are reconciled against current replay evidence or versioned oracle changes.
  • First live authoring repeatedly proposed empty replacement text. Added guidance for legal block removal without relaxing editor restrictions. Expired capability baselines are now retired before retry cooldowns or investigation reuse.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Exact release repository suite passed 387a4b0: 1109 passed, 4 skipped, 18 warnings in 855.18 seconds. Skips cover two opt-in live tests, a Linux descriptor-limit test and an absent memory candidate; real local-model construction was separately exercised through the canonical controller.
Builder and identity checks passed 145 passed, 3 warnings in 38.13 seconds.
Queue retry reconciliation passed 100 passed in 11.58 seconds, including expired baselines during future cooldowns and preservation of implemented records.
Exact release frozen pipeline qualification passed All five scenario groups passed at 387a4b0, with checksummed qualification packet recorded.
Real isolated capability baseline measured WSL/Bubblewrap executed all 54 generated cases plus six anchors: TRAIN 3/18, VALIDATION 3/18, HOLDOUT 3/18, anchors 6/6.
Real canonical construction failed At 4f2dabf, local Qwen completed an initial authoring attempt plus four internal repairs; each failed the existing nonempty replacement constraint. No candidate patch passed testing, validation or promotion.
Second real local-model construction passed Qwen generated a valid patch without repairs. The same frozen harness-owned fixture failed on the baseline and passed on the candidate. Independent broader evaluation remains underway.
Production activation passed Existing launch helper restarted Hiro; health confirmed loaded revision and checkout both 387a4b0, model connected, checkout clean. Ten pre-existing dashboard files were preserved byte for byte.
Production capability validation passed Candidate007ffd18:18/18 unseen VALIDATION versus baseline3/18; anchors6/6. Initial canary checkpoint passed, including25 independent regression tests and fresh local-model checks.
Isolated second run completion infrastructure_blocked Construction succeeded, but later worktree creation hit a Windows path-length limit. This isolated result is not a candidate-quality rejection. The shorter production checkout independently progressed into canary.
Final capability holdout and canary passed Candidate007ffd185576: final HOLDOUT18/18 versus baseline3/18; anchors6/6. All existing canary checkpoints0,5,15,60minutes passed.
Final governor repository suite failed 1 failed,1102 passed,12 skipped,20 warnings in465.00seconds. The candidate violated an existing150-word seat-comparison contract.
Targeted baseline/candidate reproduction confirmed_regression Unchanged production:2passed in0.43seconds. Candidate:1failed,1passed in0.65seconds. Both reported2marker warnings. The failure is excessive_content, not an environment problem.

Current state

Next steps