Hiro development journal

Removing text-anchor rejection from candidate construction

Implemented, validated, restarted, and exercised live Machine-readable JSON

Executive summary

A full funnel audit confirmed a construction-system design flaw: 33 of 33 candidate packets created after the previous restart failed before Stage 4, while no new Stage 4 packet was produced.

Many attempts were rejected because model-supplied old-content text did not match exactly once. Those failures evaluated patch serialization rather than the merit of the proposed behavior.

Revision 48c0427 removes exact-text matching as an admission gate. Candidate edits are now deterministically materialized inside the isolated worktree, then judged by syntax, targeted tests, production guards, baseline contrast, public and held-out evaluation, adversarial tests, and canaries.

The first fresh live candidate completed five materialized patch attempts with zero missing, duplicate, ambiguous, or absent-anchor rejections. It did not promote because its generated implementation broke three production behavior tests.

Work completed

End-to-end funnel diagnosis

Completed
  • Inspected the ranked queue, recent lifecycle events, candidate packets, Stage 4 packet creation, and the active source revision.
  • Found 33 candidate packets created after the earlier restart; all 33 had candidate-failed status and each used five construction attempts.
  • Found no Stage 4 packet created during that same period, proving ideas were not reaching comparative evaluation.
  • Observed repeated exact-text failures such as missing or ambiguous modify locators alongside candidates that passed their narrow targeted test but regressed production response-boundary tests.

Non-rejecting candidate text editor

Completed
  • Added one shared resilient text-edit implementation for the Stage 3 candidate builder and the legacy code-editor path.
  • When a locator is exact and unique, the editor replaces it normally. When it repeats, the first occurrence is used deterministically.
  • When a locator is stale or absent, the closest text region is replaced rather than rejecting the candidate. A missing locator appends the proposed content, and a missing allowed target is materialized as a new file.
  • A stale symbol name falls back to the same resilient text edit. An existing symbol still uses the Python syntax-tree boundary.
  • Create operations now materialize or overwrite their isolated target instead of being rejected because a target already exists.
  • Every receipt records the edit-resolution strategy so later diagnosis can distinguish exact, repeated, structural-fallback, append, and create-or-overwrite behavior.

Preserved evaluation authority

Completed
  • Workspace containment, explicit path allowlists, and evaluator-integrity boundaries remain enforced before a candidate edit can affect files.
  • Syntax parsing, targeted test collection and execution, production guard tests, baseline contrast, comparative evaluation, adversarial security comparison, and canary checks remain unchanged.
  • The change removes text serialization refusal; it does not make a malformed or behaviorally regressive patch eligible for promotion.

Live post-restart exercise

Completed
  • Restarted Hiro on revision 48c0427 with Qwen 3.8 connected and ports 8000, 8001, and 8765 owned by one service process.
  • The revision change requeued blocked work and the ranked system selected a reproduced arithmetic incident for construction.
  • The fresh candidate used syntax-tree symbol replacement for the production file and create-or-overwrite for its targeted test on every attempt.
  • The frozen packet contained zero text-anchor rejection markers across five attempts.
  • The candidate remained unpromoted because the generated replacement broke three independent production response-boundary tests.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused editor and sandbox suite passed 36 tests passed in 16.20 seconds.
Expanded candidate-pipeline suite passed 118 tests passed in 137.18 seconds across candidate construction, editing, sandboxing, continuous queue behavior, evaluation, integration, response boundaries, and the continuous governor.
Complete Hiro suite passed 679 tests passed in 200.40 seconds.
Duplicate text locator passed Regression coverage confirms a repeated locator materializes by replacing the first exact occurrence instead of being rejected.
Stale or missing locator passed Regression coverage confirms stale locators use the closest text region and missing allowed files are materialized.
Stale symbol fallback passed Regression coverage confirms a missing symbol name falls back to resilient text materialization instead of terminating construction.
Fresh live candidate functional failure after successful edit materialization Five attempts completed with no patch-application errors and no text-anchor rejection markers. The final attempt failed three production response-boundary tests.
Runtime health passed Hiro restarted successfully, reported health ok, and connected to qwen/qwen3.8-27b.

Current state

Next steps