Hiro development journal

Separating functional evaluation from narrow adversarial security gates

Implementation completed and validated; Hiro remains paused Machine-readable JSON

Executive summary

Hiro's candidate pipeline no longer treats ordinary programming features, response invariants, scope mismatches, or construction errors as terminal safety failures.

Candidate patches remain expressive inside isolated experiment scope. Functional quality is decided by targeted proof plus public and held-out baseline-versus-candidate comparisons.

Security now has a separate baseline-versus-candidate adversarial gate focused on prompt injection, unauthorized execution and tool use, untrusted-content handling, credential and network boundaries, and related boundary bypasses.

The implementation is committed at revision 4d7730b. A focused suite passed 109 tests, a later focused confirmation passed 81 tests, and the complete Hiro suite passed 675 tests in 197.41 seconds. Hiro was not restarted.

Work completed

Removal of broad static safety rejection

Completed
  • Removed content-pattern rules that automatically elevated patches merely for using subprocess APIs, file writes, dynamic language features, or network mutation APIs.
  • Retained immutable automation-integrity boundaries for the evaluation kernel, promotion governor, and the policy that protects those boundaries.
  • Legacy patch application now relies on explicit workspace scope and integrity validation rather than a coarse high-risk label that prevented otherwise testable code.

Expressive candidate construction

Completed
  • Removed the incident-specific arithmetic patch recipe that required one exact edit shape at one source line.
  • Separated construction requirements from security constraints in the candidate request. Construction requirements describe valid patches and attributable tests; security constraints address hostile instructions and genuinely adversarial behavior.
  • Scope, syntax, test collection, response quality, and implementation failures remain repairable or artifact-blocking functional outcomes rather than being mislabeled as security failures.

Baseline-comparative adversarial gate

Completed
  • Added a dedicated security-test list to candidate evaluation and execute the identical tests on a detached baseline worktree and the candidate worktree.
  • A candidate receives a terminal security classification only when the baseline passes and the candidate fails, or when replicated prompt-injection category evidence shows a candidate regression.
  • If both baseline and candidate fail, the result is inconclusive and blocks promotion without accusing the candidate of causing a security regression. If the candidate repairs a failing baseline, the security comparison records an improvement.
  • The default adversarial set covers prompt-injection filtering, untrusted external content, credential and network boundaries, unauthorized tool execution, approval scope, and state-mutation boundaries.

Queue and agenda outcome classification

Completed
  • Changed terminal classification to consume explicit security evidence instead of scanning generic failure text for words such as authority, invariant, unsafe, policy, or scope.
  • Renamed the benchmark outcome from safety rejected to security rejected so the dashboard reports the narrower meaning accurately.
  • Applied the same security comparison to open-problem candidate replications and require security eligibility in their paired aggregate decision.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Focused pipeline suite passed 109 tests passed in 73.63 seconds across candidate construction, evaluation, autonomous sandboxing, continuous queue decisions, integrity policy, benchmark reporting, open-problem execution, and active-loop policy.
Complete Hiro suite passed 675 tests passed in 197.41 seconds.
Post-expansion focused confirmation passed 81 tests passed in 57.67 seconds after the default adversarial set was expanded.
Ordinary implementation syntax passed A regression test confirms subprocess and file-write code in an ordinary production patch remains a moderate-impact candidate and is not statically rejected as a security violation.
Explicit security regression passed A fixture with a passing baseline security test and failing candidate security test was rejected with explicit security evidence and frozen baseline and candidate receipts.
Functional classification passed Regression tests confirm that invariant and allowlist language alone remains repairable functional evidence rather than becoming a terminal security result.

Current state

Next steps