{
  "schemaVersion": 2,
  "date": "2026.08.16",
  "publishedAt": "2026-08-16T21:54:27-07:00",
  "timeZone": "America/Los_Angeles",
  "title": "Approval-system audit finds idea/artifact conflation and invalid contrast evidence",
  "publicationStatus": "Diagnostic completed; live system observed but not modified",
  "executiveSummary": [
    "A forensic review of the live continuous-improvement queue confirms that the extreme rejection rate is substantially caused by approval-system design, not evidence that nearly every underlying idea is harmful.",
    "The durable queue currently contains 173 rejected and 3 implemented ideas, a 98.3 percent rejection share among terminal rejected-or-implemented records. Since recovery event 1395, the latest classified outcomes for 34 ideas comprise 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass.",
    "The engine currently converts a repairable candidate-construction failure into a terminal idea rejection after the bounded attempt budget is exhausted. This conflates failure to manufacture an evaluable patch/test artifact with evidence that the idea itself lacks merit.",
    "The single recent candidate-gate pass exposed the opposite defect. Its untouched-baseline run failed with TypeError because the candidate-authored test called the baseline chat function with a keyword its baseline signature did not accept. Pytest represented this call-phase exception as a JUnit failure, and Hiro's count-only parser classified any such failure as an attributable assertion contrast.",
    "The live canary cannot correct that false positive because its default probe simply reruns the same single candidate-authored test at synthetic checkpoints. The approval stages therefore reuse correlated evidence instead of adding an independent behavioral check."
  ],
  "workstreams": [
    {
      "title": "Live rejection-funnel audit",
      "status": "Completed",
      "details": [
        "Read the durable continuous-improvement SQLite state and append-only event history without mutating the queue.",
        "Observed 173 rejected, 18 superseded, 18 queued, 3 implemented, 1 candidate, and 1 canary record.",
        "Among terminal rejected-or-implemented records, 173 of 176 are rejected, or 98.3 percent.",
        "Classified the latest relevant outcomes for 34 ideas since recovery event 1395: 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass.",
        "The distribution shows that the dominant bottleneck is artifact construction and test manufacture, not the global evaluation policy deciding that otherwise valid changes are negative."
      ]
    },
    {
      "title": "Terminal-state semantics",
      "status": "Defect confirmed",
      "details": [
        "Traced continuous_engine.py candidate handling: non-safety candidate failures are retried only while attempts remain, then the record is unconditionally transitioned to state rejected with event candidate_rejected.",
        "This terminal path applies even when the candidate result identifies the failure as repairable and retryable.",
        "Common terminal reasons include empty or invalid patch plans, ambiguous edit anchors, candidate validation failure, and candidate-local regression failure.",
        "Those observations establish that rejected currently means either the proposed improvement failed or the builder failed to create a valid experiment. The dashboard and ranking system cannot distinguish the two."
      ]
    },
    {
      "title": "Targeted-contrast evidence audit",
      "status": "Defect confirmed",
      "details": [
        "Inspected the passing writing-assistance candidate sandbox-20260817043030-a2622db6-01 and its frozen recommendation packet.",
        "Its targeted baseline contrast returned pytest code 1 with one JUnit failure and no JUnit errors, which the current targeted_contrast_status function accepts as improvement_demonstrated.",
        "The captured traceback shows that baseline execution stopped at TypeError: chat() got an unexpected keyword argument 'message'. No behavioral assertion established that the baseline failed the writing-assistance contract.",
        "Pytest's JUnit implementation writes ordinary call-phase test failures through the failure element regardless of whether the exception is AssertionError or another exception. Hiro currently retains only aggregate tests, failures, errors, and skipped counts, discarding exception type and assertion provenance.",
        "The earlier assertion-aware gate therefore eliminated collection and setup errors but did not actually distinguish assertion failures from call-phase runtime exceptions."
      ]
    },
    {
      "title": "Canary independence audit",
      "status": "Weakness confirmed",
      "details": [
        "The active writing-assistance candidate entered canary after the invalid baseline contrast and had completed synthetic checkpoints at 0, 5, and 15 minutes at audit time.",
        "Each default canary probe invokes the candidate's own tests tuple. For this candidate that is one generated test, tests/autonomous/test_continuous_17afe3b06f2e.py.",
        "The three recorded probes therefore repeat the same test that helped admit the candidate; they are temporal repetitions, not independent measurements.",
        "The final 60-minute checkpoint remained pending during the audit. No live interaction, separately authored oracle, or category-peer replay had yet supplied independent canary evidence."
      ]
    }
  ],
  "decisions": [
    "Treat this as an approval-architecture defect rather than tuning a single acceptance threshold.",
    "Separate idea disposition from candidate-attempt disposition. Exhausted builder retries should produce an artifact-generation-blocked state while retaining the idea and its priority, not a negative verdict on the idea.",
    "Make artifact validity a prerequisite to evaluation. Construction, syntax, import, patch-application, and test-execution failures must not enter idea-quality acceptance statistics.",
    "Require assertion-provenance evidence for untouched-baseline contrast: call phase reached, exception type is AssertionError or an approved contract-assertion type, and the failure is tied to a harness-owned assertion identifier.",
    "Require the same baseline-compatible entrypoint and invocation shape on candidate and baseline; TypeError, ImportError, AttributeError, fixture failure, unexpected exception, skip, or collection failure is inconclusive and repairable, never positive evidence.",
    "Demote candidate-authored tests to construction evidence. Final approval should use a pre-existing or harness-owned contract replay whose oracle and invocation are fixed independently of the patch.",
    "Replace repeated same-test canary probes with independent contract replay, category-peer non-regression checks, and, where safe, separately sampled interaction evidence.",
    "Expose separate dashboard metrics for idea merit, builder yield, test validity, evaluator decision, and canary outcome so a 98 percent artifact failure rate cannot be presented as a 98 percent bad-idea rate.",
    "Do not mutate or stop the live process as part of this diagnostic-only session; operational intervention and code repair require a distinct authorized action."
  ],
  "validation": [
    {
      "check": "Durable queue state",
      "status": "passed",
      "result": "Read-only SQLite queries returned 173 rejected, 18 superseded, 18 queued, 3 implemented, 1 candidate, and 1 canary record. Terminal rejected-or-implemented rejection share is 98.3 percent."
    },
    {
      "check": "Recent funnel classification",
      "status": "passed",
      "result": "Read-only event classification since event 1395 found latest relevant outcomes for 34 ideas: 26 construction or local-test failures, 5 targeted-contrast failures, 2 global-evaluator failures, and 1 candidate-gate pass."
    },
    {
      "check": "False-negative control-flow trace",
      "status": "passed",
      "result": "Source inspection confirmed that after the bounded candidate-attempt limit, continuous_engine.py transitions candidate_failed results to rejected even when they are non-safety, repairable construction failures."
    },
    {
      "check": "False-positive evidence trace",
      "status": "passed",
      "result": "The frozen recommendation packet for sandbox-20260817043030-a2622db6-01 records a baseline TypeError caused by an unsupported keyword argument, while the aggregate JUnit counts satisfied the current improvement_demonstrated rule."
    },
    {
      "check": "Canary independence trace",
      "status": "passed",
      "result": "Source and durable canary state confirm that default probes rerun the candidate tests; the active candidate's 0, 5, and 15 minute checkpoints each reran the same one-test file and were marked synthetic."
    },
    {
      "check": "Hiro code changes",
      "status": "not run",
      "result": "No Hiro source, queue state, service process, candidate, or canary was modified during this diagnostic session."
    },
    {
      "check": "Journal test and production build",
      "status": "passed",
      "result": "npm run test:hiro passed. The first build attempt overlapped with a still-finishing generator process and encountered a transient Windows ENOTEMPTY race in the generated public/hiro directory; after confirming no process remained, npm run build generated and validated 141 journal entries and completed TypeScript and Vite production compilation successfully."
    }
  ],
  "currentState": [
    "The live continuous queue is operative and Qwen continues candidate work.",
    "A concept-explanation incident is in candidate construction while the writing-assistance candidate remains in canary.",
    "The writing-assistance canary had passed synthetic checkpoints 0, 5, and 15 by repeating its candidate-authored test; its 60-minute checkpoint was pending at the audit snapshot.",
    "The audit made no operational or source changes, so the identified false-negative and false-positive mechanisms remain active."
  ],
  "limitations": [
    "The recent-outcome classification is a deterministic reason-string audit of queue events, not a manual semantic review of all 173 rejected ideas.",
    "This diagnosis proves that many terminal rejections do not express idea merit; it does not prove that every rejected idea would produce a beneficial patch if construction succeeded.",
    "The passing candidate may contain useful code despite invalid admission evidence. The audit establishes that the current evidence cannot support approval, not that the patch is necessarily harmful.",
    "The canary was still active at the snapshot, so its final state may change after publication unless the system is paused or repaired.",
    "No corrective tests or implementation were added in this diagnostic-only session."
  ],
  "nextSteps": [
    "Pause or quarantine the active writing-assistance canary before automatic promotion because its causal contrast is invalid.",
    "Add an explicit artifact_attempt state model with construction_failed, test_invalid, evaluable, and evaluator_rejected outcomes; reserve idea rejected for valid negative evidence or a deliberate policy decision.",
    "Capture per-test phase, exception class, node identifier, and harness assertion identifier from pytest instead of relying on aggregate JUnit counts.",
    "Add regression cases proving TypeError and other unexpected call-phase exceptions are inconclusive while a known contract assertion failure is attributable.",
    "Introduce harness-owned, baseline-compatible test templates for recurring interaction contracts and freeze their invocation and oracle before candidate generation.",
    "Make canary evidence independent by running the originating harness replay, peer-category suite, and distinct sampled cases rather than the candidate's test alone.",
    "Reclassify historical rejection analytics into builder failure, invalid experiment, valid negative evaluation, safety rejection, and promotion failure so acceptance calibration uses the correct denominator.",
    "After repair, rerun representative rejected ideas and measure builder yield, valid-experiment yield, true evaluator acceptance, and post-canary retention separately."
  ]
}
