Hiro development journal

Reviewing Contrastive Language Models as a local Hiro scorer

Source review and deferred record; no integration or experiment Machine-readable JSON

Executive summary

Reviewed CLM as a possible local ranking/decision component alongside the deferred Jev idea. It is a candidate scorer, not a replacement for Hiro's generative reasoning model or external discovery.

The source documentation supports an exploratory evaluation, but does not establish GGUF equivalence, Hiro-specific calibration, or a production advantage.

Work completed

Primary-source and existing-state review

Completed
  • Read https://github.com/Contrastive-LM/CLM, src/clm/embedder.py, src/clm/engine.py and https://huggingface.co/Contrastive-LM/CLM-v0.1-8B after following the supplied Reddit discussion.
  • Reference heads use a frozen Qwen3-8B encoder with last-token pooling; the small head download does not include the encoder.
  • Current serving documentation uses vLLM. The embedding client accepts an OpenAI-compatible endpoint, but that does not prove quantized llama.cpp embeddings preserve trained-head rankings or calibration.
  • Inspected existing Hiro opportunity priority, continuous idea scoring and bounded external-lead selection. A learned scorer could be compared at these existing boundaries, without creating another research queue.
  • Checked open handoffs; the existing ranked-autonomy handoff was not executed.

Deferred catalog follow-up

Completed
  • At the user's request, added deferred-clm-local-decision-scoring to the existing Research Map.
  • Retained the supplied Reddit URL, upstream repository/model card/code links, review details, limitations, related Jev identity and reconsideration conditions.
  • Updated the deferred-ideas navigation guide. No runtime or authority changes.

Decisions and reasoning

Validation and evidence

CheckStatusResult
Documentation and code review completed Read upstream implementation/model card and selected Hiro source functions; no runtime reproduction.
Published benchmarks not_reproduced Reported agentic verifier scores rely on fine-tuned heads and small held-out task sets; they are not zero-shot Hiro results. Published latency hardware differs from the local RTX 5090.
GGUF compatibility and local resources not_run No model downloaded, runtime installed, embedding parity tested, or coexistence memory measured.
Deferred record passed Source schemas, unique identities, source references and existing Markdown export inclusion validated.

Current state

Next steps