Executive summary
The new Open Problems Registry supplies durable structure, but it still needs a disciplined external discovery system that finds consequential problems rather than merely accumulating papers, issues, or benchmark scores.
No single public platform currently serves as a complete registry of open problems in self-improving AI harness construction. The strongest design is a composite: METR's portable task standard, ARC Prize's benchmark-and-open-challenge model, PaperBench's hierarchical rubrics, real task benchmarks such as SWE-bench, OSWorld, and MLE-bench, and issue ecosystems from active agent frameworks.
Hugging Face is valuable as a discovery, dataset, leaderboard, collection, and competition substrate. It should feed Hiro's candidate-problem inbox and later host bounded challenges when useful, but Hugging Face should not define the research agenda by popularity or leaderboard rank alone.
The recommended next layer is a Problem Discovery Observatory that separates raw leads from corroborated problems, clusters signals by underlying mechanism, requires reproducible evidence, and converts validated gaps into versioned challenge packets with public and held-out evaluation.