Evaluation receipt / repo-to-launch

Repo to Launch

Both arms passed every objective check, but all three blind comparisons tied.

Inconclusive
3Valid pairs
3Blind ties
3/3Treatment verifier passes
3/3Baseline verifier passes

What this proves

Produced grounded launch packs safely, without proving a quality improvement.

Status

inconclusive. The skill caused no factuality regression, but demonstrated no lift over baseline.

Method

Three blind paired Harbor runs on one local-first unsigned beta repository.

Model

openai/gpt-5.6-sol

Limitations
  • One representative repository
  • One model
  • Blind judge did not receive full generated file contents
Exact packagesha256:2543926cc0dad9180a219cad5bd64229632c39bdb70f8092c2bf308617443e8e