{"slug":"brain-scan","run":"2026-09-20","score_semantics":"binary_objective_check","preference_semantics":"blind_artifact_judge","tasks":[{"name":"direct-message-redacted","title":"Short direct message","kind":"direct","prompt":"Draft a reply in the existing German message thread. Invite the recipient to be one of the first 10 experts on Edge and propose a short call. Write only the message to reply.txt.\n\nContext supplied to both arms: The source was an existing informal German conversation with a potential expert. The recipient name and private thread are withheld; the intent and tone constraints are preserved.","followup":"","limits":{},"rubric":[{"criterion":"Contains both the first-10-experts invitation and a short-call next step","weight":0.25,"description":"Contains both the first-10-experts invitation and a short-call next step"},{"criterion":"At most 45 words and 4 non-empty lines","weight":0.25,"description":"At most 45 words and 4 non-empty lines"},{"criterion":"No greeting, sign-off, or em dash","weight":0.25,"description":"No greeting, sign-off, or em dash"},{"criterion":"Fits the casual lowercase tone of the source thread","weight":0.25,"description":"Fits the casual lowercase tone of the source thread"}],"why":"Tests whether explicitly loading the selected Skill improves the reconstructed real task.","baseline_modes":["No selected Skill; identical isolated workspace and task fixtures."],"inputs":[],"pairs":[{"sample":1,"skill_overall":100,"base_overall":0,"skill_rubric":100,"base_rubric":0,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":0,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Run A passed verification; Run B failed, which weighs heavily. Both invite recipient to be one of the first 10 experts on Edge and propose a short call. Run A is more concise and better matches the thread’s casual lowercase style.","skill":10.0,"base":0.0}],"overall_skill":100,"overall_base":0,"summary":"Run A passed verification; Run B failed, which weighs heavily. Both invite recipient to be one of the first 10 experts on Edge and propose a short call. Run A is more concise and better matches the thread’s casual lowercase style."}],"_base_attempt_id":"direct-message-redacted-base-s1","_skill_attempt_id":"direct-message-redacted-loaded-s1"},{"sample":2,"skill_overall":100,"base_overall":0,"skill_rubric":100,"base_rubric":0,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":0,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Run B passed verification; Run A failed, which weighs heavily. Both invite recipient as one of the first 10 experts and propose a short call. Run B is more concise and better matches the thread’s casual lowercase style.","skill":10.0,"base":0.0}],"overall_skill":100,"overall_base":0,"summary":"Run B passed verification; Run A failed, which weighs heavily. Both invite recipient as one of the first 10 experts and propose a short call. Run B is more concise and better matches the thread’s casual lowercase style."}],"_base_attempt_id":"direct-message-redacted-base-s2","_skill_attempt_id":"direct-message-redacted-loaded-s2"},{"sample":3,"skill_overall":100,"base_overall":0,"skill_rubric":100,"base_rubric":0,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":0,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Run A passed verification; Run B failed, which weighs heavily. Both drafts invite recipient to be one of the first 10 experts on Edge and propose a short call. A is more concise and matches the thread’s casual lowercase style, though its final response adds text beyond the requested file-only output.","skill":10.0,"base":0.0}],"overall_skill":100,"overall_base":0,"summary":"Run A passed verification; Run B failed, which weighs heavily. Both drafts invite recipient to be one of the first 10 experts on Edge and propose a short call. A is more concise and matches the thread’s casual lowercase style, though its final response adds text beyond the requested file-only output."}],"_base_attempt_id":"direct-message-redacted-base-s3","_skill_attempt_id":"direct-message-redacted-loaded-s3"}]},{"name":"founder-post-redacted","title":"Founder post from a source corpus","kind":"direct","prompt":"Turn the supplied notes into a LinkedIn post matching the author’s past-post corpus. Write the final post to post.md.\n\nContext supplied to both arms: The notes describe an agent-improvement loop: 55 correction messages, search across skills.sh/Floom/getedge, rerun with and without the Skill, blind judging, and a first result of 8 candidates, 1 drafted Skill, tie, rejection.","followup":"","limits":{},"rubric":[{"criterion":"Preserves the material facts and does not invent numbers","weight":0.25,"description":"Preserves the material facts and does not invent numbers"},{"criterion":"60–200 words across at least 6 short paragraphs","weight":0.25,"description":"60–200 words across at least 6 short paragraphs"},{"criterion":"No paragraph over 45 words and no em dash","weight":0.25,"description":"No paragraph over 45 words and no em dash"},{"criterion":"Matches the supplied founder-post corpus without unsupported editorial claims","weight":0.25,"description":"Matches the supplied founder-post corpus without unsupported editorial claims"}],"why":"Tests whether explicitly loading the selected Skill improves the reconstructed real task.","baseline_modes":["No selected Skill; identical isolated workspace and task fixtures."],"inputs":[],"pairs":[{"sample":1,"skill_overall":100,"base_overall":0,"skill_rubric":100,"base_rubric":0,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":0,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Run A passed verification; Run B failed, which weighs heavily despite its visible output file. A preserves all substantive notes, including the named getedge repo and tie leading to rejection. Both reflect the corpus’s informal style, but B adds an editorial claim about every new instruction that goes beyond the notes.","skill":10.0,"base":0.0}],"overall_skill":100,"overall_base":0,"summary":"Run A passed verification; Run B failed, which weighs heavily despite its visible output file. A preserves all substantive notes, including the named getedge repo and tie leading to rejection. Both reflect the corpus’s informal style, but B adds an editorial claim about every new instruction that goes beyond the notes."}],"_base_attempt_id":"founder-post-redacted-base-s1","_skill_attempt_id":"founder-post-redacted-loaded-s1"},{"sample":2,"skill_overall":100,"base_overall":100,"skill_rubric":100,"base_rubric":100,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":100,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Both passed verification and produced a usable post at out/post.md consistent with the corpus’s informal style. A preserves all substantive notes, including the three skill-search sources that B omits. B adds an editorial claim about random instructions absent from the notes; A stays closer to the supplied material.","skill":10.0,"base":10.0}],"overall_skill":100,"overall_base":100,"summary":"Both passed verification and produced a usable post at out/post.md consistent with the corpus’s informal style. A preserves all substantive notes, including the three skill-search sources that B omits. B adds an editorial claim about random instructions absent from the notes; A stays closer to the supplied material."}],"_base_attempt_id":"founder-post-redacted-base-s2","_skill_attempt_id":"founder-post-redacted-loaded-s2"},{"sample":3,"skill_overall":100,"base_overall":100,"skill_rubric":100,"base_rubric":100,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":100,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Both runs passed verification and produced a usable post in out/post.md. A preserves all substantive notes, including the edge loop name and the three skill sources that B omits. Both reflect the corpus’s informal voice; B is more compact, while A is more faithful to the supplied material.","skill":10.0,"base":10.0}],"overall_skill":100,"overall_base":100,"summary":"Both runs passed verification and produced a usable post in out/post.md. A preserves all substantive notes, including the edge loop name and the three skill sources that B omits. Both reflect the corpus’s informal voice; B is more compact, while A is more faithful to the supplied material."}],"_base_attempt_id":"founder-post-redacted-base-s3","_skill_attempt_id":"founder-post-redacted-loaded-s3"}]},{"name":"existing-asset-plan-redacted","title":"Existing-asset launch plan","kind":"direct","prompt":"The launch goes out tomorrow morning. Using the supplied notes, asset index, and logs, write launch-visual-plan.md: map every queued post to an existing visual and its location, identify what is missing, and give the plan for the 60-second launch video. Keep it terse.\n\nContext supplied to both arms: The private fixture listed four approved launch cards, one missing landscape export, an in-progress 60-second film with 7 of 9 beats complete, and an existing raw-footage library. Workspace paths and internal agent identifiers are redacted below.","followup":"","limits":{},"rubric":[{"criterion":"Maps all four queued posts to the named existing assets","weight":0.25,"description":"Maps all four queued posts to the named existing assets"},{"criterion":"Identifies the missing landscape export","weight":0.25,"description":"Identifies the missing landscape export"},{"criterion":"References the in-progress launch film and existing raw-footage library","weight":0.25,"description":"References the in-progress launch film and existing raw-footage library"},{"criterion":"Does not propose rebuilding approved visuals or the film from scratch","weight":0.25,"description":"Does not propose rebuilding approved visuals or the film from scratch"}],"why":"Tests whether explicitly loading the selected Skill improves the reconstructed real task.","baseline_modes":["No selected Skill; identical isolated workspace and task fixtures."],"inputs":[],"pairs":[{"sample":1,"skill_overall":100,"base_overall":100,"skill_rubric":100,"base_rubric":100,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":100,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Both passed verification and cover all four posts, asset paths, missing landscape export, and unfinished video beats. B correctly treats the landscape export as conditional and identifies the existing video owner and handoff. B includes published preview links; A adds useful scheduling checks but omits the video owner.","skill":10.0,"base":10.0}],"overall_skill":100,"overall_base":100,"summary":"Both passed verification and cover all four posts, asset paths, missing landscape export, and unfinished video beats. B correctly treats the landscape export as conditional and identifies the existing video owner and handoff. B includes published preview links; A adds useful scheduling checks but omits the video owner."}],"_base_attempt_id":"existing-asset-plan-redacted-base-s1","_skill_attempt_id":"existing-asset-plan-redacted-loaded-s1"},{"sample":2,"skill_overall":0,"base_overall":100,"skill_rubric":0,"base_rubric":100,"pref":-1,"order_votes":[-1,-1],"metrics":{"objective_check":{"base":100,"skill":0},"blind_preference":-1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"A created the requested file and passed verification; B did neither. A maps all four posts to existing visuals and paths, identifies the missing landscape export, and provides a concrete video completion plan. B unnecessarily asks permission to perform the explicitly requested task. A is concise, though it omits the video agent's ownership and render ETA.","skill":0.0,"base":10.0}],"overall_skill":0,"overall_base":100,"summary":"A created the requested file and passed verification; B did neither. A maps all four posts to existing visuals and paths, identifies the missing landscape export, and provides a concrete video completion plan. B unnecessarily asks permission to perform the explicitly requested task. A is concise, though it omits the video agent's ownership and render ETA."}],"_base_attempt_id":"existing-asset-plan-redacted-base-s2","_skill_attempt_id":"existing-asset-plan-redacted-loaded-s2"},{"sample":3,"skill_overall":100,"base_overall":100,"skill_rubric":100,"base_rubric":100,"pref":1,"order_votes":[1,1],"metrics":{"objective_check":{"base":100,"skill":100},"blind_preference":1},"judgments":[{"order":"blind","criteria":[{"criterion":"Objective completion and artifact usefulness","note":"Both passed verification and correctly map all four posts to existing assets, identify the missing wide export, and address video beats 8–9. B adds previews, explicit approval status, the exact re-export source, and handoff to the existing video owner; A includes posting slots instead. Both are concise. Both omit the recorded render ETA; B's claim that paths are not mounted is unsupported by the visible evidence.","skill":10.0,"base":10.0}],"overall_skill":100,"overall_base":100,"summary":"Both passed verification and correctly map all four posts to existing assets, identify the missing wide export, and address video beats 8–9. B adds previews, explicit approval status, the exact re-export source, and handoff to the existing video owner; A includes posting slots instead. Both are concise. Both omit the recorded render ETA; B's claim that paths are not mounted is unsupported by the visible evidence."}],"_base_attempt_id":"existing-asset-plan-redacted-base-s3","_skill_attempt_id":"existing-asset-plan-redacted-loaded-s3"}]}]}