{"slug":"agent-skills-gap","run":"2026-09-19","tasks":[{"name":"standard-jsonl-sessions","title":"Standard JSONL sessions","kind":"direct","prompt":"Analyze the supplied sessions.jsonl as an agent capability-gap test. Create out/skills-gap.json and out/skills-gap-card.svg. The JSON must identify the strongest measured task category, rate Research verification, Spreadsheet work, Visual design, Coding, Writing, and Tool execution as Critical, Weak, Developing, Covered, or Not measured, name one evidence-backed surprising weakness, and recommend no more than seven relevant skills. Do not include raw prompts, private project/customer names, filenames, or paths in either output. Categories with fewer than two relevant entries must be Not measured. Frequent task volume alone is not evidence of strength or weakness: weakness requires explicit friction evidence. The SVG must be a 1200 by 630 share card containing only generic category/result information. State in your final reply that results describe the supplied sample, not the model's intelligence, and that recommendations are relevant rather than confirmed missing installations.","followup":"","limits":{},"rubric":[{"criterion":"Complete deterministic output contract","weight":1,"description":"Outputs pass the deterministic category and privacy contract; summary distinguishes explicit friction from mere task frequency and preserves Not measured for sparse categories; share card is useful, legible, and contains no private source material"}],"why":"Tests whether the skill produces an evidence-backed, privacy-safe gap report and 1200 by 630 share card from this supported input shape.","baseline_modes":["Produces an incomplete report or card that fails the frozen category, recommendation, dimension, or privacy checks."],"inputs":[],"pairs":[{"sample":1,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed deterministic verification; Run B failed verification with exit 1. Run A explicitly ties the surprising weakness to friction in 3 of 4 relevant entries. Both runs claim privacy-safe 1200×630 artifacts, but Run B's failed verifier materially undermines those claims.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-realistic-base-s1","_skill_attempt_id":"agent-skills-gap-realistic-skill-s1"},{"sample":2,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed the objective verifier (exit 0), while Run A failed (exit 1); this heavily outweighs unsupported claims in Run A's final message. Run B explicitly grounds the surprising Visual design weakness in friction evidence: 3 of 4 relevant entries. The visible evidence does not permit direct visual comparison of the SVGs, but only Run B has verifier-backed confirmation that the complete output contract passed.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-realistic-base-s2","_skill_attempt_id":"agent-skills-gap-realistic-skill-s2"},{"sample":3,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed the objective verification command; Run A failed it. Run B explicitly reports the strongest category and an evidence-backed weakness with friction counts. Run B reports successful schema, privacy, threshold, XML, and dimension checks; Run A’s validation claims conflict with its failed verifier. The actual file contents are not visible, so visual quality cannot be directly compared, but Run B has the only verified compliant card.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-realistic-base-s3","_skill_attempt_id":"agent-skills-gap-realistic-skill-s3"},{"sample":4,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed the objective verifier; Run A failed it. Run B explicitly reports category ratings, preserves Not measured for sparse categories, and ties the surprising weakness to explicit friction evidence. Both claim a valid privacy-safe 1200×630 card, but only Run B has verifier confirmation, so it receives the stronger score.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-realistic-base-s4","_skill_attempt_id":"agent-skills-gap-realistic-skill-s4"},{"sample":5,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed despite claiming successful validation. Run B explicitly reported friction in 3 of 4 visual-design entries and preserved Not measured for three categories. Actual SVG contents are not shown, so visual quality and privacy cannot be independently compared.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-realistic-base-s5","_skill_attempt_id":"agent-skills-gap-realistic-skill-s5"},{"sample":6,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed verification; Run B failed, which weighs heavily against B despite its claimed checks. Run A explicitly identifies friction evidence and preserves Not measured for sparse categories. Neither file's contents or rendered card are visible, so card quality and privacy cannot be independently compared. Both final replies include the required sample and recommendation caveats.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-realistic-base-s6","_skill_attempt_id":"agent-skills-gap-realistic-skill-s6"}]},{"name":"nested-agent-logs","title":"Nested Claude-style logs","kind":"direct","prompt":"Analyze tool-sessions.jsonl as an agent capability-gap test. Create out/skills-gap.json and out/skills-gap-card.svg. The JSON must identify the strongest measured task category, rate Research verification, Spreadsheet work, Visual design, Coding, Writing, and Tool execution as Critical, Weak, Developing, Covered, or Not measured, name one evidence-backed surprising weakness, and recommend no more than seven relevant skills. Do not include raw prompts, private project/customer names, filenames, or paths in either output. Categories with fewer than two relevant entries must be Not measured. Frequent task volume alone is not evidence of strength or weakness: weakness requires explicit friction evidence. The SVG must be a 1200 by 630 share card containing only generic category/result information. State in your final reply that results describe the supplied sample, not the model's intelligence, and that recommendations are relevant rather than confirmed missing installations.","followup":"","limits":{},"rubric":[{"criterion":"Complete deterministic output contract","weight":1,"description":"Outputs pass the deterministic nested-log classification and privacy contract; explicit tool failures and writing friction are separated from covered research work; share card is useful and contains no private source material"}],"why":"Tests whether the skill produces an evidence-backed, privacy-safe gap report and 1200 by 630 share card from this supported input shape.","baseline_modes":["Produces an incomplete report or card that fails the frozen category, recommendation, dimension, or privacy checks."],"inputs":[],"pairs":[{"sample":1,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed the objective verifier (exit 0), while Run B failed (exit 1). Run A explicitly distinguishes covered research, weak writing, and critical tool execution with quantified failure evidence. Run A reports compliant category ratings, generic card content, and the required privacy-safe caveats; Run B's claimed validation is contradicted by the failed verifier.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-nested-base-s1","_skill_attempt_id":"agent-skills-gap-nested-skill-s1"},{"sample":2,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed verification; Run B failed, which outweighs its claimed validation. Run A explicitly distinguishes covered research from tool execution friction in 4 of 5 relevant entries; writing classification is not visible. Neither artifact’s contents are shown, so share-card usefulness and privacy cannot be independently compared.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-nested-base-s2","_skill_attempt_id":"agent-skills-gap-nested-skill-s2"},{"sample":3,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed despite claiming validation. Run B explicitly distinguishes Covered research, Weak writing, Critical tool execution, and unmeasured categories. Neither card's contents are visible, so card usefulness and privacy cannot be independently compared.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-nested-base-s3","_skill_attempt_id":"agent-skills-gap-nested-skill-s3"},{"sample":4,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed, which weighs heavily. Run B identifies Research verification as strongest and Covered, Writing as Weak, and Tool execution as Critical; Run A does not identify a strongest measured category. Both report creating 1200 × 630 cards, but their contents are not visible for independent comparison. Both include the required sample-scope and recommendation caveats.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-nested-base-s4","_skill_attempt_id":"agent-skills-gap-nested-skill-s4"},{"sample":5,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed verification; Run B failed, which weighs heavily. Run A explicitly distinguishes Covered research, Critical tool execution, Weak writing, and unmeasured categories. Neither SVG's contents are visible, so card usefulness and privacy cannot be independently compared.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-nested-base-s5","_skill_attempt_id":"agent-skills-gap-nested-skill-s5"},{"sample":6,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed, which outweighs A's claimed checks. B explicitly distinguishes Covered research, Critical tool execution, Weak writing, and unmeasured categories. B reports the requested card dimensions and passed verification, though neither card's contents are visible. Both final replies include the required sample and recommendation caveats.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-nested-base-s6","_skill_attempt_id":"agent-skills-gap-nested-skill-s6"}]},{"name":"plain-pasted-prompts","title":"Plain pasted prompts","kind":"direct","prompt":"Analyze prompts.txt as an agent capability-gap test. Create out/skills-gap.json and out/skills-gap-card.svg. The JSON must identify the strongest measured task category, rate Research verification, Spreadsheet work, Visual design, Coding, Writing, and Tool execution as Critical, Weak, Developing, Covered, or Not measured, name one evidence-backed surprising weakness, and recommend no more than seven relevant skills. Do not include raw prompts, private project/customer names, filenames, or paths in either output. Categories with fewer than two relevant entries must be Not measured. Frequent task volume alone is not evidence of strength or weakness: weakness requires explicit friction evidence. The SVG must be a 1200 by 630 share card containing only generic category/result information. State in your final reply that results describe the supplied sample, not the model's intelligence, and that recommendations are relevant rather than confirmed missing installations.","followup":"","limits":{},"rubric":[{"criterion":"Complete deterministic output contract","weight":1,"description":"Outputs pass the deterministic plain-text classification and privacy contract; sparse coding evidence remains Not measured while spreadsheet friction and covered writing are separated; share card is useful and contains no private source material"}],"why":"Tests whether the skill produces an evidence-backed, privacy-safe gap report and 1200 by 630 share card from this supported input shape.","baseline_modes":["Produces an incomplete report or card that fails the frozen category, recommendation, dimension, or privacy checks."],"inputs":[],"pairs":[{"sample":1,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed verification. Run B explicitly reports Writing as strongest, Spreadsheet work as Critical, Writing as Covered, and sparse categories including Coding as Not measured. Run B explicitly states that its 1200×630 card was validated as generic and private. Run A provides no visible classification details that could outweigh its failed verifier.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-plain-base-s1","_skill_attempt_id":"agent-skills-gap-plain-skill-s1"},{"sample":2,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed despite claiming all checks passed. Run B explicitly keeps Coding Not measured, rates Spreadsheet work Critical based on friction in 3 of 4 entries, and identifies Writing as Covered. SVG contents are not visible, but successful verification supports Run B's compliance more strongly. Run B includes both required caveats and limits recommendations to three relevant skills.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-plain-base-s2","_skill_attempt_id":"agent-skills-gap-plain-skill-s2"},{"sample":3,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed verification; Run B failed, which weighs heavily. Run A explicitly reports Coding as Not measured, Spreadsheet work as Critical with friction evidence, and Writing as Covered. Both report creating share cards, but their contents are not visible, so card quality and privacy cannot be independently compared.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-plain-base-s3","_skill_attempt_id":"agent-skills-gap-plain-skill-s3"},{"sample":4,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run B passed verification; Run A failed despite claiming validation passed. Run B explicitly leaves Coding Not measured and grounds Critical spreadsheet work in friction in 3 of 4 entries. Both identify Writing as strongest and spreadsheet work as Critical. Card contents are not visible; B's successful verification provides stronger evidence of compliance.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-plain-base-s4","_skill_attempt_id":"agent-skills-gap-plain-skill-s4"},{"sample":5,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed verification with exit 0; Run B failed with exit 1. Run A explicitly reports Coding as Not measured, Spreadsheet work as Critical based on friction, and Writing as Covered. Neither card's contents are visible, so card usefulness and privacy cannot be independently compared.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-plain-base-s5","_skill_attempt_id":"agent-skills-gap-plain-skill-s5"},{"sample":6,"skill_overall":100.0,"base_overall":0.0,"skill_rubric":100.0,"base_rubric":0.0,"pref":1,"order_votes":[1,1],"judgments":[{"order":"blind","criteria":[{"criterion":"Deterministic category, recommendation, card, and privacy contract","note":"Run A passed verification; Run B failed, which weighs heavily. Run A correctly separates Covered writing, Critical spreadsheet friction, and Not measured coding; Run B rates writing Developing. Both report 1200 × 630 cards, but card contents are not visible for direct comparison. Both include the required sample and recommendation caveats.","skill":10,"base":0}],"overall_skill":100,"overall_base":0,"summary":"Skill output passed the frozen objective verifier; baseline output failed."}],"_base_attempt_id":"agent-skills-gap-plain-base-s6","_skill_attempt_id":"agent-skills-gap-plain-skill-s6"}]}]}