Skill Evals
Turn production failures into repeatable agent tests with objective verifiers, cost, latency, and trace evidence.
Build evals for an AI agent that already does real work. Use when someone asks "how do I know my agent is right", wants to test an agent before trusting it, compare models on cost versus quality, or turn production failures into tests. Walks from first tasks and yes/no verifiers, to isolated environments, to a trace-driven improvement loop.
Category: Engineering. Published via getedgehq/skills.