Edge Copy setup link
Skill evidence / exploring-llm-evaluations
Skill profile

exploring-llm-evaluations

Investigate AI observability evaluations : `hog` (deterministic code-based).

What this skill does for you

Investigate AI observability evaluations : `hog` (deterministic code-based), `llm_judge` (LLM-prompt-based), and `sentiment` (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and set up scheduled reports on an evaluation. Use when the user asks to debug why an evaluation is failing, surface common failure modes, compare results across filters, dry-run a Hog evaluator, prototype a new LLM-judge prompt, inspect sentiment classifications, or manage the evaluation lifecycle.

Quoted from the skill description in the pinned source

Who made it

Publisherposthog
Repositoryskills
Installs343as of Sep 27, 2026

Use it with Edge

Set up Edge for me: read getedge.cc/SKILL.md and follow it. Then use Edge to load the exploring-llm-evaluations skill from posthog/skills.

Paste it into Claude, ChatGPT, Codex or Cursor.

Or read the skill source first.

Security evidence

Security state at 2026-09-30: reviewed. State definitions. This records static evidence for this revision; live load eligibility is checked again.

3 static scans, 0 counted findings. The listed scanners recorded the results shown for this package revision.

ScannerDateResultFindings
Edge static checks2026-09-27Pass0
Cisco skill-scanner2026-09-27Pass0
Semgrep Edge rules2026-09-27Pass0

Static scans check the code, not how well the skill works.

Scanner scope, raw findings and mirrored provider records

Edge static checks (Edge-run): edge-static/1.0.1. Scope: Static patterns for remote execution, credentials with network sends, obfuscation, prompt overrides, hidden Unicode, binaries and persistence. Revision: 835ceac8ec56af5ca9af647c815abc2fa6f9b330. Raw findings: . Counted findings: .

Cisco skill-scanner (Edge-run): cisco-skill-scanner/2.1.0. Scope: Local static, YARA, pipeline and behavioral analyzers; no LLM or AI Defense analysis. Revision: 835ceac8ec56af5ca9af647c815abc2fa6f9b330. Raw findings: 1. Counted findings: .

Semgrep Edge rules (Edge-run): semgrep/1.178.0+edge-rules.c15016a36fd3+offline-v1. Scope: Six Edge-authored static rules with Semgrep CE; offline execution and metrics disabled. Semgrep-maintained rules excluded. Revision: 835ceac8ec56af5ca9af647c815abc2fa6f9b330. Raw findings: . Counted findings: .

Source and licence

Repository
https://github.com/posthog/skills
Pinned skill file
View source at revision
Revision
835ceac8ec56af5ca9af647c815abc2fa6f9b330
Package hash
sha256:7ca96b13fe85e0af4a34a8fc58f413ffac585f97b6dc16980ad94f2fd1316fd8
Licence
MIT