Featured Skill
Eval-Driven Development
by eval-driven-dev
Instrument Python LLM apps, build golden datasets, write eval-based tests, and root-cause failures for the full eval cycle.
About This Skill
A comprehensive skill for eval-driven development of LLM applications. Covers the full evaluation lifecycle: instrumenting Python LLM apps for observability, building golden datasets from production traffic, writing evaluation-based test suites, running evaluations at scale, and systematic root-cause analysis of failures. Essential for teams building production AI applications.
How to Install
# Clone or download the skill to your skills directory
git clone https://github.com/levnikolaevich/claude-code-skills ~/.claude/skills/eval-driven-developmentOr download the SKILL.md file directly and place it in your project's .claude/skills/ directory.
Tags
evalllmtestingpythonmachine-learningobservability