Skip to main content
Featured Skill

Eval-Driven Development

by eval-driven-dev

Instrument Python LLM apps, build golden datasets, write eval-based tests, and root-cause failures for the full eval cycle.

About This Skill

A comprehensive skill for eval-driven development of LLM applications. Covers the full evaluation lifecycle: instrumenting Python LLM apps for observability, building golden datasets from production traffic, writing evaluation-based test suites, running evaluations at scale, and systematic root-cause analysis of failures. Essential for teams building production AI applications.

How to Install

# Clone or download the skill to your skills directory git clone https://github.com/levnikolaevich/claude-code-skills ~/.claude/skills/eval-driven-development

Or download the SKILL.md file directly and place it in your project's .claude/skills/ directory.

Tags

evalllmtestingpythonmachine-learningobservability