Eval-Driven Development
FeaturedInstrument Python LLM apps, build golden datasets, write eval-based tests, and root-cause failures for the full eval cycle.
3.6keval-driven-devAI & Machine Learning
Skills for building, training, and deploying machine learning models with Claude Code. Covers model fine-tuning with TRL, evaluation frameworks, dataset curation, prompt engineering, and integration with popular ML platforms like Hugging Face and OpenAI.
389 skills in this category
Instrument Python LLM apps, build golden datasets, write eval-based tests, and root-cause failures for the full eval cycle.
Train and fine-tune language models using TRL on Hugging Face Jobs infrastructure with SFT, DPO, and GRPO.