google-agents-cli-eval
2,094This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failur
llm-evaluation
565Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testi
Eval Harness Skill
561Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
Agentic Evaluation Patterns
73Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loop
Eval Before Switch
Blocks a production model change until a private eval has been run. Use before any model version bump or provider switch.