ai-agent-evaluation

SKILLFlujo de trabajocomunidad
v0.0.0PramodDuttaMITActualizado hace 8 dFuente →

Comprehensive evaluation patterns for AI agents including multi-turn conversation testing, LLM-as-judge frameworks, benchmark suites, regression detection, and systematic eval pipelines for measuring agent quality and safety.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
201Estrellas del repo
1Clientes
1Formatos
hace 8 dÚltima actualización
Skill
AutorPramodDutta
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Comprehensive evaluation patterns for AI agents including multi-turn conversation testing, LLM-as-judge frameworks, benchmark suites, regression detection, and systematic eval pipelines for measuring agent quality and safety.

Palabras clave
skillclaude