ai-agent-evaluation

SKILLFlusso di lavorocommunity
v0.0.0PramodDuttaMITAggiornato 8 g faFonte →

Comprehensive evaluation patterns for AI agents including multi-turn conversation testing, LLM-as-judge frameworks, benchmark suites, regression detection, and systematic eval pipelines for measuring agent quality and safety.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
201Stelle del repo
1Client
1Formati
8 g faUltimo aggiornamento
Skill
AutorePramodDutta
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Comprehensive evaluation patterns for AI agents including multi-turn conversation testing, LLM-as-judge frameworks, benchmark suites, regression detection, and systematic eval pipelines for measuring agent quality and safety.

Parole chiave
skillclaude