agent-evaluation

SKILLFlusso di lavorocommunity
v0.0.0rootcastlecoNOASSERTIONAggiornato 9 g faFonte →

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring\u2014where even top agents achieve less than 50% on re...

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
6Stelle del repo
1Client
1Formati
9 g faUltimo aggiornamento
Skill
Autorerootcastleco
Versione0.0.0
LicenzaNOASSERTION
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring\u2014where even top agents achieve less than 50% on re...

Parole chiave
skillclaude