agent-evaluation

SKILLWorkflowcommunity
v0.0.0rootcastlecoNOASSERTIONUpdated 4d agoSource →

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring\u2014where even top agents achieve less than 50% on re...

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
6Repo stars
1Clients
1Formats
4d agoLast update
Skill
Authorrootcastleco
Version0.0.0
LicenseNOASSERTION
CategoryWorkflow
Formatsskill.md
PromptNot published
Compatibility
Claude✓ Supported
Cursor
Copilot
ChatGPT
Gemini
About

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring\u2014where even top agents achieve less than 50% on re...

Keywords
skillclaude