auto-arena

SKILLFlujo de trabajocomunidad
v0.0.0agentscope-aiApache-2.0Actualizado hace 17 dFuente →

Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Estrellas del repo
1Clientes
1Formatos
hace 17 dÚltima actualización
Skill
Autoragentscope-ai
Versión0.0.0
LicenciaApache-2.0
CategoríaFlujo de trabajo
Formatosskill.md
PromptAbrir (ver la pestaña Prompt)
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model

Palabras clave
skillclaude