auto-arena

SKILLWorkflowCommunity
v0.0.0agentscope-aiApache-2.0Aktualisiert vor 17 TQuelle →

Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Repo-Sterne
1Clients
1Formate
vor 17 TLetzte Aktualisierung
Skill
Autoragentscope-ai
Version0.0.0
LizenzApache-2.0
KategorieWorkflow
Formateskill.md
PromptÖffnen (siehe Tab „Prompt“)
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model

Schlagwörter
skillclaude