eval-design

SKILLFlujo de trabajocomunidad
v0.0.0agentscope-aiApache-2.0Actualizado hace 17 dFuente →

Use when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a labeled evaluation set. Also use when the user mentions test data design, eval coverage, difficulty stratification, synthet

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Estrellas del repo
1Clientes
1Formatos
hace 17 dÚltima actualización
Skill
Autoragentscope-ai
Versión0.0.0
LicenciaApache-2.0
CategoríaFlujo de trabajo
Formatosskill.md
PromptAbrir (ver la pestaña Prompt)
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Use when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a labeled evaluation set. Also use when the user mentions test data design, eval coverage, difficulty stratification, synthetic data generation for eval, or "how to create good evaluation data." Outputs datasets in OpenJudge-

Palabras clave
skillclaude