nasde-benchmark-runner

SKILLFlujo de trabajocomunidad
v0.0.0NoesisVisionMITActualizado hace 22 dFuente →

Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: - Run a benchmark (all tasks, single task, specific variant) - Re-run assessment evaluation on existing trial results - Check or verify results in Opik (traces, feedback scores, experiments) - Troublesh

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
11Estrellas del repo
1Clientes
1Formatos
hace 22 dÚltima actualización
Skill
AutorNoesisVision
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Run coding agent benchmarks and verify results with nasde. Use this skill when the user wants to: - Run a benchmark (all tasks, single task, specific variant) - Re-run assessment evaluation on existing trial results - Check or verify results in Opik (traces, feedback scores, experiments) - Troubleshoot a failed benchmark run - View or compare trial results Even if the user doesn't say "benchmark"

Palabras clave
skillclaude