design-resumable-model-evaluation

SKILLFlujo de trabajocomunidad
v0.0.0bastosMITActualizado hace 14 dFuente →

Designs strict model benchmarks that persist per-case evidence, resume without repeating completed work, separate first-attempt behavior from remediation, and stop early only when failure is mathematically certain. Use for slow, costly, or interruptible evaluations.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
6Estrellas del repo
1Clientes
1Formatos
hace 14 dÚltima actualización
Skill
Autorbastos
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Designs strict model benchmarks that persist per-case evidence, resume without repeating completed work, separate first-attempt behavior from remediation, and stop early only when failure is mathematically certain. Use for slow, costly, or interruptible evaluations.

Palabras clave
skillclaude