design-resumable-model-evaluation

SKILLWorkflowcommunauté
v0.0.0bastosMITMis à jour il y a 14 jSource →

Designs strict model benchmarks that persist per-case evidence, resume without repeating completed work, separate first-attempt behavior from remediation, and stop early only when failure is mathematically certain. Use for slow, costly, or interruptible evaluations.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
6Étoiles du dépôt
1Clients
1Formats
il y a 14 jDernière mise à jour
Skill
Auteurbastos
Version0.0.0
LicenceMIT
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Designs strict model benchmarks that persist per-case evidence, resume without repeating completed work, separate first-attempt behavior from remediation, and stop early only when failure is mathematically certain. Use for slow, costly, or interruptible evaluations.

Mots-clés
skillclaude