research-loop

SKILLWorkflowcommunauté
v0.0.0EmlembowMITMis à jour il y a 8 jSource →

Run autonomous, metric-driven experiments on a version-controlled implementation against a fixed evaluation harness. Use when the user asks to improve eval pass rate, benchmark score, prompt or policy quality, performance, cost, or another measurable outcome through repeated hypothesis, change, eval

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
2Étoiles du dépôt
1Clients
1Formats
il y a 8 jDernière mise à jour
Skill
AuteurEmlembow
Version0.0.0
LicenceMIT
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Run autonomous, metric-driven experiments on a version-controlled implementation against a fixed evaluation harness. Use when the user asks to improve eval pass rate, benchmark score, prompt or policy quality, performance, cost, or another measurable outcome through repeated hypothesis, change, evaluate, keep-or-discard cycles. Protect generalization with holdout gates and reject hardcoded cases,

Mots-clés
skillclaude