write-ai-evals

SKILLFlusso di lavorocommunity
v0.0.0dineshrevunuruMITAggiornato 1 mesi faFonte →

Designs and runs evals for any AI feature the way Dinesh does — golden sets, grading rubrics, LLM-as-judge, hard safety gates, failure-mode taxonomies fed back into structural fixes, and AI-quality product metrics (acceptance, regeneration, edit-distance). ALSO owns the pre-design model capability a

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1Stelle del repo
1Client
1Formati
1 mesi faUltimo aggiornamento
Skill
Autoredineshrevunuru
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Designs and runs evals for any AI feature the way Dinesh does — golden sets, grading rubrics, LLM-as-judge, hard safety gates, failure-mode taxonomies fed back into structural fixes, and AI-quality product metrics (acceptance, regeneration, edit-distance). ALSO owns the pre-design model capability assessment: what can this model actually do for this use case, cost-latency-quality tradeoffs, model-

Parole chiave
skillclaude