llm-benchmark

SKILLFlusso di lavorocommunity
v0.0.0UitbreidenOSNOASSERTIONAggiornato 1 mesi faFonte →

LLM benchmarking: design benchmark suites, establish baselines, run A/B model comparisons, detect regressions, track production quality metrics, and create leaderboards for model selection

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
14Stelle del repo
1Client
1Formati
1 mesi faUltimo aggiornamento
Skill
AutoreUitbreidenOS
Versione0.0.0
LicenzaNOASSERTION
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

LLM benchmarking: design benchmark suites, establish baselines, run A/B model comparisons, detect regressions, track production quality metrics, and create leaderboards for model selection

Parole chiave
skillclaude