build-eval-dataset

SKILLFlujo de trabajocomunidad
v0.0.0ContextJet-aiNOASSERTIONActualizado hace 11 dFuente →

Use this to build a good evaluation dataset for an LLM app, the part everyone underestimates. Trigger on "make an eval set", "what should I test my LLM on", "I don't have test data for my prompt", "build a golden dataset", or before setting up evals. A great eval set beats a great metric; garbage-in

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
29Estrellas del repo
1Clientes
1Formatos
hace 11 dÚltima actualización
Skill
AutorContextJet-ai
Versión0.0.0
LicenciaNOASSERTION
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Use this to build a good evaluation dataset for an LLM app, the part everyone underestimates. Trigger on "make an eval set", "what should I test my LLM on", "I don't have test data for my prompt", "build a golden dataset", or before setting up evals. A great eval set beats a great metric; garbage-in means your evals lie to you.

Palabras clave
skillclaude