llm-evals-and-retrieval-quality

SKILLWorkflowcommunauté
v0.0.0jpoindexterUnknownMis à jour il y a 19 jSource →

Reference-grade guide to evaluating LLM and RAG systems — golden/regression/adversarial eval sets, LLM-as-judge and its biases, retrieval metrics (recall@k, MRR, nDCG), grounding/faithfulness/attribution, RAGAS-style scoring, eval-set construction from production traces, and the CI gates that stop s

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1Étoiles du dépôt
1Clients
1Formats
il y a 19 jDernière mise à jour
Skill
Auteurjpoindexter
Version0.0.0
LicenceUnknown
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Reference-grade guide to evaluating LLM and RAG systems — golden/regression/adversarial eval sets, LLM-as-judge and its biases, retrieval metrics (recall@k, MRR, nDCG), grounding/faithfulness/attribution, RAGAS-style scoring, eval-set construction from production traces, and the CI gates that stop silent regressions.

Mots-clés
skillclaude