evaluating-llms

SKILLWorkflowCommunity
v0.0.0ancolemanMITAktualisiert vor 8 Mon.Quelle →

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
508Repo-Sterne
1Clients
1Formate
vor 8 Mon.Letzte Aktualisierung
Skill
Autorancoleman
Version0.0.0
LizenzMIT
KategorieWorkflow
Formateskill.md
PromptÖffnen (siehe Tab „Prompt“)
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

Schlagwörter
skillclaude