evaluating-llms

SKILLWorkflowcommunity
v0.0.0ancolemanMITUpdated 9mo agoSource →

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
508Repo stars
1Clients
1Formats
9mo agoLast update
Skill
Authorancoleman
Version0.0.0
LicenseMIT
CategoryWorkflow
Formatsskill.md
PromptOpen (see Prompt tab)
Compatibility
Claude✓ Supported
Cursor—
Copilot—
ChatGPT—
Gemini—
About

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

Keywords
skillclaude

No dependency coverage

This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.