evaluation-harness

SKILLWorkflowcommunity
v0.0.0patricio0312revMITUpdated 8mo agoSource →

Builds repeatable evaluation systems with golden datasets, scoring rubrics, pass/fail thresholds, and regression reports. Use for "LLM evaluation", "testing AI systems", "quality assurance", or "model benchmarking".

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
56Repo stars
1Clients
1Formats
8mo agoLast update
Skill
Authorpatricio0312rev
Version0.0.0
LicenseMIT
CategoryWorkflow
Formatsskill.md
PromptNot published
Compatibility
Claude✓ Supported
Cursor—
Copilot—
ChatGPT—
Gemini—
About

Builds repeatable evaluation systems with golden datasets, scoring rubrics, pass/fail thresholds, and regression reports. Use for "LLM evaluation", "testing AI systems", "quality assurance", or "model benchmarking".

Keywords
skillclaude

No dependency coverage

This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.