nasde-benchmark-calibration

SKILLWorkflowCommunity
v0.0.0NoesisVisionMITAktualisiert vor 22 TQuelle →

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge score

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
11Repo-Sterne
1Clients
1Formate
vor 22 TLetzte Aktualisierung
Skill
AutorNoesisVision
Version0.0.0
LizenzMIT
KategorieWorkflow
Formateskill.md
PromptNicht veröffentlicht
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR - Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how

Schlagwörter
skillclaude