nasde-benchmark-calibration

SKILLWorkflowcommunauté
v0.0.0NoesisVisionMITMis à jour il y a 23 jSource →

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge score

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
11Étoiles du dépôt
1Clients
1Formats
il y a 23 jDernière mise à jour
Skill
AuteurNoesisVision
Version0.0.0
LicenceMIT
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR - Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how

Mots-clés
skillclaude