nasde-benchmark-calibration

SKILLFlujo de trabajocomunidad
v0.0.0NoesisVisionMITActualizado hace 23 dFuente →

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge score

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
11Estrellas del repo
1Clientes
1Formatos
hace 23 dÚltima actualización
Skill
AutorNoesisVision
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR - Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how

Palabras clave
skillclaude