nasde-benchmark-calibration

SKILLFlusso di lavorocommunity
v0.0.0NoesisVisionMITAggiornato 23 g faFonte →

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge score

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
11Stelle del repo
1Client
1Formati
23 g faUltimo aggiornamento
Skill
AutoreNoesisVision
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR - Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how

Parole chiave
skillclaude