eval-integrity

SKILLFlusso di lavorocommunity
v0.0.0conorbronsdonMITAggiornato 14 g faFonte →

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo fo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
9Stelle del repo
1Client
1Formati
14 g faUltimo aggiornamento
Skill
Autoreconorbronsdon
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo for evidence across seven dimensions (pre-registration, contamination, holdout hygiene, judge validity

Parole chiave
skillclaude