eval-integrity

SKILLFlujo de trabajocomunidad
v0.0.0conorbronsdonMITActualizado hace 14 dFuente →

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo fo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
9Estrellas del repo
1Clientes
1Formatos
hace 14 dÚltima actualización
Skill
Autorconorbronsdon
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo for evidence across seven dimensions (pre-registration, contamination, holdout hygiene, judge validity

Palabras clave
skillclaude