eval-integrity

SKILLWorkflowCommunity
v0.0.0conorbronsdonMITAktualisiert vor 14 TQuelle →

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo fo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
9Repo-Sterne
1Clients
1Formate
vor 14 TLetzte Aktualisierung
Skill
Autorconorbronsdon
Version0.0.0
LizenzMIT
KategorieWorkflow
Formateskill.md
PromptNicht veröffentlicht
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo for evidence across seven dimensions (pre-registration, contamination, holdout hygiene, judge validity

Schlagwörter
skillclaude