eval-integrity

SKILLWorkflowcommunauté
v0.0.0conorbronsdonMITMis à jour il y a 14 jSource →

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo fo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
9Étoiles du dépôt
1Clients
1Formats
il y a 14 jDernière mise à jour
Skill
Auteurconorbronsdon
Version0.0.0
LicenceMIT
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo for evidence across seven dimensions (pre-registration, contamination, holdout hygiene, judge validity

Mots-clés
skillclaude