eval-integrity

SKILLWorkflowcommunity
v0.0.0conorbronsdonMITUpdated 1mo agoSource →

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo fo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
9Repo stars
1Clients
1Formats
1mo agoLast update
Skill
Authorconorbronsdon
Version0.0.0
LicenseMIT
CategoryWorkflow
Formatsskill.md
PromptNot published
Compatibility
Claude✓ Supported
Cursor—
Copilot—
ChatGPT—
Gemini—
About

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo for evidence across seven dimensions (pre-registration, contamination, holdout hygiene, judge validity

Keywords
skillclaude

No dependency coverage

This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.