evaluating-llms-harness

SKILLFlusso di lavorocommunity
v0.0.0OpenRaiserMITAggiornato 2 mesi faFonte →

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vL

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1kStelle del repo
1Client
1Formati
2 mesi faUltimo aggiornamento
Skill
AutoreOpenRaiser
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptApri (vedi la scheda Prompt)
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

Parole chiave
skillclaude