ref-hallucination-arena

SKILLFlusso di lavorocommunity
v0.0.0agentscope-aiApache-2.0Aggiornato 18 g faFonte →

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Stelle del repo
1Client
1Formati
18 g faUltimo aggiornamento
Skill
Autoreagentscope-ai
Versione0.0.0
LicenzaApache-2.0
CategoriaFlusso di lavoro
Formatiskill.md
PromptApri (vedi la scheda Prompt)
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucina

Parole chiave
skillclaude