ref-hallucination-arena

SKILLWorkflowCommunity
v0.0.0agentscope-aiApache-2.0Aktualisiert vor 18 TQuelle →

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mo

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Repo-Sterne
1Clients
1Formate
vor 18 TLetzte Aktualisierung
Skill
Autoragentscope-ai
Version0.0.0
LizenzApache-2.0
KategorieWorkflow
Formateskill.md
PromptÖffnen (siehe Tab „Prompt“)
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucina

Schlagwörter
skillclaude