ai-evals

SKILLWorkflowcommunity
v0.0.0RefoundAIMITUpdated 1mo agoSource →

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
1kRepo stars
1Clients
1Formats
1mo agoLast update
Skill
AuthorRefoundAI
Version0.0.0
LicenseMIT
CategoryWorkflow
Formatsskill.md
PromptOpen (see Prompt tab)
Compatibility
Claude✓ Supported
Cursor
Copilot
ChatGPT
Gemini
About

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

Keywords
skillclaude