pydantic-evals

SKILLWorkflowcommunity
v0.0.0FuenfgeldMITUpdated 3mo agoSource →

Test and evaluate AI agents and LLM outputs using code-first evaluation framework with strong typing. Use when the user wants to: (1) Create evaluation datasets with test cases for AI agents, (2) Define evaluators (deterministic, LLM-as-Judge, custom, or span-based), (3) Run evaluations and generate

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
10Repo stars
1Clients
1Formats
3mo agoLast update
Skill
AuthorFuenfgeld
Version0.0.0
LicenseMIT
CategoryWorkflow
Formatsskill.md
PromptNot published
Compatibility
Claude✓ Supported
Cursor—
Copilot—
ChatGPT—
Gemini—
About

Test and evaluate AI agents and LLM outputs using code-first evaluation framework with strong typing. Use when the user wants to: (1) Create evaluation datasets with test cases for AI agents, (2) Define evaluators (deterministic, LLM-as-Judge, custom, or span-based), (3) Run evaluations and generate reports, (4) Compare model performance across experiments, (5) Integrate evaluations with Pydantic

Keywords
skillclaude

No dependency coverage

This entry publishes no npm package, so Forge has no dependency tree for it. That is a gap in coverage — not a statement that it has no dependencies.