pydantic-evals

SKILLFlusso di lavorocommunity
v0.0.0FuenfgeldMITAggiornato 3 mesi faFonte →

Test and evaluate AI agents and LLM outputs using code-first evaluation framework with strong typing. Use when the user wants to: (1) Create evaluation datasets with test cases for AI agents, (2) Define evaluators (deterministic, LLM-as-Judge, custom, or span-based), (3) Run evaluations and generate

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
10Stelle del repo
1Client
1Formati
3 mesi faUltimo aggiornamento
Skill
AutoreFuenfgeld
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor—
Copilot—
ChatGPT—
Gemini—
Descrizione

Test and evaluate AI agents and LLM outputs using code-first evaluation framework with strong typing. Use when the user wants to: (1) Create evaluation datasets with test cases for AI agents, (2) Define evaluators (deterministic, LLM-as-Judge, custom, or span-based), (3) Run evaluations and generate reports, (4) Compare model performance across experiments, (5) Integrate evaluations with Pydantic

Parole chiave
skillclaude

Nessuna copertura delle dipendenze

Questa voce non pubblica alcun pacchetto npm, quindi Forge non ha un albero delle dipendenze per essa. È una lacuna di copertura, non l'affermazione che non abbia dipendenze.