pmstack-transcript-review

SKILLFlusso di lavorocommunity
v0.0.0RyanAlbertsMITAggiornato 1 mesi faFonte →

Walks a PM through Anthropic's Step 6 ritual — reading transcripts from many trials to diagnose every failed eval task as one of three things — model mistake, grader mistake, or task-spec error. Implements the practice Anthropic describes as "critical" — without it, badly-calibrated graders mask rea

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
8Stelle del repo
1Client
1Formati
1 mesi faUltimo aggiornamento
Skill
AutoreRyanAlberts
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Walks a PM through Anthropic's Step 6 ritual — reading transcripts from many trials to diagnose every failed eval task as one of three things — model mistake, grader mistake, or task-spec error. Implements the practice Anthropic describes as "critical" — without it, badly-calibrated graders mask real model improvements. Use when the user has a /run-eval result with failures, asks "why did this fai

Parole chiave
skillclaude