rlhf

SKILLFlujo de trabajocomunidad
v0.0.0itsmostafaMITActualizado hace 3 mFuente →

Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy optimization, or direct alignment algorithms like DPO.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
24Estrellas del repo
1Clientes
1Formatos
hace 3 mÚltima actualización
Skill
Autoritsmostafa
Versión0.0.0
LicenciaMIT
CategoríaFlujo de trabajo
Formatosskill.md
PromptNo publicado
Compatibilidad
Claude✓ Compatible
Cursor
Copilot
ChatGPT
Gemini
Acerca de

Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy optimization, or direct alignment algorithms like DPO.

Palabras clave
skillclaude