rlhf

SKILLFlusso di lavorocommunity
v0.0.0itsmostafaMITAggiornato 3 mesi faFonte →

Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy optimization, or direct alignment algorithms like DPO.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
24Stelle del repo
1Client
1Formati
3 mesi faUltimo aggiornamento
Skill
Autoreitsmostafa
Versione0.0.0
LicenzaMIT
CategoriaFlusso di lavoro
Formatiskill.md
PromptNon pubblicato
Compatibilità
Claude✓ Supportato
Cursor
Copilot
ChatGPT
Gemini
Descrizione

Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy optimization, or direct alignment algorithms like DPO.

Parole chiave
skillclaude