rlhf

SKILLWorkflowcommunauté
v0.0.0itsmostafaMITMis à jour il y a 3 moisSource →

Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy optimization, or direct alignment algorithms like DPO.

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
24Étoiles du dépôt
1Clients
1Formats
il y a 3 moisDernière mise à jour
Skill
Auteuritsmostafa
Version0.0.0
LicenceMIT
CatégorieWorkflow
Formatsskill.md
PromptNon publié
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Understanding Reinforcement Learning from Human Feedback (RLHF) for aligning language models. Use when learning about preference data, reward modeling, policy optimization, or direct alignment algorithms like DPO.

Mots-clés
skillclaude