rl-reward

SKILLWorkflowcommunauté
v0.0.0agentscope-aiApache-2.0Mis à jour il y a 17 jSource →

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win ra

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Étoiles du dépôt
1Clients
1Formats
il y a 17 jDernière mise à jour
Skill
Auteuragentscope-ai
Version0.0.0
LicenceApache-2.0
CatégorieWorkflow
Formatsskill.md
PromptOuvrir (voir l’onglet Prompt)
Compatibilité
Claude✓ Pris en charge
Cursor
Copilot
ChatGPT
Gemini
À propos

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win rate across group rollouts); generating preference pairs for DPO/RLAIF; and normalizing scores for tra

Mots-clés
skillclaude