rl-reward

SKILLWorkflowCommunity
v0.0.0agentscope-aiApache-2.0Aktualisiert vor 17 TQuelle →

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win ra

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Repo-Sterne
1Clients
1Formate
vor 17 TLetzte Aktualisierung
Skill
Autoragentscope-ai
Version0.0.0
LizenzApache-2.0
KategorieWorkflow
Formateskill.md
PromptÖffnen (siehe Tab „Prompt“)
Kompatibilität
Claude✓ Unterstützt
Cursor
Copilot
ChatGPT
Gemini
Über

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win rate across group rollouts); generating preference pairs for DPO/RLAIF; and normalizing scores for tra

Schlagwörter
skillclaude