rl-reward

SKILLFlusso di lavorocommunity
v0.0.0agentscope-aiApache-2.0Aggiornato 2 mesi faFonte →

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win ra

Community-submitted skill. Not yet reviewed by the Forge team. Full prompt content may not be available.Request review →
789Stelle del repo
1Client
1Formati
2 mesi faUltimo aggiornamento
Skill
Autoreagentscope-ai
Versione0.0.0
LicenzaApache-2.0
CategoriaFlusso di lavoro
Formatiskill.md
PromptApri (vedi la scheda Prompt)
Compatibilità
Claude✓ Supportato
Cursor—
Copilot—
ChatGPT—
Gemini—
Descrizione

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win rate across group rollouts); generating preference pairs for DPO/RLAIF; and normalizing scores for tra

Parole chiave
skillclaude

Nessuna copertura delle dipendenze

Questa voce non pubblica alcun pacchetto npm, quindi Forge non ha un albero delle dipendenze per essa. È una lacuna di copertura, non l'affermazione che non abbia dipendenze.