BigHugger
sk Skill · agentscope-ai

rl-reward

Build RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and cost; aggregating multi-dimensional pointwise scores into a scalar reward; pairwise tournament reward for GRPO on subjective tasks (net win rate across group rollouts); generating preference pairs for DPO/RLAIF; and normalizing scores for training stability.…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythonbashPython
Host repository
agentscope-ai/OpenJudge
Host stars
838
Host language
Python