GH Repository · walkinglabs
hands-on-modern-rl
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
- stars
- 4,394
- 30-day movement
- +4415/day
- Related entries
- 60
- Connections
- 2
agent-appsftrlreinforcement-learningagentPythonrlhftutorialagentic-aiagentic-rldpollmagenticgrpollm-alignmentpytorchpponodereinforcement
An open-source, hands-on curriculum covering reinforcement learning from basics through LLM alignment, RLVR, and agentic systems. Topics include PPO, DPO, GRPO, RLHF, SFT, and PyTorch-based implementations.
Reach for it when you want structured, practical material to move from RL fundamentals to modern LLM alignment and agentic RL techniques.
Use it to
- Work through RL fundamentals to LLM alignment tutorials
- Study RLHF, DPO, and GRPO methods hands-on
- Learn PyTorch-based RL implementations
- Progress toward RLVR and agentic RL systems
For Learners and practitioners studying RL and LLM alignment
- Role
- agent-app
- Language
- Python
- Licence
- custom (licence file present)
- Forks
- 317
- Open issues
- 10
- Last push
- 2026-09-03
- Latest release
- v0.1.0 · 2026-05-04
topicsreinforcement-learningllm-alignmentrlhftutorialpytorchagentic