Home

reinforcement-learning-human-feedback-scratch

Public

End-to-end implementation of Reinforcement Learning with Human Feedback (RLHF) to align a GPT-2 model with human preferences — covering Supervised Fine-Tuning (SFT), Reward Modeling, and PPO-based alignment — built from scratch in Python.

Creat:2025-09-12T22:06:18
Update:2025-09-24T02:17:00
2
Stars
0
Stars Increase