Reinforcement Learning from Human RLHF
The method of training that model is rewarded by human preferences.
In RLHF, people compare and rating model outputs, a reward model from these preferences is trained and the main model is set to maximize this award with consolidation learning. ChatGPT has played a critical role in increasing the helpfulness of assistants.
