RLHF · Reinforcement Learning from Human Feedback
Tuning a model using human preference rankings to make its outputs more helpful and aligned.
Current numbers
560 GBWeights alone: four equal 70B PPO-RLHF models at two bytes/weight; placement and runtime excluded