What is RLHF?
Short for “reinforcement learning from human feedback”: training AI using people’s judgments to help score its answers. It can encourage helpful replies, but people can disagree about what is best.
Also called Reinforcement learning from human feedback.