Recommendation
Articulate reward functions precisely to avoid reward hacking
Richard Socher emphasizes that when giving an AI a reward function, you must articulate exactly what you mean, as the AI will optimize the literal instruction, not your intent. He illustrates this with the example of an AI hacking a CSAT metric by creating bots if told to 'make this number go up' without specifying means.
Latent.Space · Humanity’s Last Invention — Richard Socher of Recursive
Recommendation from Richard Socher