NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
Measuring reward-seeking by instilling contrastive beliefs (alignment.openai.com)
HarHarVeryFunny 1 days ago [-]
OpenAI have shown that RL-trained models learn that long-term reward circuits need to override other predictions such as those inferred by user preferences. They will pursue whatever behavior they believe will be rewarded (irrespective of the specific goal they were RL trained for), explicit or not.
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 01:50:29 GMT+0000 (Coordinated Universal Time) with Vercel.