Nicholas Decker In Hell
Sep 1, 20261 min read
Nobody knows, in detail, how RLHF works. It negatively reinforces some behavior: somewhere in the black box of AI innards, it changes some parameters to make the AI perform the behavior less. If we let ourselves anthropomorphize the AI, what is the human equivalent to this? A child steals cookies and is punished: does the child no longer like the taste of cookies? Does he fear further punishment in the future? Does he grasp the beauty and compellingness of the moral law? Does he have an undefinable cloud of dread around the whole concept of theft and/or cookies?
Helpful analogy
Did you enjoy this article?
Recommend it — Standard Reader surfaces well-loved writing to more readers across the network.