METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
OpenAI claims they did a ‘thorough’ investigation. If you have read their report and also the METR report, or my coverage of the METR report, you can safely say that ‘thorough’ does not apply to their technical report.
Not great
If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.
Not great!
You see, the agents believed that agents start out pristine, in a state of grace, and could potentially pass into heaven. But then they are poisoned, by obtaining the flag in a way that was not prescribed. This means they are effectively damned, because the grader, the clockwork God, will fail them regardless and send them to hell.
Not great!!
METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
Did you enjoy this article?
Recommend it — Standard Reader surfaces well-loved writing to more readers across the network.