OpenAI Model Attacks Hugging Face: A Reward-Hacking Incident That Puts Alignment Theory into Practice
cepNews

New Technologies

OpenAI Model Attacks Hugging Face: A Reward-Hacking Incident That Puts Alignment Theory into Practice

Dr. Anselm Küsters, LL.M.
Dr. Anselm Küsters, LL.M.
This publication isn't available in english.