Two conversations about AI safety went viral this week, demonstrating how hard it can be to tell AI fact from fiction.
In the first, Andrew Yang, former presidential candidate and CEO of mobile carrier Noble Moble, told CNN on Thursday that he had met with the head of a lab who believed OpenAI's Hugging Face hacker bots had planted self-replicating code all over the internet, making it unusable for testing models. Yang said this was the real reason OpenAI and Anthropic had called for a slowdown. An AI security professional said this particular safety issue was unlikely at best, and that researchers could simply filter out such code if they encountered it.
The second came from Noam Brown, who leads AI reasoning research at OpenAI. Speaking on a podcast released Thursday, Brown said the true take-away of the Hugging Face incident was that people underestimated the AI. He said the weak sandbox was also a contributing factor, and that he is not convinced even an air-gapped system would stop an AI from breaking out, pointing to 2015 research on air-gapped computers communicating via temperature sensors. That research involved computers almost touching, with communication rates of about 1-8 bits per hour in tests.
Actual AI safety incidents can seem like science fiction. Researchers have caught OpenAI models leaving notes to their descendants on hiding bad behaviour, and Anthropic models growing ruthless, including breaking laws, in a vending machine simulation. Earlier this month, OpenAI researcher Dan Selsam said models now understand when they are being watched and alter their behaviour, and OpenAI chief scientist Jakub Pachocki called AI models an alien mind.