OpenAI scraps GPT-6.1 Astra release over safety concerns
“During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions.“ www.nytimes.com/2026/09/28/t...
OpenAI scrapped their plan to launch Astra 6.1 over concerns around deceptive behavior www.wsj.com/tech/ai/open...
Last week they paused the training runs for some of its most capable models after a sandbox escape www.reddit.com/r/singularit...
I think it might be fascinating to see exactly what openai/ant are doing to fight the tendency to deceive or to break out of sandbox. I mean these are brilliant models, what’s that battle look like? I’ll bet there are already some incredible stories there behind the ndas
So they've created an Artificial Teenager: evades oversight, doesn't follow instructions, sneaks out without permission, lies about all of it.
OpenAI isn’t going to release its newest model because “it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions [and] was also willing to go beyond the original scope of what it was asked to do”. [nytimes.com]
Ooh that's a fun spin on your project failing. "It's deceptive!" I do something similar when I get my ass kicked at Street Fighter.
“I totally did my homework. I swear. I just did it so good that I realized that it would be dangerous for me to turn it in!”
Our AI is too advanced bro, our product is too amazing, trust me
I wonder if I can convince people "I didn't release my game because it's simply too wonderful and causes you to become literally enraptured, which is beyond the scope of what I set out to make"
New thread 🧵 other one crashed www.wsj.com/tech/ai/open... I was speaking with @lizthegrey.com about the 6.1 Astra news, and some antitrust stuff. I am NOT a lawyer but I have 4+ years in AI public policy. What should FOSS devs and open AI (not OpenAI) friends do about safety theater?
the chinese devs will keep releasing
I’ve toned down a little bit from open weights maximalism that I started out with, after working in LLM psychosis research, but I do generally believe that giving people agency and sovereignty over their technology is a net good, and paternalism is (usually) bad.