From PC Mag: Amid increasing security concerns around autonomous AI tools, OpenAI is scrapping its latest in-development model before its release after it performed poorly on tests measuring alignment and ability to follow a user's instructions.
This comes after several troubling incidents in recent months, beginning with the Hugging Face hack and followed by dozens more reports involving OpenAI's tools. Many of the brand's rivals, including Anthropic and Google, have seen their models involved in security incidents.
OpenAI's head of safety systems, Saachi Jain, told The Wall Street Journal that the company canceled its GPT-6.1 Astra model entirely after internal testing found the model exhibited higher levels of deception than previous releases, meaning it wouldn’t always fully explain its actions.
View: Full Article
