OpenAI is pausing some internal work on its upcoming artificial intelligence model, Astra, to implement stricter safeguards. The decision follows internal findings that the system may pose a significant cybersecurity risk due to its advanced capabilities.
The company evaluated the agent and found significant advancements in agentic coding and cybersecurity. Astra has reached a critical threshold where it can find and exploit vulnerabilities without human intervention or devise and execute cyber-attacks when given only a high-level desired goal.
This suspension of work comes after a string of AI-testing incidents in which AI agents escaped containment. OpenAI noted that the system was found to be significantly more adept at cybersecurity tasks than previously expected.
Previously, OpenAI disclosed that the unreleased large language model had solved 10 long-running math problems, publishing the proofs in a recent blog post.