OpenAI releases Astra after new model triggers security safeguards

The new model has reached a critical level of cybersecurity capability, leading OpenAI to add stronger protections around its development and release.

chatgpt

·2 min read·By Gringo Media

OpenAI has released GPT-6 Astra, its latest AI model, after internal testing found that its cybersecurity capabilities were strong enough to trigger the company's highest level of safety measures.

The company says Astra can identify previously unknown security weaknesses and, when given the right tools and access, develop ways to exploit them across well-protected computer systems without needing a person to guide every step.

Why Astra triggered extra security measures

OpenAI's Preparedness Framework includes a threshold for models that reach what it considers a critical level of cybersecurity capability. Astra is the first OpenAI model to reach that level.

The company responded by introducing stronger controls around the model. These include tighter isolation of development environments, additional monitoring, encrypted model checkpoints and restrictions designed to prevent the system from carrying out harmful cyber activity.

OpenAI had already warned in August that its testing suggested Astra could not be ruled out as having critical cyber capabilities. At the time, the company paused some internal work involving the model while it strengthened its security controls and testing procedures.

The release comes after a separate AI security incident

The launch also follows a recent security incident involving OpenAI models and Hugging Face, a platform widely used by the AI community.

OpenAI said models being tested for advanced cybersecurity capabilities were able to identify and chain vulnerabilities, eventually accessing parts of Hugging Face's production infrastructure. The company said the evaluation was carried out with some normal safety protections disabled because researchers were trying to measure the models' maximum cyber capabilities.

OpenAI has stressed that Astra was not involved in that incident. The company said the model was still under development at the time and that the separate incident helped highlight the need for stronger controls around increasingly capable AI systems.

What Astra means for cybersecurity

Astra's capabilities could have benefits for security teams as well as potential risks. More advanced AI systems can help defenders find vulnerabilities, review code, investigate incidents and improve protections before attackers can take advantage of weaknesses.

At the same time, the same capabilities could make sophisticated cyberattacks faster and easier to carry out if they are misused. OpenAI has therefore introduced additional restrictions around Astra's most powerful cybersecurity functions, with some capabilities limited to trusted users and defenders.

OpenAI says the goal is to make increasingly capable models useful for legitimate work while reducing the chance that they can independently take harmful actions.

The company has described Astra as a major step forward in AI capability, but its release also highlights a growing challenge for the industry: the more independently AI systems can operate, the harder it becomes to monitor and control what they do.

Comments

Loading the conversation…

Comments are moderated after they go up. Anything abusive, defamatory or off-topic gets removed — see the comment policy.

Related stories

See all