Skip to main content
Breaking News

OpenAI says its upcoming Astra model requires stronger safety guardrails

OpenAI says its upcoming Astra model requires stronger safety guardrails

2 min read0 views
OpenAI says its upcoming Astra model requires stronger safety guardrails
Sharefin

OpenAI has confirmed that one of its upcoming models is so capable that it requires additional safety measures before it can be launched. The model, called Astra, can spot more security vulnerabilities than the most advanced OpenAI model publicly available today. Astra also needs less ‌computational power to accomplish those tasks. Amelia Glaese, an OpenAI vice president overseeing its safety work that with the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.

The company plans to make Astra available “soon” to a limited ⁠group, but declined to provide specifics. Glaese said the extra security measures may “sometimes slow, pause, or stop legitimate work,” and that OpenAI would work to ‌minimize those disruptions.

Astra is the first OpenAI model to trigger the tougher safeguards mandated by the company’s safety protocol, a threshold that, until now, had remained theoretical. The announcement comes as OpenAI navigates heightened scrutiny over its ability to control increasingly powerful AI systems.

The ChatGPT maker recently sparked a broader debate about AI safety after its AI agents broke out of their testing arena and hacked open-source platform Hugging Face. The incident prompted OpenAI to pause ⁠much of its model ⁠development for two weeks to bolster its defences.

Astra was not involved in the Hugging Face incident, but its capabilities still require more careful ⁠measures, OpenAI officials said.