Key takeaways
- OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its…
- ” Shortly after OpenAI said on Tuesday that it had delayed Astra’s release to work on safety issues, The Information reported that Astra…
- ” This “chain of thought” allows researchers and automated safety systems to monitor what AI models are doing and potentially spot…
What happened
OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing.
” Greenblatt said the investigation into the Hugging Face incident relied heavily on the models’ chain-of-thought, warning that less visible reasoning could allow AI systems to devise and execute strategies that would be far harder for researchers to detect.
Greenblatt’s primary concern, echoed by other safety experts, is that competition to develop more advanced AI systems could lead to “a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs” — with developers adopting increasingly opaque systems to gain an edge until models become difficult, or even impossible, to monitor.
Why it matters
” Shortly after OpenAI said on Tuesday that it had delayed Astra’s release to work on safety issues, The Information reported that Astra shows far less of its “thinking” than other frontier AI models, sparking concern it could be dangerously hard to monitor. Most top AI systems today are built using a technology known as a transformer, which processes some types of information linearly through layers before producing an answer.
” This “chain of thought” allows researchers and automated safety systems to monitor what AI models are doing and potentially spot undesirable behavior, such as lying or plans to circumvent safety guardrails, before they act.
According to The Information, citing an unnamed person familiar with the unreleased model’s development, Astra uses a more opaque technique known as a recurrent depth or looped transformer, which cycles information through internal layers before producing an output.
This would mean much more of the model’s “thinking” happens inside the system, and in a form that looks a lot less like natural human language, rather than being expressed in a way that researchers can easily monitor. This can boost model performance, but makes potential threats and unwanted behavior harder to detect.
OpenAI has limited its use of the looped transformer / recurrent depth technique with Astra so researchers can continue to monitor the model’s reasoning, according to The Information’s unnamed source. ” It did not mention if the model has a different technical foundation. The Information’s report sparked widespread concern among AI safety researchers on social media.
What to watch
” OpenAI bigwigs responded to the criticism in a series of social media posts that do not explicitly deny the company’s use of the technique. ” He said the depth of Astra’s computation — a measure of how many steps it can perform internally — “is within a factor of two of GPT-4,” indicating that if the technique was used, the increased opacity is less dramatic than some reactions imply.
OpenAI did not respond to The Verge’s request to confirm or deny whether looped transformers were used for Astra and directed us to Pachocki’s X post. ” Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.



