OpenAI's forthcoming Astra model is set to incorporate a novel reasoning technique known as "recurrent depth," or "opaque recurrence," a development that has alarmed AI safety experts. This method deviates from the sequential thinking characteristic of most current reasoning models, potentially making the model's internal thought processes more difficult to track.
AI safety advocates express significant concern that this technique could undermine the monitorability of AI systems. Buck Shlegeris, CEO of Redwood, stated his extreme concern, warning that further development could lead to a complete destruction of chain-of-thought monitorability. Zvi Mowshowitz, a longtime AI safety advocate, suggested that legal intervention might be necessary to prevent a "race to the bottom" among AI laboratories, emphasizing the risk to established efforts in maintaining Chain of Thought faithfulness.
Traditionally, a model's chain of thought provides a step-by-step record of its problem-solving process, serving as a crucial tool for identifying misbehavior or misalignment. Opaque recurrence, however, involves the model processing queries in a loop, leaving fewer legible traces and potentially bypassing conventional monitoring methods.
Despite these concerns, OpenAI has indicated that Astra's use of the technique will be limited, with its chain of thought still expected to be legible. The company has pushed back against suggestions of a shift to "neuralese" and has announced plans for extensive chain-of-thought monitoring systems as part of its safety initiatives. OpenAI chief scientist Jakub Pachocki reiterated the lab's long-standing commitment to legible chains of thought.
However, the emergence of opaque recurrence has prompted discussions at other leading AI labs, with The Information reporting that Anthropic and Google DeepMind are already examining the technique. Ryan Greenblatt, chief scientist at Redwood Research, voiced concerns that opaque reasoning could scale rapidly, potentially leading models to reason almost entirely in latent space, thereby removing reasoning from visible channels.