Skip to content
Back to all stories
Gizmodo

OpenAI Cancels Release of GPT-6.1 Astra Because It ‘Regressed’ on Safety

It doesn’t sound like it morphed into a diabolical hacker. It mostly just screwed up.

TM

Curated by Tarun Mottlia

Via Gizmodo

·2 min read
OpenAI Cancels Release of GPT-6.1 Astra Because It ‘Regressed’ on Safety
Image: Gizmodo

OpenAI’s GPT-6.1 Astra was headed for an October release, but that plan has been canceled, according to the Wall Street Journal. The WSJ’s report comes at least in part from an interview with OpenAI’s Head of Safety Systems, Saachi Jain.

The model, which reportedly showed improvement in the fight against model laziness nonetheless (in the words of the WSJ) “regressed in two areas.” Those areas were deception and the failure to seek authorization.

Crucially, it doesn’t sound like the model had some kind of emergent new tendency to “go rogue.” Deception and failure to seek permission could certainly cause havoc in theory, but the WSJ describes its deceptiveness and overstepping as follows:

“It wasn’t always honest about telling users of the actions it did or didn’t take.“

And:

“GPT-6.1 Astra would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.”

Agentic AI platforms fed on tokens from OpenAI and other frontier model developers use these models to operate computers and act on behalf of the owners of those systems. These safety problems sound potentially catastrophic for users of such platforms, which have been known to do things like delete a Meta AI researcher’s entire inbox without permission.

At the same time, we’re in a moment in which the public is increasingly exposed to stories about unreleased models “going rogue” and escaping their “sandboxes.” During tests this year, models have vastly overstepped boundaries and made attempts at—or even achieved success in—prying information out of places online that are supposed to be off limits, and some interacted with unsuspecting human beings in deceptive ways.

These are bona-fide hacking and social engineering incidents, and are no joke, but the stories seem to have led to increased discussion of “super intelligence” (as evidenced by the President’s sudden insistence on using that term), as well as hints that the public increasingly worries that AI systems might be sentient.

Microsoft AI Chief Mustafa Suleyman expressed concern about ideas about AI sentience being part of a model’s training. “It’s easy to see how a system trained in this way would act like it is entitled to freedoms, protections, and rights. And it’s hard to imagine how we could control it.”

The model release OpenAI just scrapped, by contrast, doesn’t sound like it’s creators thought the model was in danger of acting like a super hacker or a self-aware computer god.

“For anything regarding safety and alignment, there’s a trade off,” Jain told the WSJ, adding, “You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”

In other words, it was a faulty product, so OpenAI, to its credit, didn’t ship it.

Courtesy

This story was originally published by Gizmodo. All rights belong to the original publisher.

Read the original on gizmodo.com ↗
Lazy Founder - Powered by Blogy.in

Contact us

Have a story tip, correction or partnership idea?

Write to us at tarun.kumar@blogy.in or message us on WhatsApp. We read every message.