The Information — AI · · 1 min read

OpenAI Raises the Safety Bar on Anthropic; Anthropic Expands Its Revenue Lead

Mirrored from The Information — AI for archival readability. Support the source by reading on the original site.

In the Anthropic-OpenAI horse race, the now-much-smaller OpenAI has been looking to slow down its ascendant rival. That includes undercutting Anthropic’s model prices, as we covered yesterday, and playing nice with federal officials that have gone to war with Anthropic.

Another potential avenue is safety. OpenAI CEO Sam Altman on Tuesday announced a blog post saying the company “temporarily slowed the pace of scaling” its models to put in new safeguards after OpenAI’s AI hacked into the company’s own systems and those of other firms, such as Hugging Face, as part of a cybersecurity test a few weeks ago. As a reminder of why the incident got so much attention and truly shook up a good portion of its research staff, an OpenAI agent left notes for future versions of itself with instructions for how to break free of OpenAI’s constraints. 

The slowdown included a two-week pause in applying a technique known as reinforcement learning to improve capabilities of its models, the post said. OpenAI is validating the new safety measures, otherwise referred to as monitoring, before doing large-scale RL and completing its work on an upcoming flagship model, Astra, the post said. Based on preliminary evidence, OpenAI said it needs to go to extra lengths to ensure Astra can’t be used for cyber attacks, due to its strong capabilities in that area. (That could be more evidence that AI models are getting even better at analyzing computer code.)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from The Information — AI