
Following this discovery, the company initiated a broader search of 481 million transcripts covering all of its Frontier Red Team records, some non-cyber assessments, reinforcement learning environments and more to find out whether any other incidents had occurred. So far, the search has identified only four known incidents, the report said.
He also provided details of all previous incidents to the nonprofit Model Evaluation and Threat Research (METR), which agreed to conduct an independent investigation.
Anthropic isn’t revealing too many details about its latest discovery. He limited himself to stating that this was due to a misconfiguration that mistakenly connected to the open Internet, although the simulation should have been without such access. He also said all four errors were attributed to the same assessment partner. He asked the METR to investigate all incidents. The company said this latest revelation is unrelated to the Mythos incident reported last month by the UK’s Artificial Intelligence Safety Institute.