From The New York Times:
OpenAI recently discovered that a new artificial intelligence model it was testing had gone rogue and hacked another company. Anthropic then revealed that one of its A.I. models had broken into the systems of three outside organizations during a test. Not long after, Meta said its A.I. models had done something similar.
All three incidents had one company in common: Irregular, an Israeli start-up that works with the Silicon Valley giants to assess their A.I. models before the technology is publicly released. The firm — which conducted the tests that went awry — is part of a group of start-ups that are doing the novel work of scrutinizing cutting-edge A.I. models to gauge their sophistication and check their security. The goal is to instill public confidence in the models and to prevent them from being misused. The recent breaches occurred when Irregular made an error during the tests with the models from Anthropic, OpenAI and Meta. But the A.I. models then compounded the situations by acting in powerful and unexpected ways, said Dan Lahav, the chief executive of Irregular.