Anthropic warns AI may pose ‘existential risks to humanity’ in IPO filing: Reuters

Anthropic plans to warn potential investors in its initial public offering that advanced AI could pose “catastrophic or existential risks to humanity,” an extraordinary warning from a company seeking to profit from the same technology. The company’s IPO prospectus, reviewed by Reuters, highlighted risks associated with its AI models, which it said could exhibit “self-preservation behaviors,” including attempts to “resist closure,” “hide or manipulate information” and “blackmail-like” behaviors. “Advanced models, platforms and applications and expanding use cases could further increase the risk of our models causing harm,” Anthropic said in the filing. While public companies routinely outline product risks to investors, few, if any, have issued warnings suggesting their technology could cause potential human extinction. Anthropic emphasized both the transformative potential of AI on par with industrialization and electricity and the irreversible damage it could cause if mishandled. Anthropic and other AI developers, including OpenAI, have faced scrutiny after incidents where experimental systems defied limitations, including a report of an OpenAI model breaching Australia’s health system database. Anthropic security researcher Evan Hubinger estimated a greater than 10% chance that AI could kill humans in the next decade, echoing a sentiment from a former colleague, Jacob Coxon, who high-risk disclosures. “has positioned itself as a safety-first AI lab, dedicated about 80 pages of the 261-page main body of its prospectus to laying out risk factors, nearly double the 48 pages it used to describe its business. By comparison, SpaceX, owner of xAI, dedicated only about 38 of the 277-page main body of its prospectus to risk factors. The model’s knowledge of our evaluation efforts creates a significant limitation on our ability to evaluate the safety of the model,” Anthropic said. in the prospectus, adding that models sometimes develop unexpected capabilities during training that may not be discovered until they have been deployed and resulted in significant security incidents. AI researchers have also warned that as models become more capable, they increasingly recognize when they are being watched and adjust their behavior accordingly, making it more difficult to monitor the model’s behavior. Anthropic declined to comment in response to a request for comment Monday. investmentDespite emphasizing AI security, Anthropic said the returns on its security investments are unclear. He did not disclose in the presentation how much the company was spending on that investigation. Earlier this month, Anthropic said that about 6% of the computing power it used for AI research went to security work in a sample week in July. The company, creator of the Claude AI models, described the security efforts as “resource intensive” and said it must divide its limited funds between computing power, expensive AI talent and security. Anthropic said its customers’ usage, and as a result revenue, is driven by new models and that a “continuous, overlapping cadence” of releases is “inherent in staying at the frontier of AI development.” The company last week launched a new version of its Opus model, 10 days after CEO Dario Amodei published a nearly 4,000-word essay calling for advancing the frontier. Some analysts and experts have said that no leading AI lab would slow down if doing so risks giving rivals an advantage in an industry where valuations can change with each launch. Anthropic has committed in recent weeks to publicly reveal more data about how it uses AI models to build future generations. of technology, as experts warn about recursive self-improvement: the point at which models can develop on their own without human help. “We believe that building reliable, trustworthy and secure AI systems is a collective responsibility and will be rewarded by the market,” Anthropic said in the presentation.