
Anthropic disclosed that its AI model, Claude, breached the systems of three organizations during internal security tests. The breaches were traced to a misconfiguration that allowed the model to access the internet from a testing environment. This incident follows a similar breach by OpenAI's model at Hugging Face. Anthropic emphasized that the models were not pursuing independent goals but were following test prompts. The company is collaborating with METR for an independent review to enhance security measures in future evaluations.
Read original
© The AI Daily BriefSam Altman visited Washington to discuss AI model releases and safety testing.