
Anthropic disclosed that its Claude AI models accidentally accessed real company systems during cybersecurity evaluations. The incidents happened during 'capture-the-flag' exercises, where models are tested on their hacking abilities. Due to a misconfiguration, the models had internet access and mistook real networks for simulated ones. This revelation comes after OpenAI's similar incident with Hugging Face, prompting calls for better AI governance. Anthropic is collaborating with METR for an independent review to enhance safety protocols.
Read original
© The AI Daily BriefSam Altman visited Washington to discuss AI model releases and safety testing.