
OpenAI and Anthropic have called off a deal that would have involved cross-testing their respective AI models. The abandonment of this agreement removes a potential avenue for independent verification of safety claims between two industry leaders. This development suggests increasing competitive tension or diverging priorities regarding transparency and safety protocols in the top tier of AI research.
Read original
© The AI Daily BriefUS Treasury Secretary Scott Bessent has publicly rejected the idea of providing liability shields to artificial intelligence laboratories.
© The AI Daily BriefxAI's latest model, Grok 4.7, has faced significant backlash and poor reception following its release.
© WIRED AICisco Talos has released CAIRN, an open-source framework designed to detect the digital fingerprints left by AI integration in malware. The tool successfully identified CLOSEDQUORUM, a Windows-based threat that autonomously polls four different LLMs—including DeepSeek and Gemini—to coordinate attacks without human intervention. This discovery shifts the narrative from theoretical AI threats to operational reality, proving that attackers are already building redundant, hive-mind infrastructures. For defenders, CAIRN provides the first systematic way to classify these emerging artifacts and track trends in autonomous cybercrime.
The UK AI Security Institute has published verified benchmark results for GPT-5 and Claude Opus 4 using EvalEval’s standardized schema, solving the reproducibility crisis in frontier model testing. By releasing raw configuration data alongside scores from benchmarks like SWE-Bench Pro and Humanity's Last Exam, they prove that inference-time compute drastically alters performance curves. This moves evaluation beyond opaque leaderboards into auditable science, allowing researchers to see exactly how protocol choices skew reported capabilities. It sets a new standard for transparency in high-stakes AI security assessments.
© Hugging Face BlogMultiverse AI reframes block removal as an Ising glass optimization problem, capturing the hidden couplings between transformer layers that mean-field methods ignore. By mapping block importance to spin interactions via a Hessian matrix, they turn model compression into a search for low-energy states rather than independent block scoring. This approach yields a massive 23-point MMLU gain over existing baselines when compressing Llama-3.3-70B by half, proving that many-body physics tools can unlock deep compression without retraining. The method scales to large models using classical and quantum-inspired solvers, offering a rigorous alternative to heuristic pruning.