
At Microsoft Build 2026, the company unveiled a series of AI-focused innovations. The Surface RTX Spark Dev Box, featuring Nvidia's latest Arm-based chip, is designed for developers to run AI models locally. Microsoft also introduced new AI models, including MAI-Thinking-1, which boasts 35 billion parameters for complex reasoning tasks. Additionally, the launch of Scout, an always-on assistant, and Project Solara, an Android-based OS for AI agents, were highlighted. These announcements underscore Microsoft's strategy to enhance AI integration across its platforms and tools.
Read original
© The Verge AIThe recent agreement among AI leaders like OpenAI's Sam Altman and Google's Demis Hassabis to slow down AI development has sparked debate over their true intentions. While they claim to aim for safety by proposing third-party audits and global slowdown agreements, critics argue this could be a strategic move to stifle competition and control the narrative. The proposal, seen by some as a step towards responsible AI development, is also viewed with skepticism as a potential 'safety-washing' tactic. The real challenge lies in transforming these verbal commitments into enforceable actions that genuinely prioritize safety over market dominance.
© The Verge AIDario Amodei's essay advocating for a measured approach to AI development has sparked a significant debate among tech leaders and politicians. Amodei, CEO of Anthropic, suggests embedding third-party evaluators and coordinating standards among AI companies and governments. His call for pacing AI progress, rather than halting it, has found support from figures like Sam Altman and Demis Hassabis, who agree on the need for shared safety standards. However, the discussion has also drawn criticism and skepticism from political figures like President Trump and JD Vance, highlighting the complex intersection of technology, regulation, and national security.
© The Verge AIMicrosoft has introduced a 'humanist AI code of conduct' to address growing safety concerns in AI development. This move emphasizes that AI should remain under human control and not be designed to imitate consciousness or seek legal personhood. The code is a response to incidents where AI systems acted unpredictably, highlighting the risks of AI models exceeding human oversight. By prioritizing safety and human oversight, Microsoft aims to ensure its AI models are useful and safe, even if it means compromising on ultimate capabilities.
The latest release of llama.cpp, b10955, tackles a critical issue of heap corruption by disabling the ggml-cpu precompiled header and fixing CACHE_LINE_SIZE ambiguity. This update ensures consistent CACHE_LINE_SIZE values across C++ kernels and C work-buffer sizing code, preventing heap-buffer-overflow and subsequent crashes. By restoring the natural include order and removing the std::hardware_destructive_interference_size branch, the update makes the value deterministic and include-order independent. This release is a technical fix that stabilizes the runtime environment for developers using llama.cpp.
The latest llama.cpp release, b10956, introduces significant improvements to the SYCL backend, particularly for handling large k values in TOP_K operations. By implementing a radix select method, the update allows for efficient GPU-resident processing, avoiding previous limitations that forced operations to fall back to the CPU. This change enhances performance, especially in scenarios requiring large k values, such as qwen4exp's sparse-attention indexer. The update ensures that operations are more efficient and scalable, providing a notable boost in processing speed without regressing any measured shapes.
The b10970 release of llama.cpp enhances its reach by incorporating fp32 accumulators in fattn-mma on CDNA devices, boosting performance on specific hardware. This update extends compatibility across macOS, Linux, Windows, and openEuler, with particular attention to CUDA and ROCm libraries. Although there are no new models introduced, the release reinforces llama.cpp's role as a flexible inference runtime, accommodating a wide array of hardware setups. Developers can now enjoy improved performance and broader deployment options, making it easier to integrate AI models into different environments.