llama.cpp has released version b11113, expanding hardware support to include ROCm 10.0 for Linux and Windows, alongside CUDA 13 binaries. The update also introduces experimental support for Linux arm64 on Snapdragon platforms, leveraging CPU, Adreno GPU, and Hexagon NPU acceleration with an accompanying setup guide. Existing builds for Apple Silicon, Vulkan, OpenVINO, and SYCL remain available, though KleidiAI support on macOS has been temporarily disabled. This release reinforces the project's role as a cross-platform standard for local AI inference.
Read originalThe llama-server now binds to multiple addresses, a practical upgrade for anyone running local inference behind reverse proxies or complex network setups. This change removes the previous single-address limitation, allowing flexible routing without external workarounds. While the release includes standard binaries for CUDA 13 and ROCm 10.0, the networking feature is the real differentiator here. It makes self-hosted deployments slightly more robust for power users who need granular control over traffic flow.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing full ROCm 10.0 support to both Linux and Windows, closing a long-standing gap for AMD GPU users who previously had to rely on workarounds or older versions. The inclusion of CUDA 13 builds alongside CUDA 12 ensures compatibility with the latest NVIDIA driver stacks without forcing users into beta territory. Perhaps most notably, the addition of native Snapdragon support on Linux marks a significant step toward efficient AI inference on ARM-based mobile and edge devices, expanding the hardware ecosystem beyond traditional x86 and NVIDIA dominance.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing ROCm 10.0 to both Linux and Windows builds. The real surprise is the inclusion of a setup guide for Linux arm64 on Snapdragon devices, bridging the gap between mobile silicon and local LLMs. CUDA 13 support is now standard across platforms, ensuring compatibility with the latest NVIDIA stacks. While KleidiAI on macOS is disabled in this specific build, the broader platform coverage means developers no longer need to compile from source for most hardware configurations.
Claude Code just got its first major model upgrade with Opus 5.5 as the new default, bringing a 1M context window and aggressive pricing that reshapes local inference economics. Beyond the headline model swap, this release quietly stabilizes the background subagent system, fixing critical issues where tool lists were rebuilt instead of cached and reports were silently lost during compaction. The UI layer also sees significant polish, with mouse support in fullscreen mode and fixes for Windows terminal rendering that had plagued power users. This is less about new features and more about making the agent runtime reliable enough for heavy, multi-step workflows.
© GitHub ChangelogOpenAI is widening the GPT-6 lineup inside GitHub Copilot by adding two distinct models: Sol for complex agentic coding and Luna for fast, cheap tasks. This moves Copilot away from a single default toward a tiered strategy where developers can explicitly choose between heavy reasoning and lightweight efficiency. Sol targets multistep validation in Pro+ and Enterprise tiers, while Luna opens up to the broader Pro user base as the lowest-cost option. The gradual rollout across IDEs like VS Code and JetBrains means teams can now tune their coding assistants for specific workload profiles rather than accepting a one-size-fits-all approach.
© GitHub ChangelogGitHub is killing the convenience of a single download for CodeQL CLI. Starting with version 2.27.0, the monolithic all-platform bundle is deprecated and will vanish by March 2027. Users must now switch to platform-specific archives, a move that forces CI/CD pipelines to manage separate artifacts for Linux ARM64, x86_64, and macOS. This ends the era of 'one size fits all' distribution for static analysis tooling on GitHub, requiring developers to explicitly target their build environment's architecture.