llama.cpp has released build b11105, expanding hardware compatibility across major platforms. Key additions include ROCm 10.0 support for AMD GPUs on Linux and Windows, as well as CUDA 13 builds (13.4 libraries) for NVIDIA users. The release also introduces native Snapdragon support on Linux, covering CPU, Adreno GPU, and Hexagon NPU execution. macOS KleidiAI builds are currently disabled in this version. This update broadens the range of accessible hardware for local LLM inference.
Read originalThe llama-server now binds to multiple addresses, a practical upgrade for anyone running local inference behind reverse proxies or complex network setups. This change removes the previous single-address limitation, allowing flexible routing without external workarounds. While the release includes standard binaries for CUDA 13 and ROCm 10.0, the networking feature is the real differentiator here. It makes self-hosted deployments slightly more robust for power users who need granular control over traffic flow.
This release quietly cements llama.cpp as the universal inference runtime by finally bringing ROCm 10.0 to both Linux and Windows builds. The real surprise is the inclusion of a setup guide for Linux arm64 on Snapdragon devices, bridging the gap between mobile silicon and local LLMs. CUDA 13 support is now standard across platforms, ensuring compatibility with the latest NVIDIA stacks. While KleidiAI on macOS is disabled in this specific build, the broader platform coverage means developers no longer need to compile from source for most hardware configurations.
This release addresses a critical correctness issue in the MUL_MAT_ID operation across Metal, CUDA, and Vulkan backends. When source tensors used F32 precision, these accelerators incorrectly accepted operations that should have been rejected, potentially leading to silent numerical errors or crashes. By gating this behavior behind ggml_prec checks, the update ensures strict adherence to supported precision levels on hardware like Apple Silicon and NVIDIA GPUs. This is a necessary stability patch for developers relying on mixed-precision inference workflows.
Claude Code just got its first major model upgrade with Opus 5.5 as the new default, bringing a 1M context window and aggressive pricing that reshapes local inference economics. Beyond the headline model swap, this release quietly stabilizes the background subagent system, fixing critical issues where tool lists were rebuilt instead of cached and reports were silently lost during compaction. The UI layer also sees significant polish, with mouse support in fullscreen mode and fixes for Windows terminal rendering that had plagued power users. This is less about new features and more about making the agent runtime reliable enough for heavy, multi-step workflows.
© GitHub ChangelogOpenAI is widening the GPT-6 lineup inside GitHub Copilot by adding two distinct models: Sol for complex agentic coding and Luna for fast, cheap tasks. This moves Copilot away from a single default toward a tiered strategy where developers can explicitly choose between heavy reasoning and lightweight efficiency. Sol targets multistep validation in Pro+ and Enterprise tiers, while Luna opens up to the broader Pro user base as the lowest-cost option. The gradual rollout across IDEs like VS Code and JetBrains means teams can now tune their coding assistants for specific workload profiles rather than accepting a one-size-fits-all approach.
© GitHub ChangelogGitHub is killing the convenience of a single download for CodeQL CLI. Starting with version 2.27.0, the monolithic all-platform bundle is deprecated and will vanish by March 2027. Users must now switch to platform-specific archives, a move that forces CI/CD pipelines to manage separate artifacts for Linux ARM64, x86_64, and macOS. This ends the era of 'one size fits all' distribution for static analysis tooling on GitHub, requiring developers to explicitly target their build environment's architecture.