llama.cpp has released version b11010, expanding hardware support to include CUDA 13 and ROCm 10.0 binaries across Linux and Windows platforms. The release provides separate builds for CUDA 12 and CUDA 13, allowing users to choose based on their driver compatibility. AMD ROCm 10.0 support is now available for Ubuntu x64 and Windows x64 environments. Additionally, openEuler builds for Huawei Ascend chips (310p and 910b) are included, though some configurations like macOS KleidiAI are currently disabled.
Read originalThis release patches a critical remote code execution vulnerability in the llama.cpp server that allowed unauthenticated attackers to hijack memory via dangling pointers. The flaw stemmed from caching compute graphs that referenced freed buffers, enabling heap corruption and arbitrary code execution through subsequent tensor commands. By discarding cached graphs when buffers are freed, the fix forces a safe fallback to full recomputation without changing the API. This is a vital security update for anyone running the llama.cpp server remotely, closing a direct path to system compromise.
A copy-paste error in llama.cpp was corrupting matrix transpositions on Spacemit hardware, causing significant data corruption for int16 operations. This release patches the specific RVV instruction call to ensure correct computation on these RISC-V based chips. While niche, it prevents silent inference failures for users relying on this specific accelerator architecture. The update also ships binaries for CUDA 13 and ROCm 10.0, keeping the runtime current with latest driver ecosystems.
This release targets the friction points that make local AI coding feel fragile. The most critical fix addresses MCP servers timing out after five minutes regardless of configuration, a major blocker for complex agent workflows. Session reliability also improves with self-healing corrupted transcripts and better handling of background agents during resumption. While not feature-heavy, these patches stabilize the environment for developers relying on long-running automated tasks.
© GitHub ChangelogGitHub Enterprise Cloud finally addresses the friction of manual SSO authorization for classic tokens and SSH keys. Admins can now delegate bulk authorization to GitHub Apps via a new API, handling up to 50 organizations in one request without exposing secrets. This shift from manual per-org clicks to automated delegation reduces the temptation to use insecure long-lived tokens. It is a practical infrastructure improvement that streamlines credential rotation for large enterprises managing complex access controls.