New2 developments · 1 source · over 10 days
llama.cpp b11081 release with CUDA 13 and ROCm 10 support
How it developed
- llama.cpp ReleasesWhere it started
llama.cpp b11081 release with CUDA 13 and ROCm 10 support
- llama.cpp Releases
llama.cpp optimizes NVIDIA V100 inference performance
llama.cpp adds a 3% throughput boost for NVIDIA V100 users by routing sm_70 to the Turing MMVQ nwarps table