Fixes per-request reasoning budget handling by ensuring caller-supplied values reach the sampling layer instead of defaulting to server sett
llama.cpp b10003 enhances tokenization via unified argument parsing, Windows UTF-8 support, and exposed model-sourcing flags