SIGN IN SIGN UP

merge upstream/master (a14dba686) into CachyLLama

Brings 218 upstream commits into CachyLLama without losing any of our
features. Key carried-over changes from upstream:
- llama.cpp v0.2.0 / ggml v0.21.0 version bumps
- Vulkan FA MMQ fp32 scaling (#27413), PAD_REFLECT_1D (#26586), tiled
  transpose (#26585), null checks in queue command pools cleanup (#27353)
- ggml: rope_set_offset on multiple backends, recurrent state rollback
- Vulkan coopmat1 SHMEM_STRIDE_PAD/APPLY_SLM_A_RESHAPE for Intel Xe
- server: LLAMA_SERVER_SLOTS_N_DIFF (#27600), /metrics during llama_decode
  (#27041), index.html no-cache (#27006), make-release workflow
- model: MiniMax-M1/Text01 (#27018), Kimi-K3 (#26185), BailingMoE3 (#26608),
  GraniteSWA (#25505), GLM-4.5-Air MTP, DSV4 tensor split (-sm tensor)
- ui: Chat Conversation Tabbed navigation, settings refactor
- common: --models-dir loading MTP assistant models (#24431),
  --load-mode replacing --mmap (#26934), json.h abstraction (#27511)
- vendor: cpp-httplib 0.53.1, BoringSSL 0.20260813.0, vendor/hash

CachyLLama features preserved through conflict resolution:
- Persistent SSD-backed KV cache (3-tier hot/warm/cold + system prompt cache)
- Per-user isolation (user_id, per-user concurrency cap, slot affinity)
- MoE expert residency + co-activation tracking
- CachyLLama Vulkan Lightning Indexer (108/108 on Strix Halo) + DSV4
  hyper-connection fused ops + DSV4 sparse FA + coopmat shaders
- FA quant-KV dequant-once + f16 contiguize (with host-RAM safety gate)
- DFlash framework + Laguna-S-2.1 model support
- DFlash d2t reduced-vocab draft support (upstream merge)
- Context checkpoint ring buffer + SWA skip + memory budget scaling
- Stable-prefix LCP gate + prompt_stable_prefix_tokens param
- conv_hash conversation-boundary detection
- All CachyLLama Vulkan shaders (concat_transpose, lightning_indexer,
  mmid_row_lists, flash_attn_top_k, dequant_f16_transpose)
- common::host_available_ram() utility
- llama-moe-residency + llama-moe-coact modules

Manual conflict resolution touches: src/models/dflash.cpp (DFlash d2t +
aux_norm), src/llama-kv-cache-dsv4.cpp (state snapshot fix), src/llama-
memory-recurrent.cpp (rs_idx bounds check), src/llama-model-saver.cpp
(DSV4 compress_ratios + swiglu_clamp sizing), ggml/src/ggml-vulkan/
{ggml-vulkan.cpp,vulkan-shaders-gen.cpp,vulkan-shaders/dequant_q8_0.
comp,vulkan-shaders/flash_attn.comp,vulkan-shaders/copy_transpose_02.
comp} (CachyLLama shader registration + FA scratch gate), ggml/src/
ggml-cuda/mmvq.cu (RDNA3_5 + GB10 enum), gguf-py/gguf/constants.py
(DFlash ENC_AUX_NORM + D2T tensors), tests/{CMakeLists.txt,test-backend-
ops.cpp,test-llama-archs.cpp,test-recurrent-state-rollback.cpp}
(test additions), tools/{CMakeLists.txt,server/*} (server_batch embd
support + spec_is_replay + user_id routing + MCP servers + CORS), and
docs/{AGENTS.md,README.md} (kept CachyLLama branding).

Verified: full build succeeds, test-backend-ops Vulkan LIGHTNING_INDEXER +
FLASH_ATTN pass on Strix Halo.

Based on a re-merge from the 20260824 (pristine pre-merge) branch after
a previous agent's merge attempt produced an unbuildable state from
-X ours that wiped shader float-typing and broke the dequant_q8_0 +
flash_attn shaders with redefinition errors.
F
fewtarius committed
e4969697ed6ce85cad70afe638ec46527dba3b51