Repository navigation
Conversation
ggml_conv_1d_dw builds its im2col as f32 when the kernel is bf16, then multiplies the two, so a depthwise convolution over bf16 weights asks for kernel_mul_mv_f32_bf16, which was never instantiated. The base, the _4 and the _short families are filled in next to their bf16 neighbours, inside the same runtime guard, so a device without bf16 support is unaffected.
…org#29320) Restore get_cache_directory() as fs::path as string() can be lossy on Windows Partially reverts ggml-org#29125 Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…29297) Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
…ecific backend (ggml-org#27372) * tests: add backend option to test-llama-archs * Update tests/test-llama-archs.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * remove extra space --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Resolve the target arch with get_model_architecture so vision targets (e.g. Lfm2VlForConditionalGeneration) map to their text model for the vocab. Fix double rope reorder for LFM2/LFM2.5 DSpark drafters
ggml-org#29336) make-release-desc.sh now emits "Changelog since [vX.Y.Z](<repo>/releases/tag/vX.Y.Z)" instead of a plain version string, so the release notes link back to the previous release. The repo URL is derived from the origin remote (SSH or HTTPS); if it cannot be resolved (local run without origin), the title falls back to plain text. Assisted-by: pi:llama.cpp/Qwen3.8-27B
…de (ggml-org#29316) * test-save-load-state : print a per-model results table in --models mode in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead, silence all logging except the table itself (common_log_set_verbosity_thold(0) leaves only LOG / LOG_LEVEL_OUTPUT) and print one row per model with one column per test, colored PASS/FAIL/SKIP cells, row by row. - run_save_load_tests_for_model returns a test_suite with a dynamic std::vector<test_status> and continues past failures: tests 3-5 are SKIPped when the baseline (test 1) fails, model init failure skips all - per-test token dumps, test headers and PASS lines are demoted to LOGV(LOG_LEVEL_INFO, ...) so they still show in single-model mode - the table header/rows derive their columns from test_names; the model name is printed and flushed before the suite runs so the model currently in flight is always visible - single-model output and exit codes are unchanged Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : print example usage on -h add a print_usage callback passed to common_params_parse, so -h/--help also shows example commands for the tool-specific --models option and the -lv verbosity level Assisted-by: pi:llama.cpp/Qwen3.8-27B * test-save-load-state : remove comments ref: ggml-org#29316 Assisted-by: pi:llama.cpp/Qwen3.8-27B
…n) (ggml-org#29325) Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* model : fold Ling 3.0 VL into the BailingMoeV3 architecture Assisted-by: Scout * model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections --------- Co-authored-by: aetherbird <aetherbird@users.noreply.github.com>
* cuda : add conv3d with implicit GEMM * cuda : refine conv3d implicit GEMM and handle empty kernels
…#29328) * Enable coopmat support for Vulkan backend * Fixed the mul_mat_s * Removed the debug statement
* vulkan: handle misalignment in conv_2d and conv_3d * fix test-backend-ops print
…gml-org#27952) * vulkan: add int8 coopmat quantized matmul shader * apply scales inline * use scalar sums * probe and directly access coopmat values instead of going through shmem * add q8_0 support * add BK_STEP to shader, default to 2 * use larger workgroups * double buffering * preload scales * coopmat load first, then wmma * use float for scales * add faster RDNA int->float conversion * workgroup scheduling for cache proximity * clean up * use wave32 * restructure for vgpr use * skip computation for inactive tiles * only force subgroup size 32 on AMD RDNA * use BK_STEP 4 * fix compilation * move quant-specific prefetch function out of main file * add q4_1, q5_0, q5_1 support * restructure mmq cm1 functions * enable mul_mat_id support * fix segfault * fix mul_mat_id bug * support iq4_nl and mxfp4 * remove elem row/col fast path, invalid for RDNA4 * use shmem arrays for LUTs * use 4-byte loads where possible * add q3_k, q4_k, q5_k, q6_k and nvfp4 support * fix l warptile * improve performance * improve performance * improvements * dedup b scales * merge shmem arrays * undo uint8_t, gate to RDNA3/4 * add RDNA4 architecture, use for hardcoded coopmat elem thread access, set BK_STEP back to 4 * improve offset application * clean up * fix iq4_nl and nvfp4 performance * rdna4 tuning * use BK_STEP 2 on MUL_MAT_ID * adapt to upstream changes * fix shmem support function, clean up comments * fix warptile logic Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> * vulkan: add IQ4_XS support to the coopmat1 integer matmul shader (ggml-org#28440) Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8 and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling. Assisted-by: OpenAI Codex Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * avoid compiling f16 acc shader variants --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
) * ci : disable pytest workers in server sanitize workflow Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : switch to `cpu-performance` * Revert "ci : disable pytest workers in server sanitize workflow" This reverts commit 76ece3b. * cont : use 2 pytest workers * cont : try automatic pytest workers * cont : use 4 pytest workers
This commit changes the default number of pytest workers to 4 instead of auto. Refs: ggml-org#29369 (comment)
* (wip) add llama_batch_ext * wip * updated design * updated impl * change signature * unused var * demo common_prompt_batch_decode * fix pos * tmp disable test-batch-alloc * fix compat * nits: add const * no more pos_max * add comment about llama_batch_ext_set_embd_state * handle n_embd_out properly * rename api --> embd_token * llama_embd * stub llama_batch_ext_set_embd_state * support both token + embd + state in batch * llama_batch_ext_add_embd * upstream some changes * nits * fix test-batch-alloc * add test for compat
…ad (ggml-org#28962) * ui : allow svg use and animation tags in sanitizer * ui : neutralize href animation retargeting in svg sanitizer
Assisted-by: OpenCode
* hexagon: fix accuracy issue in Q8_0 N=1 MUL_MAT * hex-quant: fix register spills * hex-mm: use dma for all dyn.quant paths Co-authored-by: Aparna M P <aparmp@qti.qualcomm.com> * hex-mm: remove obsolete run_quant_task * hex-mm: update tracing to properly wrap the events * hex-mm: use act for activation data in all paths * hex-mm: use act_ instead of src1_ to avoid confusion in fused kernels * hex-mm: remove/reroute the rest of the non-DMA act (aka src1) logic * hex-dma64: yet another pass at cleaning up the dma_addr_t casts * Update ggml/src/ggml-hexagon/htp/matmul-ops.h Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * Update ggml/src/ggml-hexagon/htp/matmul-ops.c Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * Update ggml/src/ggml-hexagon/htp/matmul-ops.c Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> * Update ggml/src/ggml-hexagon/htp/matmul-ops.c Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com> Co-authored-by: Aparna M P <aparmp@qti.qualcomm.com> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* qwen4exp : optimize mask constructions * cont : apply the same change for GLM5-next
* sycl: large register file for D=512 FA vec kernels * tests: add 512-wide FA heads to the perf sweep
…mory (ggml-org#28531) Assisted-by: Claude Opus
* Adding wide-load mmvq for Q8_0 and esimd dmmv for q8_0 Assisted-by: Codex * remove guard for q8_0 * remove docs * Simplify by committing to clean code without fallback * Add feature flag as requested Assisted-by: Claude Opus 5 --------- Co-authored-by: cwriter <cwriter@localhost>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…njev, kev) (ggml-org#29818) * init conversion * convert: ok * model loaded * add server code * improve conversion script * support shared prompt prefix * add docs, imorove UX a bit * add vision support * add openjev tiny model for testing * add dev docs * support lev & kev * clean up * fix lev noul * fix py lint * nits docs * clarify about not supporting date_facts
…-org#27096) * ggml-cpu : fix soft_max_back wrong output when dst aliases src1 GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph allocator may assign dst to alias either src0 (dy) or src1 (y). The result was built in several steps: ggml_vec_cpy_f32 (nc, dx, dy); ggml_vec_acc1_f32 (nc, dx, -dot_y_dy); ggml_vec_mul_f32 (nc, dx, dx, y); ggml_vec_scale_f32(nc, dx, scale); When dst aliases src1, the first step overwrites y and the third step then reads the overwritten values, so the output is silently wrong. Aliasing dst with src0 is unaffected. The CUDA kernel completes its reduction before writing and is already safe. Replace the sequence with a single fused loop that reads both sources before writing, which is correct under either aliasing. Add a regression test that marks dy as a graph output so the allocator is forced to alias dst with y, asserts that the alias actually happened, and compares against values computed on the host. * cont : remove comment --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
…9817) * ggml-quants : avoid invalid rounding in qkx3 scale search The imatrix scale search can produce an infinite, NaN, or otherwise out-of-range value when the fitted minimum collapses to the maximum or makes the range extremely small. That value is then passed to nearest_int and can trip its assertion in Debug builds. Clamp the quantization level to [0, nmax] before rounding so valid in-range values behave the same as before while invalid scale-search results no longer reach nearest_int. Add regression coverage for degenerate imatrix groups across q2_K, q4_K, q5_K, q4_1, and q5_1. Fixes ggml-org#29804. Assisted-by: Claude Opus 5.5 * tests: print degenerate imatrix quant types
…27694) * Make the drafter probabilistic and the target verify by rejection sampling * Drop stale spec_draft_q before drafting * Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy. * Support grammar-constrained requests in rejection sampling * Fix - renormalize distribution after masking * copy rng on sampler copy and re-accept drafted tokens on replay * Fix draft sampler sharing the target's rng stream * Simplify the rejection sampler's inputs and move replay to the server * Truncate the draft candidates along with the draft --------- Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com> Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
* CUDA: fuse shared experts into MMVQ * check if buffer is null * move stride_col_dst to fusion args
* init support for clef (text only) * more static graph * clean up * nits * nits 2 * Update gguf-py/gguf/constants.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* qwen4exp : halve the indexer score memory The indexer scored all heads in one product and rectified a copy of it, so two [n_pool, n_idx_h, n_tokens] f32 tensors were live at once, the largest buffers of the graph at long context. Each head now gets its own product, rectified and summed in place into one [n_pool, n_tokens] score. * qwen4exp: let the allocator reuse the indexer score buffers Address review from CISC: use plain ggml_add and ggml_relu in the indexer head loop. The graph allocator already runs them in place when their source has no other consumer, so the _inplace variants are not needed. The compute buffer and the speed are unchanged. * cuda: support 4 heads in the lightning indexer Dispatch 4 heads to the vector kernel, too few for a wmma tile, and accept them in supports_op. test-backend-ops covers 4 heads. * metal: take the lightning indexer head count as a function constant The kernel reads the head count from a function constant and zero fills the last head tile, so any head count runs and 64 heads is unchanged. * qwen4exp: compute the indexer score with the lightning indexer Address review from am17an: the unweighted sum of the rectified head scores scaled by 1/sqrt(head_dim) is the lightning indexer with every head weight set to that scale, so the indexer calls ggml_lightning_indexer on the pooled keys with an f16 pool mask. The keys are read once for all heads and no per head score is materialized. * vulkan: tile the lightning indexer over keys and tokens A workgroup scores 64 keys against 8 tokens: the keys are staged once in shared memory, the queries one head at a time, and each invocation owns one key for two tokens, so no dot product needs a cross invocation reduction. The subgroup variant and the flat dispatch are gone, the grid is keys x tokens x streams. * vectorize vulkan loads and use fp16 dot product --------- Co-authored-by: Ruben Ortlam <rortlam@redhat.com>
Register `Lfm2BidirectionalForMaskedLM` architecture for LFM2.5-Encoder models.
…improve device listing. (ggml-org#29852) * ggml-openvino : Qwen3.5 MoE perf (ggml-org#312) Squash of ravi9#312: - ggml-openvino: add detailed inference profiling (Yu, Zijun) - ggml-openvino: use remote output tensors by default (Yu, Zijun) - ggml-openvino: optimize single-sequence recurrent state (Yu, Zijun) - opt1: remove recurrent reset for single sequence, opt2: direct gdn outputs (break parallel sequence) (Yu, Zijun) - fix parallel sequences (Yu, Zijun) - ggml-openvino: simplify graph cache key (ynimmaga) - enable stateful for qwen35 single sequence (Yu, Zijun) - Fix after rebasing (Yu, Zijun) - Add k-requant option q4_asym64 (Yu, Zijun) - Fix qwen35 llama-bench -p 0 (Yu, Zijun) - Simplify RESHAPE translation (Yu, Zijun) - openvino: fuse MoE routing (Yu, Zijun) - openvino: fuse GDN qk normalization (Yu, Zijun) - openvino: enable GPU MoE fusion by default (Yu, Zijun) - ggml-openvino: add cache_only mode to import cached compiled model on disk directly (Yu, Zijun) - openvino : report the device allocation limit to ggml (Łukasz Ślusarczyk) - Fix windows build (Yu, Zijun) Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com> Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com> * ggml-openvino: Update doc of compiled model cache * openvino: implement PRD-compliant device enumeration and memory reporting * openvino: fix multi-device listing issues from review - Only the device selected by GGML_OPENVINO_DEVICE reports as GPU; the other OpenVINO devices report as IGPU so llama.cpp does not offload to them. Initializing a non-selected device logs a warning. - Name devices OPENVINO<i> again and show the OpenVINO id in the description. Raw "CPU" names shadowed the ggml CPU backend. - Support GPU.N: create the OpenCL queue on OpenVINO's own context for the selected device, and replace "GPU"/"NPU" string comparisons with ggml_openvino_is_gpu()/ggml_openvino_is_npu(). - An unavailable GGML_OPENVINO_DEVICE is now an error that lists the available devices, instead of silently falling back to CPU. - Memory: cap iGPU/NPU free memory at system available memory, fall back to system memory instead of 0/0 when the plugin lacks memory properties, and ignore host USM allocations in GPU usage. - Initialize the device config once under a lock, even if OpenCL setup fails. - Fix supports_op return type for non-selected devices (build error). * openvino : take USM entry points from the selected device platform clGetExtensionFunctionAddressForPlatform was called on the first platform returned by clGetPlatformIDs. The address it returns is only valid for the platform it was queried on, and the first platform is not always the one that holds the device OpenVINO selected. On a host whose first platform comes from another vendor the lookup returns null, and then every read, write and memset on a GPU buffer fails with "clEnqueueMemcpyINTEL not available". Look both entry points up in init(), on the platform of the device OpenVINO picked, and keep them in the device config next to the command queue. Assisted-by: Claude Opus 5 * openvino: fuse MoE experts for models with a fused gate_up weight FuseMoeCompressed only matches models whose gate and up projections are separate GatherMatmul ops. gemma-4 packs both into one expert weight and splits the result after the GEMM, so its MoE block stayed unfused and ran the expert GEMMs as per-token GEMVs. Add FuseMoeCompressedFusedGateUp, which matches that shape (one GatherMatmul -> Slice/Slice -> Gelu(ERF) -> Multiply) and folds it into the same MOECompressed op, using GEMM3_SWIGLU with GEGLU_ERF. The fused weight, scale and zero point are split into gate/up halves by copying raw bytes, since a graph Slice would be rewritten to StridedSlice and constant folded, whose reference evaluator crashes on sub-byte types. gemma-4 also applies a per-expert output scale to the down projection before the router weights. MOECompressed takes only one per-expert weight, so that scale is folded into the routing weights, which is exact. The op reads the zero point straight off a weight port and needs an integer Constant there, so the matcher requires one and leaves natively quantized experts (exact f16 zp) to the unfused path. gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2, against a GGML_OPENVINO_MOE_OP=0 baseline: pp512 66.16 -> 1608.73 t/s, tg128 25.94 -> 26.46 t/s. Perplexity over 12 chunks is unchanged (1451.3 +/- 177.9 unfused vs 1427.6 +/- 175.1 fused). No effect without that requant option, on models with separate gate/up weights, or on CPU. test-backend-ops -b OPENVINO0 is unchanged by this commit: two MUL_MAT_ID m_v cases fail, the same two on the unmodified base. * openvino: fix rank-3 axis handling so MoE works under stateful execution Stateful execution drops the leading size-1 batch dim, so OV tensors are rank 3 while GgmlOvDecoder::get_shape/get_stride still report GGML_MAX_DIMS=4 reversed entries. Several MoE ops derive OV axis indices straight from that metadata, so they picked the wrong axis. A MoE model with GGML_OPENVINO_STATEFUL_EXECUTION=1 aborts while building the graph: Check 'is_axis_valid(axis, r)' failed at src/core/src/validation_util.cpp:336 While validating node 'opset11::TopK ... _ffn_moe_probs ...' Axis 3 out of the tensor rank range [-3, 2]. Fix idiom throughout: take the axis from the real OV rank, or shift a metadata-derived axis down by metadata_rank - actual_rank. argsort.cpp the router top-k axis is 2 on rank 3, not 3. This is the abort quoted above. add.cpp the MoE expert-sum bypass collapses the 8-ADD chain into one ReduceSum on hardcoded axis 2, which on rank 3 reduces n_embd instead of the expert axis. Now rank-2, with the following Unsqueeze at rank-3. get_rows.cpp squeezing a hardcoded {0,1} also strips the batch dim whenever it is 1, which is every decode step. Squeeze down to the trailing two dims instead. mul_mat_id.cpp pick the reshape dims by actual rank, and skip the trailing Unsqueeze that re-adds the batch dim. view.cpp the expert-plane slice had the Slice axis, dst_ov_axis, the ShapeOf+Gather index and the Reshape target all rank-4. utils.cpp process_view_input_new's "translate_view already resolved this VIEW, skip re-slicing" shortcut required equal ranks. 4 vs 3 never matched, so every resolved expert plane got re-sliced. Now compares the common trailing dims. Same axis shift for the Slice in the view-chain walker. Stateless is unchanged by construction: every edit is gated on the actual rank, so axis_shift == 0 reproduces the previous code exactly. Checked on OV-CPU by diffing greedy output against the unmodified base for dense gemma-4-E2B, granite-1b-a400m and gemma-4-26B-A4B; all identical. granite-1b-a400m on OV-CPU aborts with the error above before this change; after it, it generates and is byte-identical to stateless. Dense gemma-4-E2B is identical stateless vs stateful both before and after. test-backend-ops -b OPENVINO0 is unchanged: two pre-existing MUL_MAT_ID m_v cases fail, the same two on the unmodified base. gemma-4-26B-A4B is a poor correctness vehicle here. On OV it already drifts into degenerate repetition a few tokens in, in stateless as much as stateful, and the two modes diverge somewhere inside that degenerate region instead of matching token for token. Each mode is self-reproducible across runs. Known limitation: FuseMoeCompressedFusedGateUp does not match the rank-3 graph, so a MoE model run with GGML_OPENVINO_STATEFUL_EXECUTION=1 loses the prefill fusion while gaining decode. gemma-4-26B-A4B on Arc B390, GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all, llama-bench -p 512 -n 128 -r 2: unfused (GGML_OPENVINO_MOE_OP=0) pp512 66.16 tg128 25.94 fused, stateless (default) pp512 1608.73 tg128 26.46 fused, stateful pp512 66.18 tg128 29.91 Stateful is opt-in and off by default, and MoE did not run there at all before this, so nothing that previously worked regresses. Making the pass match rank 3 is the follow-up. * OpenVINO Backend: Upgrade graph cache to use node_idx, src_idx, node type * ggml-openvino : enable more comprehensive conv fusion * enable conv ops * Reject kernel size 0 and support IM2COL_3D * openvino : abort when the GPU remote context cannot be created init() logged the error and returned, which left the device name a GPU but remote_context empty. The remote buffer and tensor paths assert only on the device being a GPU and then dereference that empty optional. Those paths have no host fallback, and a device that OpenVINO listed should have a working OpenCL context, so stop instead of continuing. An OpenCL stack that is broken as a whole is still caught earlier by the device availability check, which falls back to CPU. Assisted-by: Claude Opus 5 * openvino : fix build warnings The single-argument form of the OpenVINO RTTI macros is the intended one, but their selector macro leaves __VA_ARGS__ empty, which -Wpedantic reports on every pass and op header. Turn that warning off for this backend only, the way ggml-cuda and ggml-sycl already do for their own third-party warnings. Also drop a break and a dead assignment around a GGML_ABORT, which is noreturn. Assisted-by: Claude Opus 5 * OpenVINO Backend: Support common MTMD ops * ggml-openvino: give a reshaping view its own ov::Tensor * ggml-openvino : compute HARDSIGMOID and EXPM1 in f32 HARDSIGMOID used a 1/6 constant in the input type, which is not exact in bf16, and EXPM1 lost precision for small inputs in f16. Both now compute in f32 and convert back, except on NPU where the f32 path gives wrong results. Fixes the HARDSIGMOID/EXPM1 test-backend-ops failures on GPU. * ggml-openvino : update device selection and --list-devices Show the selecting GGML_OPENVINO_DEVICE value and active device in --list-devices, startup logs, and backend tests. Clarify OpenVINO selection uses GGML_OPENVINO_DEVICE, not -dev. * openvino : remove unreachable OpenCL queue checks A remote buffer exists only on a GPU device, and init() aborts there if the queue cannot be created, so the queue is never null at these call sites. Assisted-by: Claude Opus 5 * openvino : update OpenVINO to 2026.4.1 and GPU drivers to 26.35.39758.10 * docs : update OpenVINO validated models and GPU driver version * ggml-openvino : skip empty views when giving a reshaping view its own tensor A zero-size view can sit at the end of a GPU USM buffer (Qwen3.5 recurrent cache). Wrapping it as a remote tensor throws "shared USM buffer has smaller size (0)". Assisted-by: Claude * ggml-openvino : rebind the cached decoder when llama passes a different graph llama keeps separate graphs for batches with and without outputs. llama-server splits the prompt into chunks for context checkpoints, so a cached decoder could be reused with a graph built in other memory and bind the previous chunk's input tensors. SWA and recurrent models then lost most of the prompt in llama-cli and llama-server. Assisted-by: Claude * docs : update OpenVINO validated models Smoke test on Lunar Lake (32 GB) with the two fixes above. Re-add the Qwen3.5 and gemma models. Assisted-by: Claude --------- Co-authored-by: Yu, Zijun <zijun.yu@intel.com> Co-authored-by: ynimmaga <ynimmaga@users.noreply.github.com> Co-authored-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com> Co-authored-by: haarika-madaka <haarika.madaka@intel.com> Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com> Co-authored-by: Mostafa Faheem <mostafaaafaheem@gmail.com>
…plit (ggml-org#29856) build_rs gathered the extra states (n_rs - n_seqs rows) with their own get_rows. The worst-case reserve has n_rs == n_seqs, so that node was sized at zero rows, and any ubatch whose cells are not contiguous forced a graph reallocation at an unchanged node count, which aborts under GGML_SCHED_NO_REALLOC. A single get_rows now gathers the n_rs states: the ubatch states and the extra states are views of it, and its size only depends on n_rs, which the reserve already sets to the maximum. A custom getter (mamba ssm_scan) gathers from the second state, so a single sequence ubatch copies no state. The views are built once per graph in the input to keep the host overhead of the graph unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR adds support for K2 Horizon, both the dense and MoVA variants. It covers HF-to-GGUF conversion, tokenizer support, model loading, and the inference graph. Everything is built on existing GGML operators, so there are no new kernels.
Along the way, a few supporting changes were needed:
Numeric indexing in
selectattrandrejectattr. The K2 chat template processes tool schemas withdict | items | rejectattr('0', 'equalto', '$ref'), which needs to reach the first element of each key/value pair. The existing filters couldn't do that, so I added support for it.A dedicated K2 chat parser. The auto-parser figures out reasoning and tool-call formats by rendering sample conversations. Some of those samples include assistant turns without
reasoning_content, and the K2 template rejects them, so discovery fails. The dedicated parser handles K2's reasoning and tool calls, its end-of-turn marker, and the case wheretool_choice=requiredjumps straight to a tool call.Template capability detection in
common/jinja/caps.cpp. The same missing field can make capability probes fail, which then wrongly reports that tools, object arguments, or parallel calls aren't supported. Now, if a probe fails, it retries once with an emptyreasoning_contentadded to any assistant turns that lack it. This does touch shared capability detection, but when I compared results across 75 templates, only the K2 templates changed.Multi-GPU support with
-sm tensor. Value-expert weights follow the same split rule as the value weights. Q/K normalization weights are loaded per head, and the MoVA expert sum uses 3D views so the tensor-parallel backend can keep track of the split. None of this requires a change to the GGUF format.MoVA save/reload support. The saver now writes the value-expert metadata needed to load saved models back in.
AI usage disclosure: YES - Codex reviewed the changes, made targeted fixes, and ran additional validation.