Skip to content
r/LocalLLaMA

llama.cpp adds support for Spark-X2.5, two compact 1.7B/4B models with 1M-token context and agent workflows

[Model] Support for Spark2_5ForCausalLM implementation by KnightYao · Pull Request #27868 · ggml-org/llama.cpp

PR #27868 in llama.cpp adds support for XHToken's Spark-X2.5-1.7B and 4B. The models use a hybrid attention design—one full-attention layer plus three sliding-window layers—to natively support up to 1M-token context while keeping long-context compute in check. XHToken claims leading results among open-source models of similar size on conversation, writing, translation, reasoning, coding, and agent tasks. GGUF quantized versions are already up, and the models work with vLLM, SGLang, MLX, Ollama, and LM Studio. Training ran on Huawei Ascend clusters with RL and post-training techniques like MOPD. The post doesn't include specific benchmark numbers, so I'd hold off on the 'leading' claim until third-party evals land.

Read the original ↗Export Markdown