Liquid AI says LFM2.5-VL-DSpark speeds inference by up to 3.13x
Liquid AI released LFM2.5-VL-DSpark, an experimental drafter model for its LFM2.5-VL-3B vision-language model. The company says speculative decoding speeds generation without changing output quality, with gains of up to 3.13x on-device and 2.66x on an H100, while end-to-end latency improves by up to 2.62x and 2.27x, respectively. The drafter adds about 280 million parameters, or 8.9% to the 3-billion-parameter model.
The release includes day-one support for llama.cpp, MLX-VLM and SGLang. Liquid AI evaluated six vision tasks, including image and chart question answering, captioning and complex reasoning. The technique accelerates decoding only; image encoding and prefilling remain unchanged, which can limit the overall gain when those stages account for a large share of response time.
Why it matters · editorial interpretation
The technique could reduce generation time for LFM2.5-VL-3B across different execution environments, with day-one support for llama.cpp, MLX-VLM and SGLang. Because it accelerates decoding only, the total benefit may be smaller when image encoding and prefilling account for most of the latency.
Sources
Research