September 25, 2026
AI Daily Journal.

Understand what changes. Find what to use.

Saved 0

DevelopmentNews1 source

Liquid AI says LFM2.5-VL-DSpark speeds inference by up to 3.13x

Liquid AI released LFM2.5-VL-DSpark, an experimental drafter model for its LFM2.5-VL-3B vision-language model. The company says speculative decoding speeds generation without changing output quality, with gains of up to 3.13x on-device and 2.66x on an H100, while end-to-end latency improves by up to 2.62x and 2.27x, respectively. The drafter adds about 280 million parameters, or 8.9% to the 3-billion-parameter model.

The release includes day-one support for llama.cpp, MLX-VLM and SGLang. Liquid AI evaluated six vision tasks, including image and chart question answering, captioning and complex reasoning. The technique accelerates decoding only; image encoding and prefilling remain unchanged, which can limit the overall gain when those stages account for a large share of response time.

Why it matters · editorial interpretation

The technique could reduce generation time for LFM2.5-VL-3B across different execution environments, with day-one support for llama.cpp, MLX-VLM and SGLang. Because it accelerates decoding only, the total benefit may be smaller when image encoding and prefilling account for most of the latency.

Sources

Research

One story, many sources

Coverage of the same story gathered in one item, with a link to each original.

Summary and interpretation apart

What the sources say stays in the summary; editorial context is labeled separately.

Made with AI, with sources

Summaries and translations generated with AI from the original stories, always linked.