AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Accelerating Vision-language Models With LFM2.5-VL-DSpark on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Liquid AI released LFM2.5-VL-DSpark, an experimental 280-million-parameter drafter for speculative decoding with its LFM2.5-VL-3B model. The company reports decoding speedups of up to 3.13x on Apple silicon and 2.66x on an H100, while end-to-end gains are lower; the results have not been independently verified.

Liquid AI has released LFM2.5-VL-DSpark, an experimental 280-million-parameter draft model designed to speed up token generation with its LFM2.5-VL-3B vision-language model, as detailed in the original analysis. The company reports decoding speedups of up to 3.13x on an Apple M5 Max and up to 2.66x on an NVIDIA H100, while saying the method preserves the target model’s output under greedy decoding.

The drafter uses speculative decoding: a smaller model proposes blocks of tokens, then the larger target model checks them. Liquid AI says accepted proposals let the target produce tokens more quickly without changing greedy output. The 280-million-parameter add-on increases the target model’s parameter count by about 8.9%, according to the company.

Liquid AI reports results across six vision-language tasks, including general and text-based visual question answering, image captioning, chart questions, complex reasoning and multi-turn conversation. With MLX on an M5 Max, it measured 2.30x to 3.13x faster decoding and 1.56x to 2.62x lower end-to-end latency. With llama.cpp on an M3 Ultra, the reported ranges were 1.57x to 2.14x for decoding and 1.30x to 1.77x end to end.

On an H100, the company reports end-to-end gains of 1.64x to 2.27x and a maximum decoding speedup of 2.66x. The model is available on Hugging Face in Safetensors and GGUF formats. Liquid AI says the release has support in llama.cpp, MLX-VLM and SGLang, though the source identifies required integration patches for those toolchains.

At a glance
announcementWhen: Released; the source does not give a pu…
The developmentLiquid AI released an experimental speculative-decoding drafter for its LFM2.5-VL-3B vision-language model.
At a glance
announcementWhen: announced September 2026, available now
The developmentLiquid AI announced and released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, available immediately on Hugging Face.

Faster Local Vision-Language Inference

The release targets the time users wait for a vision-language model to generate an answer, particularly when running a model locally. Images add work before generation because they are encoded and represented as visual tokens. A faster decoding stage could make interactions with a 3B-class model feel more responsive on consumer hardware, while the reported parameter increase is relatively small compared with the target model.

The measured end-to-end improvements are smaller than the decoding gains. Liquid AI says speculative decoding speeds the generation phase, not image encoding or prompt processing. Those other stages still contribute to total latency, so users should not expect the maximum decoding multiplier to apply to the full request. Actual results may also depend on hardware, prompts and image resolution.

Compatibility with established inference tools could make the drafter easier to try for developers already using those systems. However, the figures are company-reported benchmarks, and independent measurements would help establish how consistently the speedups hold across workloads.

DSpark Moves Into Multimodal Models

DSpark is Liquid AI’s speculative-decoding approach. The company previously applied it to text-only LFM2.5 models; this release extends the method to a vision-language target. According to the announcement, image patches and text tokens share a representation before the model’s tapped layers, allowing the drafter to work with hidden-state vectors of the same dimensions across both input types.

The drafter uses a four-layer attention-only architecture and a block size of nine, selected through comparisons of three-, four- and five-layer versions, the company says. Its components include a decoder stack, a hidden-state projection and a Markov head. Liquid AI reports that acceptance improved during training before gains diminished over ten epochs.

In speculative decoding, the target model verifies proposed tokens. Under greedy decoding, the company says this verification preserves the target model’s output. The release is labeled experimental, and Liquid AI has not said whether or when it will become a stable product.

“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”

— Liquid AI, in its announcement

Benchmark and Release Questions

The speed figures have not been independently verified, and performance may vary with hardware, prompt content and image resolution. The source material also includes an apparently inconsistent H100 range written as “20.4x to 2.66x”; Liquid AI has not clarified that figure, so it is not used here as a benchmark result.

Other open questions include how well token acceptance holds on vision tasks outside the reported benchmark set, how the drafter affects sampling-based generation, and what its runtime memory cost is beyond its parameter count. Training data details have not been published, and the company has not provided a timeline for moving the model beyond experimental status.

Independent Tests and Tool Support

Developers can download the model and try it with supported inference tools; the source identifies patches required for SGLang, llama.cpp and MLX-VLM. The next useful evidence will be independent benchmarks across hardware and workloads, along with details about runtime memory use and performance on tasks beyond the company’s test set. Liquid AI has not announced a stable-release date.

Key Questions

What is LFM2.5-VL-DSpark?

It is an experimental 280-million-parameter drafter intended to speed up speculative decoding with Liquid AI’s LFM2.5-VL-3B vision-language model.

How much faster is it?

Liquid AI reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an H100. Its reported end-to-end gains are lower and vary by hardware and task.

Does speculative decoding change the model’s answers?

Liquid AI says the target verifies proposed tokens and that output is unchanged under greedy decoding. The source does not establish how the drafter affects sampling-based generation.

Where can developers get the model?

Liquid AI says it is available on Hugging Face in Safetensors and GGUF formats, with integrations or required patches for llama.cpp, MLX-VLM and SGLang.

Are the reported results independently confirmed?

No independent verification is included in the source material. The speed figures are Liquid AI’s own measurements, and results may vary by setup and workload.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meta Takes Down A Critical Video About Meta AI Glasses After Filming At Meta

Meta has taken down a video criticizing its AI glasses following filming activities at its headquarters, raising questions about content moderation and transparency.

Isar Aerospace Launch Into Orbit [Video]

Isar Aerospace’s latest launch marks a significant milestone, with confirmed orbit insertion. Details are still emerging, but the event signals growing European space capabilities.

Apple Documents Confirm Powerful iPhone 18 Pro Specs

Apple’s internal documents reveal detailed specifications for the upcoming iPhone 18 Pro, highlighting significant upgrades in hardware and performance.

Drones 101: How Drones Fly and What They’re Used For

How do drones navigate the skies and transform industries? Discover their incredible capabilities and the future they promise.