Liquid AI Adds Speculative Decoding to LFM2.5-VL-3B for Faster Vision-Language Inference

Liquid AI released LFM2.5-VL-DSpark, a 280M-parameter drafter that speeds LFM2.5-VL-3B decoding on edge devices and H100 GPUs.

AI News

Liquid AI has released an experimental speculative-decoding model designed to accelerate its LFM2.5-VL-3B vision-language model, targeting a central bottleneck in multimodal inference: generating responses quickly after an image and prompt have been processed.

The model, named LFM2.5-VL-DSpark, adds a 280-million-parameter draft model to the 3-billion-parameter target model. Liquid AI says the combination delivered decoding speedups of up to 3.13x on an Apple M5 Max and up to 2.66x on an Nvidia H100 in its internal evaluation. End-to-end latency gains were smaller, reaching 2.62x on the M5 Max and 2.27x on the H100.

The release matters most for teams deploying vision-language models on local hardware, where response latency and memory limits can be more restrictive than on large inference clusters. It also extends Liquid AI’s DSpark approach from text models into multimodal workloads without requiring a different speculative-decoding algorithm.

A drafter built for multimodal inputs

Speculative decoding uses a smaller model to propose several tokens at once. The larger target model then verifies those proposals, accepting the tokens that match its own next-token distribution and generating replacements where necessary. The approach can reduce the number of expensive target-model steps without changing the target model’s output under exact verification.

According to Liquid AI’s Hugging Face release, the LFM2.5-VL-DSpark drafter takes hidden states from selected layers of LFM2.5-VL-3B and proposes blocks of candidate tokens. Images and text are first projected into a shared representation, allowing the drafter to receive vectors with the same dimensionality regardless of the input modality.

Liquid AI says the vision drafter uses four attention-only layers and a block size of nine during training. For inference, the company recommends a block size of eight or nine depending on the hardware. The drafter adds approximately 8.9% to the deployed model’s parameter count, a relatively small memory increase compared with running a separate large model for the same task.

The model was trained with a mixture of vision-language supervised fine-tuning data weighted toward the workloads Liquid AI expects it to serve. The company tested three-, four-, and five-layer designs and trained the selected configuration for 10 epochs, reporting that acceptance improved with additional tokens before reaching diminishing returns.

Reported gains vary by hardware and task

Liquid AI evaluated the system across six vision-based workloads using the MMSpec benchmark. The tasks included general visual question answering, text-focused visual question answering, image captioning, chart question answering, complex reasoning, and multi-turn conversation.

On-device results were measured with MLX on an M5 Max and with llama.cpp on an M3 Ultra. Liquid AI reports that MLX decoding was 2.30x to 3.13x faster depending on the task, while end-to-end latency improved by 1.56x to 2.62x. With llama.cpp, decoding gains ranged from 1.57x to 2.14x, and end-to-end latency improved from 1.30x to 1.77x.

The H100 evaluation produced decoding gains of up to 2.66x and end-to-end improvements of up to 2.27x, according to the company. The source material includes an inconsistent lower-bound figure for the H100 decoding range, so the broader result should be treated as a vendor-reported maximum rather than a uniform performance expectation.

These are Liquid AI’s measurements, not an independent benchmark. They also describe specific hardware, software configurations, task mixes, and a DSpark block size of eight. Actual gains will depend on the acceptance rate of proposed tokens, prompt length, image complexity, quantization, batch size, and the proportion of total latency spent on image processing and prompt ingestion.

Liquid AI says speculative decoding is exact because the target model verifies every proposed token. In its implementation, greedy output should therefore match the target model running without speculation. That property addresses one of the main deployment concerns around acceleration techniques: improving speed without silently changing model behavior.

Why end-to-end latency remains the harder metric

The reported difference between decoding speed and total latency is important for product teams. Speculative decoding accelerates token generation, but it does not speed up the vision encoder or the prefill phase that processes the image tokens and text prompt.

Vision-language models can spend substantial time before producing the first token. An image is passed through a vision encoder, after which the language model processes hundreds of visual tokens alongside the textual input. On edge hardware, the lower compute budget can make those stages a larger share of total response time. As a result, a threefold decoding improvement does not translate into a threefold reduction in user-visible latency.

This limitation is an example of Amdahl’s law: the parts of a workload that remain unchanged cap the overall gain. For an application such as image chat, document analysis, or chart interpretation, teams will need to measure time to first token and complete response time separately. A faster decode path may be most valuable for long answers, repeated turns, or workflows where output generation dominates after the image has already been encoded.

The results also suggest that hardware choice will shape the value proposition. Liquid AI’s strongest reported decoding range came from Apple silicon, while the H100 delivered lower but still material gains in the company’s test. That makes the release relevant both to local inference and to GPU-backed services, but it does not establish that every deployment will see the same improvement.

Integrations lower the testing barrier

LFM2.5-VL-DSpark is available through integrations for llama.cpp, MLX-VLM, and SGLang. Liquid AI says the draft model is offered on Hugging Face in Safetensors and GGUF formats, giving developers paths for both native and quantized deployment workflows.

The SGLang integration requires a build with DSpark support for LFM2 targets, while llama.cpp and MLX-VLM also require versions containing the relevant implementation changes. In SGLang, operators attach the drafter to the target model and query an OpenAI-compatible endpoint. The block size is read from model configuration, and response timing can expose how many draft tokens were proposed and accepted.

For builders, that integration model is significant because the release does not require replacing the target vision-language model or redesigning the application interface. Teams can compare a baseline deployment with speculative decoding using the same target model and measure acceptance, latency, memory use, and output equivalence. The extra 280 million parameters still carry a memory and loading cost, however, which may matter on smaller edge devices.

What to watch next

The clearest follow-up signal will be independent testing of LFM2.5-VL-DSpark across more image sizes, quantization levels, batch sizes, and production-style prompts. Independent results would help establish whether the reported gains persist outside Liquid AI’s selected six-task evaluation.

Developers should also watch acceptance rates and end-to-end latency rather than relying only on headline decoding multipliers. Support in llama.cpp, MLX-VLM, and SGLang will be important to track as implementations mature, particularly for users deploying on Apple silicon or constrained local hardware.

Further DSpark releases for other vision-language models could indicate whether the architecture generalizes beyond LFM2.5-VL-3B. Conversely, if gains are highly sensitive to model architecture or workload, the method may remain a targeted optimization rather than a broadly portable inference layer.

Creati.ai perspective

Liquid AI’s release is a practical inference update rather than a new capability model. Its main contribution is to show how speculative decoding can be adapted to a multimodal target while preserving exact verification and keeping the additional model relatively small.

The commercial significance will depend on total user-perceived latency, not the maximum decode figure. For teams running vision-language workloads locally, the combination of open weights, existing runtime integrations, and measurable gains could justify testing. But buyers should treat the current performance claims as vendor-reported and validate them against their own images, prompts, hardware, and latency targets.

Ads