top of page

Liquid AI LFM2.5-VL-3B-DSpark Speeds Up Vision Decoding, but the 3.13x Headline Tells Only Half the Story

Sep 27
12 min read

Liquid AI released LFM2.5-VL-3B-DSpark with a headline result of up to 3.13x faster decoding for its compact vision-language model. The improvement targets token generation, not the entire process of understanding an image and producing an answer.

That distinction defines the significance of Liquid AI LFM2.5-VL-3B-DSpark. The experimental draft model shows that speculative decoding can work across visual and textual tasks on both datacenter and consumer hardware. However, Liquid AI’s own results put the best end-to-end gain at 2.62x, below the peak decoding figure.

The release pressures developers to reconsider how they optimize local vision-language applications. Model compression and smaller architectures are no longer the only routes to lower latency. A separate drafter can accelerate an existing target model while preserving its output behavior under matched decoding settings.

The comparison is therefore not Liquid AI against one rival model. It is speculative decoding against the conventional practice of making the main vision-language model smaller, simpler, or less accurate to gain speed.

Liquid AI LFM2.5-VL-3B-DSpark Adds a Dedicated Draft Model

The release separates answer quality from one important part of the latency problem.

Liquid AI LFM2.5-VL-3B-DSpark is a 279.5-million-parameter draft model built specifically for LFM2.5-VL-3B. It does not replace the 3-billion-parameter target model or answer requests independently. Instead, it proposes several likely tokens before the larger model verifies them together.

This process is called speculative decoding, a method that uses a smaller predictor to draft tokens for parallel verification by the target model. Accepted tokens reduce the number of expensive target-model passes needed to generate an answer. Rejected proposals are corrected by the target.

Liquid AI says the drafter contains four full-attention layers, a Markov head, and a confidence head. Its training block contains nine proposed tokens. Deployment uses blocks of eight or nine, depending on the hardware and inference framework.

The company published the model in Safetensors and GGUF formats through its model repository. It supports SGLang on Nvidia GPUs, MLX-VLM on Apple silicon, and llama.cpp through the GGUF checkpoint.

That framework coverage matters because inference acceleration often remains confined to a paper or custom research implementation. Here, Liquid AI has connected its drafter to three deployment paths already used for servers, Macs, and local quantized models.

SGLang requires version 0.5.19 or newer for the published configuration. MLX-VLM requires version 0.7.2 or newer and currently runs this DSpark implementation with greedy sampling. Liquid AI instructs MLX users to set temperature to zero.

The target model arrived before the drafter. Liquid AI introduced LFM2.5-VL-3B in August 2026 as an open-weight vision-language model aimed at edge deployment. The model combines a language backbone with a vision encoder for images, documents, grounding, optical character recognition, and visual tool use.

Liquid AI previously claimed the target could decode 228 tokens per second on an M5 Max. It also reported 116 tokens per second on a Ryzen AI Max+ 395 and 20 on a Galaxy S26 Ultra. Those earlier figures came from the company’s own testing.

The DSpark release changes the deployment package rather than the underlying model’s learned capabilities. Developers attach the drafter during inference, while LFM2.5-VL-3B remains responsible for approving every generated token.

Under greedy decoding, the target accepts a draft token only when it matches the token that the target would have selected itself. The resulting answer should therefore match ordinary greedy generation. At nonzero temperatures, matched sampling can preserve the target model’s output distribution instead of a single deterministic sequence.

That property gives speculative decoding a different value proposition from quantization or distillation. Those techniques can alter numerical precision, model size, or learned behavior. DSpark instead tries to reduce the time spent reaching the target model’s original decisions.

This is why the release creates a meaningful contest between two optimization routes. Developers can reduce the work performed by the main model, or they can predict more of its work and verify those predictions efficiently.

The 3.13x Result Depends on Hardware, Workload, and Measurement

Liquid AI’s highest number describes one decoding result, not a universal application speedup.

Liquid AI evaluated the drafter across six categories from MMSpec: general visual question answering, text recognition, image captioning, chart analysis, complex reasoning, and multi-turn conversation. The company tested batch size one across Apple hardware and one H100 GPU.

On an M5 Max using MLX-VLM, Liquid AI reported decoding gains ranging from 2.30x to 3.13x. End-to-end improvements ranged from 1.56x to 2.62x. The 3.13x peak occurred on the COCO image-captioning workload.

The company also tested an M3 Ultra with llama.cpp. Decoding improved by a reported 1.57x to 2.14x, while end-to-end performance improved by 1.30x to 1.77x.

On a single H100 80GB GPU with SGLang, Liquid AI measured decoding gains between 2.04x and 2.66x. End-to-end improvements ranged from 1.64x to 2.27x. The company’s detailed benchmark disclosure provides configurations and task-level results.

These tests used 16-bit processing for the vision encoder and language backbone. The H100 configuration used BF16, a draft block of nine, batch size one, and temperature zero. Apple testing used FP16, a block of eight, and up to 2,048 generated tokens.

The median answer length in the Apple evaluation was 90 tokens. That detail is important because output length changes how much of a request is spent decoding. A system producing a long description gives the drafter more time to recover its setup overhead.

Short answers create a less favorable balance. If an application returns a label, a coordinate, or one sentence, image encoding and prompt processing can dominate. Faster generation then affects a smaller share of the total wait.

Draft acceptance helps explain the reported acceleration. Liquid AI measured roughly 3.2 to 4.5 accepted tokens per verification pass across the Apple stacks. Its H100 results ranged from 3.46 to 4.57 accepted tokens.

A higher acceptance length means the target validates more useful output during each pass. Yet acceptance does not translate directly into equal speed gains. Drafter execution, synchronization, memory access, and framework overhead still consume time.

The results also vary by task. Image captioning produced the peak MLX decoding improvement, while multi-turn conversation recorded 2.30x. End-to-end gains were 2.59x and 1.91x, respectively.

That variation prevents a responsible reading of “up to 3.13x” as an expected result for every visual assistant. It is a ceiling observed in one published setup. An application’s result will depend on its hardware, prompt length, output length, sampling settings, and visual workload.

Liquid AI says DSpark retained an advantage at higher concurrency on its H100 tests. However, the gap narrowed as concurrency increased. This suggests the drafter’s relative benefit changes when the GPU shifts from memory-bound decoding toward compute-bound execution.

For product teams, the practical question is not whether the peak figure is real within Liquid AI’s test. The question is whether their latency profile resembles the test that produced it. That requires measuring each inference phase rather than copying a headline multiplier into capacity plans.

How Liquid AI Speculative Decoding Preserves the Target Model

DSpark tries to draft farther ahead without turning prediction errors into final output.

Standard autoregressive generation produces one token after another. Every new token requires another target-model pass, even when the continuation is highly predictable. This serial structure can leave hardware underused during memory-bound decoding.

Speculative decoding inserts a smaller model into that loop. The drafter proposes a block of future tokens, and the target evaluates those positions together. Work is saved when several proposals survive verification.

The challenge is producing proposals quickly enough and accurately enough to justify the extra model. A weak drafter creates rejected tokens. A heavy drafter predicts well but consumes too much time producing its block.

DSpark combines parallel generation with a lightweight sequential component. Its Markov head introduces limited dependency between proposed positions, while the confidence head estimates whether later proposals should be verified. This design aims to preserve block coherence without making drafting fully autoregressive.

The underlying DSpark research describes the approach as confidence-scheduled speculative decoding with semi-autoregressive generation. Its central tradeoff concerns proposal quality, draft latency, and the number of tokens sent for verification.

Pure parallel drafting can generate a block quickly, but accuracy often falls for tokens farther into that block. Every position depends on a context that includes earlier guesses. Errors can therefore compound across the proposal.

A fully autoregressive drafter maintains stronger dependencies but recreates part of the serial bottleneck. DSpark’s hybrid structure tries to occupy the middle ground. It adds a small sequential head after the parallel draft operation.

The confidence mechanism addresses another source of waste. Verifying an entire fixed block makes little sense when the drafter expects its later tokens to fail. A scheduler can shorten the submitted prefix before low-confidence positions consume target capacity.

The published LFM2.5-VL-3B-DSpark configuration uses a rank-256 Markov head and a separate confidence head. Its vocabulary contains 128,000 tokens. The architecture remains tied to its designated target model and cannot serve as a general drop-in drafter for every VLM.

That model-specific relationship is both a strength and a limitation. Training against one target can improve proposal alignment. However, a team switching to another target model needs a compatible checkpoint, training process, and runtime integration.

The “lossless” description also needs precise interpretation. Under temperature-zero greedy decoding, verification preserves the exact token choices that the target would make alone. The drafter does not get permission to substitute a merely plausible alternative.

At nonzero temperatures, the goal changes from reproducing one sequence to preserving the target distribution. Correct speculative sampling can do that under matched settings, according to the foundational sampling research. Speed still depends on how often draft and target distributions align.

Liquid AI reports that increasing temperature reduced acceptance in its experiments. More probability spreads toward lower-ranked tokens, creating more chances for the drafter and target to disagree. Consequently, creative sampling can deliver smaller gains than deterministic generation.

This matters for application design. Document extraction, visual grounding, chart reading, and constrained responses often use low temperatures. Open-ended image conversations may use more sampling, making the published greedy results less representative.

Liquid AI speculative decoding therefore fits predictable output particularly well. Captioning standardized product images, reading receipts, describing interface screenshots, and extracting structured facts offer plausible workloads. Their actual gains still require local measurement.

Faster Decoding Does Not Remove the Vision-Language Bottleneck

The main uncertainty is how much of a real request remains outside the accelerated decoding phase.

A vision-language request contains more work than text generation. The system must encode the image, transform it into visual representations, process those visual tokens with the prompt, and then decode the answer.

DSpark accelerates only the final stage. It does not make image encoding faster. It also leaves prefill unchanged, meaning the target still processes the prompt and visual-token context before producing its first answer token.

That boundary explains the gap between decoding and end-to-end results. Liquid AI reported up to 3.13x faster decoding on the M5 Max, but its highest total improvement was 2.62x. Other tasks showed wider differences.

On TextVQA, the company measured a 2.69x decoding improvement and a 1.56x end-to-end gain on the M5 Max. The result implies that image processing and prefill consumed a substantial portion of the original request time.

The limitation follows Amdahl’s law, which caps overall acceleration when part of a workload remains unchanged. If decoding represents half the baseline latency, even an infinitely fast decoder cannot improve the entire request beyond 2x.

Edge hardware makes this constraint especially relevant. Consumer devices provide less compute than datacenter GPUs for vision encoding and long-context prefill. A large image or multi-image prompt can delay the first token before speculative decoding begins helping.

The independent MMSpec benchmark reinforces the need for cautious interpretation. Its authors evaluated 600 samples across six task categories and ten speculative-decoding methods. They found that throughput speedup alone did not reliably represent latency performance.

MMSpec also found that techniques designed for text-only language models can degrade in multimodal settings. Cross-modal dependencies change which proposals are likely to survive. Vision awareness becomes more important as batch sizes increase.

Liquid AI followed MMSpec’s task categories, which improves the breadth of its internal evaluation. However, Liquid AI conducted and published the DSpark performance tests itself. Independent replications across common devices and production prompts remain limited.

The comparison baseline deserves equal attention. The published multipliers compare the same LFM2.5-VL-3B target with and without its drafter under specified frameworks. They do not establish that the combined system outperforms every competing VLM.

They also do not compare the system against alternative latency strategies. Developers can quantize the target, reduce image resolution, cache visual embeddings, shorten prompts, batch requests, or select a smaller model. Those changes affect different parts of the latency budget.

Memory is another operational consideration. The drafter is small beside the 3-billion-parameter target, but it is not free. Its weights, cache, runtime state, and integration consume capacity that matters on constrained devices.

Compatibility creates further friction. SGLang, MLX-VLM, and llama.cpp now provide published paths, but teams using other serving systems cannot assume immediate support. Production adoption requires stable loading, observability, batching behavior, and failure handling.

MLX-VLM’s current greedy-only DSpark path narrows its immediate use cases. Applications that depend on sampled generation need another supported stack or must wait for broader sampling support. Even then, higher temperatures can lower draft acceptance.

The release should therefore be judged as a credible systems optimization with clearly stated boundaries. It is not evidence that visual inference has become 3.13x faster in every meaningful sense.

The Real Contest Is Better Prediction Versus Less Model Work

DSpark strengthens a route where developers keep the target intact and optimize how often it must run.

The standard edge-AI response to latency has been to reduce the target model’s workload. Teams use fewer parameters, lower precision, shorter prompts, smaller images, or task-specific distillation. Each technique can improve responsiveness, but each can introduce quality or flexibility tradeoffs.

Liquid AI LFM2.5-VL-3B-DSpark proposes another path. Keep the existing target and predict several future steps with a specialized companion. Let the target verify those guesses without surrendering control over the final output.

This route becomes attractive when a team already accepts the target model’s capabilities. Replacing it would require new evaluations, prompt changes, safety checks, and product tuning. Attaching a drafter can preserve more of that investment.

The approach also suits local deployment, where memory bandwidth frequently constrains token generation. Verifying a block can use hardware more efficiently than repeatedly loading model state for one token. Liquid AI’s Apple results make that possibility concrete.

Yet smaller models retain advantages. They simplify packaging, reduce total memory use, and accelerate stages that a decoder-only optimization cannot touch. A compact vision encoder can improve time to first token, while DSpark cannot.

Quantization can also combine with speculative decoding rather than compete exclusively against it. Liquid AI supplies a GGUF drafter matched to its GGUF target. A local stack can therefore reduce precision and add drafting, provided the runtime supports the pair correctly.

That combination shifts the engineering question from choosing one technique to assigning each technique to the right bottleneck. Quantization reduces weight size and arithmetic cost. Image preprocessing changes vision cost. Speculation targets autoregressive generation.

This is why phase-level profiling becomes essential. A document assistant that analyzes high-resolution pages may spend most of its time before decoding. A visual chat tool producing detailed descriptions may spend far more time generating output.

The same distinction applies to user experience. Time to first token shapes whether an application feels responsive at the start. Tokens per second shape whether a long answer feels fluid after generation begins.

DSpark directly improves the second measure. It improves total latency when decoding occupies enough of the request. It does not guarantee a proportionate improvement in the first.

Developers should also separate single-user speed from fleet throughput. Liquid AI observed an advantage across measured H100 concurrency levels, but the difference narrowed at higher load. Production economics depend on request mixtures, batching, and service-level targets.

The most consequential aspect of the release is therefore architectural. Liquid AI has packaged speculative decoding as part of a deployable vision-model family instead of presenting it only as research.

If that pattern spreads, model releases may increasingly include a target, multiple quantizations, and hardware-specific drafters. Inference optimization would become part of the model artifact rather than an after-market serving decision.

That direction puts pressure on other open-weight model developers. Publishing only a checkpoint leaves downstream teams responsible for acceleration. Shipping a matched drafter offers a more complete latency story, even when the measured gains remain workload-dependent.

Three Signals Will Show Whether the Speedup Matters in Practice

Independent testing, broader sampling support, and real application profiles will determine whether DSpark becomes a repeatable deployment pattern.

The first signal is independent replication across accessible hardware. Developers should watch for tests on M-series Macs, consumer GPUs, and edge systems using identical prompts with and without the drafter.

Useful reports must disclose image dimensions, prompt length, output length, temperature, quantization, and runtime version. A single tokens-per-second number cannot explain whether the application’s full response became meaningfully faster.

Replication close to Liquid AI’s ranges would strengthen the company’s case. Smaller or inconsistent gains would suggest that the published workloads favor the drafter more than everyday applications do.

The second signal is broader runtime and sampling support. MLX-VLM currently limits the published DSpark path to greedy generation, while SGLang targets Nvidia deployment and llama.cpp serves GGUF use cases.

Support in additional inference engines would reduce integration cost. Stable nonzero-temperature sampling would also make the method more relevant to visual chat and creative description tools.

Developers should examine acceptance rates as sampling changes. Liquid AI says higher temperatures reduced acceptance and throughput in its experiments. Production tests should reveal whether those declines remain acceptable for conversational products.

The third signal is whether teams report phase-level gains from real applications. The decisive metrics are time to first token, decoding rate, end-to-end latency, peak memory, and throughput under expected concurrency.

A long-form visual assistant may benefit substantially because generation dominates its session. An OCR workflow returning a few fields may gain far less because vision encoding and prefill occupy most of the request.

Teams evaluating Liquid AI LFM2.5-VL-3B-DSpark should begin with traces, not headline multipliers. Measure the baseline share consumed by image encoding, prefill, and decoding. Then attach the drafter and repeat the same workload.

Check output equivalence under greedy decoding and distributional behavior under supported sampling. Measure warm and cold starts separately. Include memory pressure and sustained thermals when testing laptops or mobile-class systems.

The release provides enough implementation detail to make those evaluations possible. It also gives developers a useful reminder: model capability and inference behavior are separate engineering problems.

Liquid AI’s 3.13x figure is best understood as evidence that a matched drafter can materially accelerate one phase of local multimodal inference. The end-to-end numbers show both the value and the limit of that claim.

Will matched draft models become standard companions for open-weight VLMs, or remain specialized optimizations for long-output workloads? The answer will come from reproducible application traces, not a single peak benchmark.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page