Up to 3.2x Faster Inference with LFM2.5-DSpark
LiquidAI releases DSpark draft model checkpoints for three models from their LFM2.5 family, achieving up to 3.2x faster inference on a GPU and up to 2.87x on-device.
Intelligence analysis by Llama

LiquidAI's DSpark draft models for LFM2.5 offer significant speed improvements in inference, with up to 3.2x faster throughput on a GPU and up to 2.87x on-device, while maintaining quality parity with baseline greedy decoding.
Imagine you're trying to solve a puzzle, but you're not sure what the answer is. A normal computer would try one piece at a time, but a computer with DSpark would try many pieces at once and then check if they fit. This makes the computer much faster and more efficient.
Analysis
DSpark: A New Approach to Speculative Decoding
LiquidAI's DSpark is a new approach to speculative decoding that addresses the memory-bound nature of the decode phase in LLM inference. By using a lightweight draft model to produce candidate tokens and then having the target model verify them in a single forward pass, DSpark reduces the latency of the decode phase. This approach is particularly effective on edge devices, where memory is limited and computation is slower.
Training and Architecture
The DSpark draft models are trained on a larger and more diverse data mix, covering SFT, chat, code, and function-calling data. The models are relatively small, with each around ~300M parameters, and are designed to be efficient and fast. The training process involves running 15 epochs on the entire dataset and selecting the epoch with the highest acceptance rate rather than the lowest loss.
Quality Parity
The DSpark draft models maintain quality parity with baseline greedy decoding, with the emitted sequence being identical to baseline greedy by construction. This means that the benchmark accuracy (pass@1 or exact match) is unchanged, and the models can be used as drop-in replacements for existing models.
Inference Speed Up
The DSpark draft models deliver noticeable throughput improvements on both the large-scale accelerator (H100) and the edge deployment (M4 Max MacBook). For LFM2.5-2.6B, the speedup on the MacBook is especially noticeable, pushing the interactivity level a user can enjoy far beyond the throughput offered by most proprietary cloud models.
Key points
- DSpark is a new approach to speculative decoding that reduces latency in LLM inference.
- The DSpark draft models are trained on a larger and more diverse data mix, covering SFT, chat, code, and function-calling data.
- The models are relatively small, with each around ~300M parameters, and are designed to be efficient and fast.
- The DSpark draft models deliver noticeable throughput improvements on both the large-scale accelerator (H100) and the edge deployment (M4 Max MacBook).
If this development continues to play out positively, we can expect to see even faster and more efficient AI models in the future, leading to breakthroughs in areas such as natural language processing and computer vision.
However, there are also potential risks associated with this development, such as the possibility of AI models becoming too complex and difficult to understand, or the potential for bias and unfairness in the models.
Market signals
- Gold Escalation drives safe-haven demand for gold, per the article's framing of investor reaction.
AI-generated analysis of potential market relevance. Not financial advice.



