vLLM v0.28.0: Sparse Attention & Kimi K3 Optimizations

Key Takeaways
- vLLM v0.28.0 introduces sparse attention and significant performance gains for Kimi K3 via an adaptive spec budget.
- The release expands ROCm coverage, enabling DeepSeek V4 on gfx11 and gfx950 and Kimi K3 via the V2 model runner.
- Development momentum remains high, with v0.28.0 featuring 584 commits from 270 contributors.
vLLM, the industry-leading high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs), has reached a pivotal milestone with the release of version 0.28.0. This update represents a strategic shift toward deeper architectural optimizations, focusing on reducing latency and expanding the hardware ecosystem to allow developers to deploy state-of-the-art AI models with unprecedented speed and efficiency.
Architectural Breakthroughs in Inference Speed
The core of the v0.28.0 release is a series of groundbreaking innovations designed to tackle the primary bottlenecks of LLM serving: memory bandwidth and token generation latency. AlphaSignal reports that the update ships with the implementation of sparse attention, a technique that reduces the computational complexity of the attention mechanism by focusing only on the most relevant tokens, thereby significantly lowering the memory footprint during long-context processing.
One of the most impactful additions is the significant boost to speculative decoding. Speculative decoding allows a smaller, faster "draft" model to predict multiple tokens, which are then verified in parallel by the larger target model. vLLM v0.28.0 introduces an adaptive spec budget, which dynamically adjusts the number of tokens speculated based on the model's confidence and current system load. AlphaSignal notes that this specific implementation has delivered approximately 60% better DSpark TTFT (Time to First Token) for Kimi K3, drastically reducing the perceived lag for end-users in real-time applications.
This push for efficiency is not limited to a single feature but is part of a broader stack-wide optimization effort. The vLLM project's X account highlights that the release includes a comprehensive optimization push for both Kimi-K3 and DeepSeek-V4. These optimizations ensure that these specific model architectures can leverage the engine's memory management capabilities—such as PagedAttention—more effectively, ensuring higher throughput and lower per-token latency.
Expanding the Hardware Frontier: ROCm and AMD Support
To democratize high-performance inference, vLLM v0.28.0 significantly extends its compatibility with non-NVIDIA hardware, specifically targeting the AMD ROCm ecosystem. AI/TLDR notes that Kimi K3 can now run on ROCm through the V2 model runner, providing a viable alternative for enterprises utilizing AMD Instinct accelerators.
Further expanding this hardware reach, DeepSeek V4 is now enabled on gfx11 and gfx950 architectures. This is a critical update for developers utilizing the latest generation of AMD GPUs, as it allows them to run one of the most capable open-weights models with native hardware acceleration. Additionally, the release includes official support for AMD Quark, further diversifying the deployment options for LLM practitioners.
The engine's flexibility is fundamentally rooted in its modular architecture. By allowing for third-party hardware plugins, vLLM avoids the "monolithic repository" trap. As detailed in the vLLM installation documentation, these plugins reside outside the main repository and follow the Hardware-Pluggable RFC. This design ensures that as new AI accelerators emerge from various vendors, they can be integrated into the vLLM ecosystem without compromising the stability of the core engine.
Deep Dive: The Kimi K3 and DeepSeek V4 Synergy
The focus on Kimi K3 and DeepSeek V4 in this release is not coincidental. These models represent the cutting edge of architectural innovation. Finance Biggo notes that Kimi K3 employs a unique Multimodal PP (Pipeline Parallelism) optimization, placing the Vision Encoder in the middle or tail of the PP pipeline rather than at the beginning. This allows the system to utilize idle resources more effectively, a characteristic that vLLM v0.28.0 is now better equipped to handle through its optimized serving pipeline.
Similarly, the integration of DeepSeek V4 benefits from vLLM's improved handling of sparse architectures. As the industry moves toward Mixture-of-Experts (MoE) and other sparse attention mechanisms—similar to the MSA Sparse Attention Architecture analyzed by vLLM in relation to MiniMax M3—the ability of the serving engine to efficiently route tokens and manage KV caches becomes the deciding factor in deployment costs and user experience.
A Community-Powered Engine
The rapid evolution of vLLM is a testament to the power of open-source collaboration. The scale of the project has seen an aggressive upward trajectory. While the previous v0.27.0 release featured 561 commits from 242 contributors, the v0.28.0 release has increased this momentum significantly. The vLLM project's X account and AlphaSignal both confirm that v0.28.0 consists of 584 commits from 270 contributors, including 76 new contributors who joined the project during this cycle.
This influx of talent has allowed vLLM to expand beyond the main inference engine. The ecosystem is evolving with specialized projects like vLLM-Omni. The vLLM-Omni v0.26.0rc1 release recently aligned with the vLLM 0.26 release line, focusing on delivering major improvements for real-time multimodal serving and streaming capabilities. This suggests a future where vLLM is not just a text-inference engine, but a comprehensive multimodal serving layer capable of handling audio, vision, and text simultaneously.
The Future of High-Throughput Serving
With the release of v0.28.0, vLLM is positioning itself as the definitive bridge between research-grade model architectures and production-grade deployment. By solving the "Time to First Token" problem through adaptive speculative decoding and expanding the hardware reach to include AMD's latest silicon, vLLM is lowering the barrier to entry for deploying massive models like DeepSeek V4 and Kimi K3.
As LLMs continue to grow in parameter count and context window size, the importance of sparse attention and memory-efficient serving will only increase. The architectural decisions made in v0.28.0—specifically the commitment to the Hardware-Pluggable RFC and the focus on model-specific performance pushes—ensure that vLLM will remain agile enough to adapt to the next wave of AI breakthroughs.
Relevant solution
Website Development
Custom website development — fast, modern, ready to sell.
Related Articles

GTA VI Release Date, Gameplay Details & PS5 Availability
Rockstar Games has officially set the release date for Grand Theft Auto VI for November 19, 2026. Discover the latest on the dual protagonists Jason and Lucia, PS5 exclusivity, and the reactive world of Vice City.

Iceland EU Referendum 2026: Will Iceland Rejoin the EU?
Iceland holds a pivotal national referendum on August 29, 2026, to decide whether to restart EU membership talks after a 13-year hiatus. Discover the economic and security drivers behind this knife-edge vote.

Cara Digital Detox Singkat: Atasi Doomscrolling & Cemas
Sering merasa cemas tanpa ponsel atau terjebak doomscrolling? Pelajari panduan lengkap digital detox singkat untuk memulihkan kesehatan mental, meningkatkan produktivitas, dan mencapai work-life balance yang ideal.

Jaguares de Córdoba vs América de Cali: Liga BetPlay Analysis
América de Cali's unbeaten streak faced a stern test against Jaguares de Córdoba in a high-stakes Liga BetPlay clash. Discover how the league leaders fared in Montería.
Dapatkan Artikel Terbaru!
Berlangganan newsletter kami untuk mendapatkan tips dan insight menarik langsung ke inbox Anda.
Kami tidak akan pernah membagikan email Anda (No Spam).