DeepSeek Launches V3.2-Exp Model With Sparse Attention, Slashes API Prices

DeepSeek releases its new V3.2-Exp experimental model featuring a novel sparse attention mechanism that cuts compute costs without sacrificing quality. The update also brings a massive 50 percent reduction in API pricing.

DeepSeek unveils DeepSeek-V3.2-Exp, a new experimental artificial intelligence model built directly on the foundation of V3.1-Terminus. The standout feature of this release is DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism designed to accelerate training and inference for long-context tasks. This technical innovation allows the model to maintain high output quality while significantly reducing the computational power required.

Benchmark tests reveal that V3.2-Exp performs on par with its predecessor, ensuring that users experience no drop in capability despite the internal efficiency upgrades. Alongside the model launch, DeepSeek cuts API prices by more than 50 percent, making this powerful technology much more accessible to developers. For users who want to compare the two versions, the company keeps the V3.1-Terminus model available through a temporary API endpoint until mid-October.

The new model is currently live across the DeepSeek App, Web interface, and API. In addition to the commercial rollout, DeepSeek embraces the open-source community by releasing the model weights on Hugging Face and publishing a detailed technical report on GitHub. Developers also gain access to key GPU kernels written in TileLang and CUDA, which enable rapid research prototyping for those looking to build upon the new DSA architecture.

Read More at the original source →