Chinese AI Lab DeepSeek Releases Powerful Open Model to Challenge Industry Giants

DeepSeek unveils V3, a massive open-weights AI model that outperforms leading competitors on coding and text benchmarks while costing only six million dollars to train.

A Chinese AI lab named DeepSeek releases a highly powerful open-weights model called DeepSeek V3. The company makes this model available under a permissive license, which allows developers to download, modify, and use it for commercial applications. DeepSeek V3 handles various text-based tasks such as coding, translating, and essay writing with remarkable proficiency.

According to internal benchmarks, DeepSeek V3 outperforms both open and closed AI models, including Meta's Llama 3.1 405B and OpenAI's GPT-4o. The model particularly excels in coding competitions on platforms like Codeforces and the Aider Polyglot test. It achieves these results with a massive architecture containing 671 billion parameters and training on a dataset of 14.8 trillion tokens.

Despite its enormous size and capability, DeepSeek V3 requires surprisingly little resources to train compared to industry norms. Former OpenAI co-founder Andrej Karpathy notes that the model trains on just 2,048 GPUs over two months for a total cost of around six million dollars. This exceptionally low budget drastically undercuts the tens of thousands of GPUs typically required to train frontier-level AI systems.

Read More at the original source →