Peking University and DeepSeek release DSpark for faster LLM inference
Peking University and DeepSeek have jointly open-sourced DSpark, a speculative decoding framework designed to improve the efficiency of large language model inference. The release focuses on accelerating model responses while maintaining performance under strict latency requirements.
DSpark boosts LLM inference speed by 60-85% and can deliver up to 661% throughput gain under strict latency constraints. The framework positions speculative decoding as a practical route to faster deployment of language models where response time and serving capacity are critical.
Originally reported by pandaily.comRead the source →
Related coverage
DeepSeek releases V4 Flash Vision Exp
2 days ago
China keeps open-weight models in play as security questions grow
6 days ago
DeepSeek raises V4 API prices as demand pressures capacity
1 week ago
AI inference token prices hit annual low
2 weeks ago