New Drop from DeepSeek: Fully Open Stack for Accelerating LLM Generation
New Drop from DeepSeek: Fully Open Stack for Accelerating LLM Generation
DeepSeek has released an open stack for LLM generation, featuring the DSpark algorithm that significantly boosts generation speed.
Inside, there are ready algorithms, training, evaluation, and even a data pipeline. Take it and use it, super practical. github.com/deepseek-ai/DeepSpec
The essence lies in the DSpark algorithm. DeepSeek is already using it for DeepSeek-V4 Flash and Pro in production, and according to their data, the generation speed for users has increased by approximately 60–85% compared to the old baseline.
How the algorithm works:
Fundamentally, it is a small model that drafts outlines for the main LLM. This is called a draft model.
This approach is currently in vogue, but DeepSeek is taking it to a new level. Their draft model operates unusually, in two stages. First, a block of tokens is gathered in parallel, and then a lightweight Markov module refines the dependencies between neighboring tokens. Thanks to this approach, the drafter works quickly and does not drop significantly in the tails.
After the drafter has sketched an outline, the main LLM checks it and accepts only the correct prefix, adjusting the rest. Meanwhile, DSpark itself decides how many tokens to send for verification based on confidence estimates for the tokens and the current load on the hardware.
As a result, we achieve acceleration of at least 1.5 times without any loss of quality. Hats off to DeepSeek for such open-source work.
Why it matters
AnalysisThis new stack from DeepSeek could significantly change approaches to LLM generation, reducing processing time without sacrificing quality, which is critical for developers and companies working in this field.
Discuss in community
Share your questions and insights with developers