Open Compact Language Model with 50,000 Stars on GitHub
A new open compact language model has gained popularity on GitHub, allowing for training with minimal costs and time.
Open Compact Language Model with 50,000 Stars on GitHub
An open small model has been found that has already gained 50,000 stars on GitHub and made it to the Trending section.
With its help, one can train a super-compact language model with only 25.8 million parameters — approximately for $3 and two hours.
The project bets on maximum compactness: the smallest version is only 1/2700 the size of GPT-3. Additionally, the training of the tokenizer has been made available, as well as the full code for all training stages — Pretrain, SFT, and more. The project also includes the multimodal visual model MiniMind-V.
All training scripts are written in pure PyTorch, compatible with MoE architectures, and essentially represent an open implementation of all stages of training large language models.
Why it matters
AnalysisThis model represents an important step in the development of compact language models, allowing for reduced training costs and resource usage. The open access to code and training materials fosters community development.
Discuss in community
Share your questions and insights with developers