AntLing’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key highlights: #1- Ling-3.0-tiny-base: 7.9B total | 1.3B active. Despite having only half as many total parameters as Ling-2.5-mini-base, Ling-3.0-tiny-base delivers comparable or superior performance on most benchmarks, with particularly strong results in coding. #2- Ling-3.0-flash-base: 124B total | 5.1B active. In evaluations, Ling-3.0-flash-base achieves strong performance across coding, reasoning, and long context tasks, even when compared with models 2 to 3 times larger. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.