r/LocalLLaMA · · 1 min read

ling 3.0 flash/tiny base models

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

ling 3.0 flash/tiny base models

https://huggingface.co/inclusionAI/Ling-3.0-tiny-base

https://huggingface.co/inclusionAI/Ling-3.0-flash-base-midtrain

https://huggingface.co/inclusionAI/Ling-3.0-flash-base-30T

https://huggingface.co/inclusionAI/Ling-3.0-tiny-base-midtrain

https://huggingface.co/inclusionAI/Ling-3.0-tiny-base-30T

These checkpoints correspond to different stages of the training process:

  • Pretrained checkpoint have completed large-scale pretraining but have not undergone mid-training, WSM merging (or learning-rate decay), or post-training.
  • Mid-trained checkpoint have completed mid-training but have not undergone WSM merging (or learning-rate decay) or post-training.
  • Merged checkpoints have undergone WSM merging (or learning-rate decay) based on the mid-training checkpoints but have not undergone post-training.

These checkpoints are released to support continued pretraining, fine-tuning, and further research. For the post-trained model, please see and see Ling-3.0-tiny and Ling-3.0-flash.

submitted by /u/jacek2023
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA