[MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model!
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey guys! Supra2-Medium is finally out! It's a 25M parameters qwen3 architecture model trained entirely from scratch (on our new rig: RTX 5060 Ti 16GB + the new RTX 5060 8GB!). Here's how it competes in benchmarks with Supra-50M-Base (which is double as large!!): Note: This is a BASE model only; instruction tuned version maybe to come in the next time. Link to our HF org: https://huggingface.co/SupraLabs --> Link to the model: https://huggingface.co/SupraLabs/Supra2-Medium-Base <-- Here's a sample from the model:
Give us a like and a follow and feel free to provide us with feedback! 🤗🔥 ...and...stay tuned: Supra3 coming soon with four models: Flash-Lite 25M, Flash 50M, Pro 75M and Ultra 100M. 👀 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.