r/MachineLearning · June 21, 2026 · 1 min read

I released a softmax-free attention model at GPT-2 Medium scale (~354M params, 11.5B tokens): structural sparsity + tile-skipping kernels for long-context VRAM savings. Open weights + custom Triton kernels [R]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

Discussion (0)

No comments yet. Sign in and be the first to say something.