DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Their summary:
Their technical blog: https://pub.sakana.ai/diffusionblocks/ From what I understand, if this works at a scale, it will allow training with 3-4x memory savings all around (weights, gradients, optimizer, activations). Also makes training more parallelizable with less communication, can improve Looped Transformer training, and can improve diffusion model inference efficiency. Really hope it scales well. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.