llama.cpp releases · · 1 min read

b10632

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

ggml-metal: add chunked SSD MMA for Mamba-2 prefill optimization (#26647)

  • metal: WIP chunked SSD SSM_SCAN kernels for multi-token prefill

  • metal: drop scalar SSD path; MMA + sequential tail

  • drop WIP ssm scan test noise

  • remove state_from_dst and rename CS and NSG constants

  • remove unrelated added whitespace padding

  • added clarity to mma_tokens calculation

  • added clarity to use_mma bool checks

  • added comments to metal ssd op constants for clarity

  • reserve K tokens for sequential kernel rollback snapshots

  • reset concurrency between mma and seq tail

  • remove print args no longer used

  • fixed comment to no longer point to specific line

  • add FC_SSM_SCAN so seq path skips token offlset unless it's mma tail

  • added changes to new ssm.metal for rebase after ggml-metal.metal refactor

  • specialize ssm_scan tail with a template instead of a function constant


Co-authored-by: dpantaleoni dominikpantaleoni@gmail.com
Co-authored-by: forforever73 690105611@qq.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases