Initial ET backend by marty1885 · Pull Request #24179 · ggml-org/llama.cpp
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| This PR is developed by AINekko and by members of AIFoundry (AINekko's OSS community) and adds the ET backend that supports the ET-SOC-1 processor. ET-SOC-1 was originally created by Esperanto Technologies which AINekko later open sourced under Apache 2.0 and would like to upstream the llama.cpp backend we developed for it as a way to integrate open source hardware into the open source inference ecosystem. The the ET processor core documentation and RTL can be found at the following links
As ET-SOC-1 is an older low power processor, the absolute performance is not impressive compared to even CPUs. But it still provides better performance per watt then my ARM R7 7700 development machine can do. Please refer to the following table for concrete performance number [link] [comments] |
More from r/LocalLLaMA
-
Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and…
Aug 30
-
Got MiniMax H3 video generation running in TensorSharp
Aug 30
-
Qwen3.8-Flash-Next NVFP4 2xDGX Spark config: 50t/s decode, 2,900t/s prefill
Aug 30
-
Don't Sleep on EXL3 Quants
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.