News / #edge Tag Edge 407 articles archived under #edge · RSS Sign in to follow r/LocalLLaMA community 23d ago what will be the future of LocalLLaMA? For a long time now, the most popular posts on LocalLLaMA have been either about using LLM in the cloud or about politics. I suspect that people using local models are about 10% now. You can say that this is very good, because now it is an inclusive sub, without gatekeeping. But… 28 arXiv — NLP / Computation & Language research 23d ago EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding arXiv:2608.05303v1 Announce Type: cross Abstract: On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers. Speculative decoding and… 21 r/LocalLLaMA community 23d ago 🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp 🐦⬛ Magpie-TTS Multilingual 🦜 Nemotron Speech Streaming EN 0.6B 🦜 Nemotron-3.5 ASR Streaming 🦜 Parakeet CTC 1.1B 🦜 Parakeet TDT 0.6B v3 🥦 NanoCodec Merged PR https://huggingface.co/nvidia/magpie_tts_multilingual_357m#run-magpietts-locally-with-nemo-speechcpp I run open… 9 r/LocalLLaMA community 23d ago I thought Deepseek was the answer since I cannot afford GPU for local LLM   submitted by   /u/HsSekhon [link]   [comments] 18 r/LocalLLaMA community 23d ago Best open-source harnesses for combining cloud and local AI model orchestration? Looking for best current solutions for combining cloud models and local models seamlessly inside a harness' orchestration Edit: Right now, we don't have harnesses (that I'm aware of) that are blending local and cloud models to work together simultaneously to accomplish tasks set… 32 r/LocalLLaMA community 23d ago 32 total local models tested head to head I ran 32 local models head to head on one fact-extraction corpus, 1,001 notes, paired bootstrap on every adjacent pair. Several weeks of compute time, all on consumer grade cards. Most of the field does not separate. Six consecutive steps from 2B to 31B, and the bootstrap cannot… 37 r/LocalLLaMA community 24d ago i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local… [Fully open source under GPL3, made from the ground up for use with local models, no subscriptions, no corporate backing] When i first started this, it was meant to be a fully lightweight, extremely modular alternative to openclaw, hermes and the like , and it still is! But i… 19 arXiv — NLP / Computation & Language research 24d ago Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high VRAM, time, and energy costs. Parameter-Efficient Fine-Tuning (PEFT) of Small… 8 r/MachineLearning community 24d ago Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P] Over the past month, I've been building LiveTranscriber, an open-source iOS app for running modern speech and language models entirely on-device. The goal was to see whether recent open-source models could be turned into a practical mobile product—not just technical demos.… 8 r/LocalLLaMA community 24d ago Ling-3.0-flash MXFP4 released and running locally on one DGX Spark. In tests: ~80 tok/s decoding 2,500–3,500 tok/s long-input prefilling Smooth use by 3–4 concurrent users Private, on-device inference for coding, agents, and offline batch jobs   submitted by   /u/niacolhealth [link]   [comments] 30 r/LocalLLaMA community 24d ago Given the MiniMax H3 LoRAs Debacle - Some Important Context for Censorship enforcement and laws in China *I felt the need to write this post because it seems like very few people on this sub are aware of Chinese laws and how they're enforced, so here's an explainer coming from a Chinese person (myself). I know that this post isn't directly about local models per se, but I'm seeing… 12 TechCrunch — AI news-outlet 24d ago MacPaw taps Liquid AI to offer on-device inference to devs building for its app store MacPaw is building a local version of its AI assistant Eney using Liquid AI's models. 33 r/LocalLLaMA community 25d ago Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or export selected pages to an MP3. Everything runs locally: Kokoro for speech, and an on-device… 5 arXiv — Machine Learning research 25d ago GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection arXiv:2608.02690v1 Announce Type: new Abstract: On-device training of deep neural networks is fundamentally constrained by the computational and memory costs of large-scale datasets. Coreset selection offers a practical solution by retaining only a compact subset of real… 22 arXiv — Machine Learning research 25d ago Design-Time Optimization of Deep Neural Networks for Intermittent Learning on Microcontrollers arXiv:2608.03589v1 Announce Type: new Abstract: We present a method for designing deep neural networks (DNNs) for intermittent, energy-autonomous, on-device learning on microcontroller units (MCUs). In mobile applications where the energy can run out, e.g., when solar-powered,… 10 arXiv — NLP / Computation & Language research 25d ago MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale arXiv:2608.02613v1 Announce Type: new Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric… 15 r/LocalLLaMA community 25d ago GPT-OSS has turned one year old today! It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that model is much slower (A10B) and has not been released in a local-friendly QAT… 9 r/LocalLLaMA community 25d ago Local LLM 35B MoE — Real-world coding benchmarks (Qwen vs Ornith vs KAT) I’ve been running a fairly opinionated evaluation loop on ~35B A3B/MoE-class models for coding over the past few months. Not synthetic benchmarks: actual dev workflows, iterative debugging, refactoring passes, and failure recovery. Here’s where things stand for me: Qwen 3.6 (35B… 9 r/LocalLLaMA community 25d ago SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth Hopefully this would let us have faster local models....but it will probably be out of our price range.   submitted by   /u/giveen [link]   [comments] 12 r/LocalLLaMA community 26d ago Is LM Studio abandoning their core product? Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that LM Studio replaced almost every link to the original app that built their brand… 7 llama.cpp releases dev-tools 26d ago b10255 Extended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0–Q8_0 and FP32) ( #25874 ) sycl: extend oneDNN SDPA to Q4_0-Q8_0 and F32 KV caches Extends the oneDNN SDPA path (PR #25222 ) to handle non-F16 KV caches by dequantizing or converting K/V to dense FP16 on-device before feeding… 4 arXiv — Machine Learning research 26d ago Kilobyte Models: Neural Networks as a Seed and a Quantized Latent arXiv:2608.00860v1 Announce Type: new Abstract: The cost of storing and transmitting a trained neural network scales with its parameter count, a bottleneck for over-the-air updates, on-device libraries, and other bandwidth-bound deployments. We study an extreme form of model… 17 arXiv — NLP / Computation & Language research 26d ago Opt.Gear Technical Report arXiv:2608.01034v1 Announce Type: new Abstract: We introduce Opt.Gear, a foundation model designed for efficient on-device deployment, real-tim inference, and strong task capability. It includes a dense model (1M, 270M, and 1B) with a context length of 64K. We designed a new… 32 arXiv — Machine Learning research 27d ago Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning arXiv:2607.29353v1 Announce Type: new Abstract: With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., custom keyword spotting) or patients (e.g., adaptive health monitoring). Yet, most edge… 38 arXiv — Machine Learning research 27d ago GQ-FSL: Green Quantized Federated Split Learning arXiv:2607.29659v1 Announce Type: new Abstract: Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. While federated split learning (FSL) mitigates on-device… 28 arXiv — NLP / Computation & Language research 27d ago Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation arXiv:2607.29250v1 Announce Type: new Abstract: Small language models (SLMs) are attractive for agentic deployment due to low latency, reduced cost, and on-device privacy, yet they struggle with tool-use tasks where training data is scarce and noisy. Unlike larger models, SLMs… 13 r/LocalLLaMA community 28d ago A collection of small domain-specific benchmarks for local models (30+ and growing) Hello fellow local AI people! I took "you must create your own benchmarks" literally, and built a website for this. How does the end result look like Let's say I want to know which model has most common sense in its responses, I did everything including evaluating responses (see… 34 r/LocalLLaMA community 28d ago Tomte - super fast harness for Gemma 4 I think people are sleeping on Gemma and local models so I built a free, very fast harness for Gemma 4 that I call Tomte. https://tomteapp.com Works on Macs with M processors, will have a companion app you can connect to anywhere. So far does everything I ever needed chatGPT… 5 r/LocalLLaMA community 28d ago I'm kinda tired of obsession for one-shot tests in coding, there are good tests for multi-step debugging with analyzing output/images/videos? Personally, i think good coding model shouldn't be focused on one-shot "everything in one html-file" tests, but should be really good on debugging, fixing and modifying its own output. Anyone know such simple tests that i would able to run with local models? May be some kind of… 15 r/LocalLLaMA community 29d ago Rule Suggestion: "Open" models without weight releases should be tagged [no weights] A lot of recent models are being announced with promised open weights, but the weights are either weeks away, or in some cases (looking at you Meta) not being released at all. This sub is about local LLMs - not "maybe local in the future" llms. These models are still useful to… 26 Hugging Face Daily Papers research 1mo ago AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition Abstract On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges… 17 arXiv — Machine Learning research 1mo ago DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement arXiv:2607.27761v1 Announce Type: new Abstract: In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across different views often suffer from misalignment, leading to the partial view… 14 arXiv — NLP / Computation & Language research 1mo ago Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning arXiv:2607.27766v1 Announce Type: new Abstract: On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local… 21 r/LocalLLaMA community 1mo ago Local LLMs for non-coding What are your top 3 uses cases? Seems that outside coding the application of local is limited?   submitted by   /u/Salt_Armadillo8884 [link]   [comments] 18 r/LocalLLaMA community 1mo ago Smallest model (& tips) for intelligent computer use via Hermes? Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive an actual machine via hermes' computer_use tool and cua_driver to click through… 7 r/LocalLLaMA community 1mo ago Mechanistic interpretability streamlined for everyday users like us😎 🧠 Context: I want to give the community an Open Research (well open under Apache 2.0 clause) - tool that allows everyday users like us to look deeper into the local models we use consistently. Mechanistic interpretability streamlined into a more visible work-flow. Easy to read for… 26 arXiv — NLP / Computation & Language research 1mo ago Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation arXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt shape: translate one source sentence into one target language. In practice,… 29 r/LocalLLaMA community 1mo ago I tested proven orchestration techniques on small local models. 90% failed. The 10% that survived roughly doubled task completion. Hey localllama brochacos, what's up? I'm u/raydestar , long time local llm fan. SWE with about 10 years xp, and I have been cranking hard trying to skill up with agentic AI recently. Since open weights got good, it's just blown my mind. What's been hard to understand is "Why… 26 llama.cpp releases dev-tools 1mo ago b10181 ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory ( #26141 ) ggml_cuda_should_use_mmq() selects MMQ purely from the quantization type. The current MMQ configurations are designed and maintained against a minimum of 48 KiB per-block shared memory, the limit… 36 Hacker News — AI on Front Page community 1mo ago Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on… 29 r/LocalLLaMA community 1mo ago A slide deck you can edit with a local model or in Chrome — the whole deck is a JSON block in one HTML file (~640KB with editor and viewer included) Over the past few months, our team has been building more and more slidedecks using web frontend technologies with coding harnesses, but a common complaint is to make even small edits we need to edit the code either manually or via the harness. To avoid this loop, I ended up… 17 r/MachineLearning community 1mo ago Vendor-agnostic ML inference on production edge devices [R] I work on PostSlate, a video editing tool, and this comes out of our own work. We run ML models on-device, face detection and embedding among other things, which means we can't assume anything about the user's GPU. NVIDIA discrete, AMD, Intel integrated, Apple Silicon, all of… 7 Hugging Face Daily Papers research 1mo ago UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Abstract Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision… 11 r/LocalLLaMA community 1mo ago What’s the maximum physical amount of intelligence we can fit into small models? I was talking to my friend the other day, who is a really avid supporter of local models. He thinks we might be able to get something as intelligent (not smart in terms of how much it knows) as Claude Fable in inside a model which is like, 20 billion parameters, even if we’d get… 29 r/MachineLearning community 1mo ago Mix local LLMs, Claude Code, Codex, Gemini and more in one SDLC pipeline (open source) [P] A lot of AI coding tools assume one model should do everything: understand the task, write the code, review it, and decide whether it is correct. I ended up building something around the opposite idea. Instead of one model doing the whole software development lifecycle, every… 22 arXiv — NLP / Computation & Language research 1mo ago From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference arXiv:2607.24585v1 Announce Type: new Abstract: We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget… 30 r/LocalLLaMA community 1mo ago What local model do you still use after the hype wore off? Every time a new model is released, I tend to check it out. The benchmarks, readme, or whatever seem pretty convincing, so I download it, test it for a few hours, and then I just go back to the same couple of ones I already had. Curious what models people here have actually kept… 20 arXiv — Machine Learning research 1mo ago FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs arXiv:2607.21624v1 Announce Type: cross Abstract: Transformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains… 20 r/LocalLLaMA community 1mo ago Unexpected use of local llm I was refreshing my youtube and found out my favourite reviewer uploaded a battery test of 78 smartphones: https://youtu.be/MpgUFrsIWSQ the author said they started using robotic arm to simulate a person using the phone but they wanted to further enhance it by using agentic ai.… 19 r/LocalLLaMA community 1mo ago [OSS] Use case only possible with local inference at its core: an on-device LLM understands your entire life, then proactively offers to get your work done through computer use! Open-source & free :D Hey r/LocalLLaMA ! :D I wanna share a really cool fully OSS thing I've been building that's only possible with local models: truly proactive AI! All your existing LLM systems waits for a prompt. Truly proactive AI has to read your entire life, every single day (every file,… 19 Page 2 of 9 · 407 articles ← Newer Older →