News / #edge Tag Edge 407 articles archived under #edge · RSS Sign in to follow r/LocalLLaMA community 1mo ago Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090? curious how people here approach buying high-end/workstation cards (RTX 6000 Ada, 5000 Ada, etc) for local LLM work, do you actively watch pricing/timing on these specifically, or is the consumer 5090 usually enough for most builds? also wondering if price alerts/tracking tools… 15 r/LocalLLaMA community 1mo ago We open-sourced Logue — a privacy-first macOS meeting-notes + writing app that runs on-device (MLX, Apple Silicon) entirely At Bitwize, we've been building Logue, a native macOS app for AI meeting notes and writing, and we just open-sourced it (MIT). We're sharing it here because the whole point is that it runs 100% on-device — we wanted something that could transcribe and summarize meetings without… 27 r/LocalLLaMA community 1mo ago Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding? I am a long time iOS app developer. In the last year I have been using Cursor+Claude/others to assist with app development. I am concerned that the current low pricing will disappear eventually. I am pricing out a new laptop with the intention of using local models instead. New… 35 r/LocalLLaMA community 1mo ago How much are you actually using your local models these days? Which ones do you reach for the most? I started tracking my local model usage about four weeks ago and was wondering if anyone else here keeps track of how much they use them. I’ve also been running some tests with the cheapest SOTA open-weight Chinese models via OpenRouter. Apart from that, I’m mainly using GPT-5.5… 4 r/LocalLLaMA community 1mo ago Who ONLY use local models? Please be honest. I would love to hear about guys really dedicated to local AI and who really reject subscriptions (especially to openai and anthropic). What do you use your model for?   submitted by   /u/takoulseum [link]   [comments] 37 Hacker News — AI on Front Page community 1mo ago Android May Soon Restrict On-Device ADB Article URL: https://kitsumed.github.io/blog/posts/android-may-soon-restrict-on-device-adb/ Comments URL: https://news.ycombinator.com/item?id=49045159 Points: 262 # Comments: 140 38 r/LocalLLaMA community 1mo ago DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report) Hi everyone! Over the past five months I've been working on DKV (DifferentialKV), an open-source project exploring KV-cache compression for long-context local LLM inference. The goal is to reduce KV-cache memory requirements through anchor-based representations, joint low-rank… 21 arXiv — Machine Learning research 1mo ago Information-Theoretically Secure Aggregation for Lightweight Federated Learning: Resilient to Dropouts and Adversaries arXiv:2607.20890v1 Announce Type: new Abstract: On-device federated learning (FL) enables privacy-preserving and personalized model training on resource-constrained devices such as smartphones and IoT nodes. To reduce communication cost, sign-based methods (e.g., signSGD)… 20 r/LocalLLaMA community 1mo ago I compared local models and different quants / config on a subset of swe-verified bench And gathered a lot of data. you can see them for yourself And For the most curious, there are additional details here In this graph, I regrouped the finetunes under their base models. but you can see the details in the page. The python code to generate those pages is obviously… 18 r/LocalLLaMA community 1mo ago CPU-only inference on a Celeron N5095 SBC: 6 models from 0.6B to 8B, benchmarked I wanted to know how cheap you can go and still run local models, so I ran Ollama CPU-only on a Youyeetoo X1S. It's a single-board x86 machine with a Celeron N5095 (Jasper Lake, 4C/4T, 15W), 16GB of RAM, and a 128GB NVMe, running Kali 2025.4. Base configs of this board go for… 28 r/LocalLLaMA community 1mo ago Cactus Hybrid: We taught Gemma 4 to know when it's wrong Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score… 28 r/LocalLLaMA community 1mo ago MindControl - llama.cpp fork to guide the reasoning process via injection during sampling The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system prompts are highly specific), their reasoning process is highly unreliable and… 7 r/LocalLLaMA community 1mo ago We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions We’re open sourcing an alpha release of NeuTTS-2E : an on-device TTS model with 125M active parameters and 7 controllable emotions. The goal was simple: when you select “angry,” “fearful,” or “happy,” the delivery should follow that instruction rather than whatever emotion the… 14 arXiv — Machine Learning research 1mo ago QScheduler: Adaptive Gradient Sampling for Zeroth-Order On-Device Training on INT8 NPUs arXiv:2607.18802v1 Announce Type: new Abstract: Zeroth-Order (ZO) optimization enables On-Device Learning (ODL) on NPU-equipped microcontrollers by estimating gradients through forward passes alone, bypassing the need for backpropagation primitives and reducing memory… 37 r/LocalLLaMA community 1mo ago Anthropic claims local models are stealing from it, meanwhile it pays $1.5B for theft $1.5 billion settlement largest known payout in U.S. copyright case Case part of a wave of lawsuits from copyright holders against AI companies Some authors and publishers opted out and continue separate cases against Anthropic   submitted by   /u/Terminator857 [link]… 18 r/LocalLLaMA community 1mo ago Today I learnt the power of LocalLlama DISCLAIMER: No Ai was prompted in the creation of this post. Today I had an experience that complely blew my mind, I just had to write it down. As a bit of background I have been dabbling prompting local models using LM studio for the better part of 18 months now, keeping up to… 38 arXiv — Machine Learning research 1mo ago Fully-sensorized smart-eyewear platform for on-device Machine Learning arXiv:2607.16222v1 Announce Type: new Abstract: This paper presents ARGO, a smart eyewear platform designed to bridge ergonomic comfort, high computational throughput, and energy efficiency. Unlike cloud-dependent solutions, ARGO leverages the STM32N6 microcontroller and its… 11 r/LocalLLaMA community 1mo ago Google has disappeared completely from the top 15 Google hasn't shipped a model recently that is capable of competing with Sol or Fable. The previous models were pretty disappointing and unreliable, it seems the more time goes on that they might have different strategies: - They might be going all-in on on-device inference for… 25 Hugging Face Daily Papers research 1mo ago Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark Abstract Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, individual sensors die, and signal dropouts cluster precisely when a… 13 r/LocalLLaMA community 1mo ago No-coding model? Mostly all of the local models these days are competing for coding benchmarks. Is there any lab that just focuses all of their attention on making a better creative writing model?   submitted by   /u/Lost_Care7289 [link]   [comments] 11 r/LocalLLaMA community 1mo ago Is it possible to run a local model focused solely on "intelligence" and outsource its "knowledge" to web searches? I'm looking to run a very lightweight local model that acts as the brain, handling the logic and comprehension, while hooking it up to a web search tool to act as its memory and knowledge base.   submitted by   /u/chucrutcito [link]   [comments] 26 arXiv — Machine Learning research 1mo ago From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation arXiv:2607.15552v1 Announce Type: new Abstract: Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability. Classical optimization enforces constraints but fails to capture… 15 r/LocalLLaMA community 1mo ago What’s your favorite underrated local model? What’s your favorite underrated local model that you actually use every day? I’m not talking about the mainstream choices like Qwen 3.6 or Gemma 4. I’m looking for the hidden gems that deserve more attention. What do you use it for, and what hardware are you running it on? I’m… 20 r/LocalLLaMA community 1mo ago Qwen vs Gemma Hi! Been doing some local LLM stuff, and I can't help but notice: despite vastly-superior benchmark scores, Qwen 3.6 35a3B feels... substantially less intelligent than Gemma 4 26a4B (QAT). In terms of prompt adherence, output coherence, and just general "sanity", Gemma seems… 4 r/LocalLLaMA community 1mo ago Sharing MiniBot v2, this is what I'm currently using I gave it a major update so I thought I'd share. I make things that do work for me, always have... and this is the latest. It a single file 20k lines ;P https://github.com/illsk1lls/MiniBot (i previously posted v1 of this, which was not WPF) This is for local models only. Although it is OpenAI compatible. Have your favorite AI scan it to make sure its safe enough for you... Tools are enabled as… 28 r/LocalLLaMA community 1mo ago TUI app building on Rust How are people building nice looking TUIs with their local models? So far I’ve had zero luck, and my TUIs in Rust that local models build look like shiet   submitted by   /u/Infinite-Ad4512 [link]   [comments] 27 r/LocalLLaMA community 1mo ago Local LLM project Is it worth running local models on this old beast? Dell PowerEdge R710 (2009-12 era) Dual Xenon 5500 (I'm pretty sure) 48Gb DDR3-1066   submitted by   /u/Motor-Independent572 [link]   [comments] 18 r/LocalLLaMA community 1mo ago A year ago you told me my open-source screen-watching app was flaky. You were right, so I spent the year fixing it with your feedback. Thank you r/LocalLLaMA c: !! TL;DR: This post is part update, mostly thank you for your support :)) Observer is an open-source app that lets local LLMs watch your screen and notify you (WhatsApp/SMS/email/Discord) when something happens. A year of your feedback later , setup went from "flaky and very… 31 arXiv — Machine Learning research 1mo ago PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference arXiv:2607.14618v1 Announce Type: new Abstract: CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs. We… 23 r/LocalLLaMA community 1mo ago I’m taking a break I’ve just formatted my MacBook Pro with an M2 Max and 32GB of RAM. I’m taking a break from local LLMs. I’ve spent the last few months tinkering with local models, MTP, quants, harnesses, Gemma, Qwen… It’s time to take a break. I noticed this was becoming less of a hobby and more… 28 r/LocalLLaMA community 1mo ago I just got my first GPU that can actually run an LLM (Laptop 5090 24 GB) what do I play with first? I figure I can run 30b or even 72b models on it, but this is my first time running my own local LLM. I want to play with prompt engineering and see the most complex things I can get it to do. Any hot tips? I know it all moves fast and this sub is wired in I have 32GB ram also… 24 Hugging Face Daily Papers research 1mo ago PalmClaw: A Native On-Device Agent Framework for Mobile Phones Abstract Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task… 19 r/LocalLLaMA community 1mo ago If you had a 384GB (4x Blackwell), what model would you put on it and why? Hey guys, Company I work for is actually very interested in spending the money to host our own local model for the team. We expect probably 2-3 super users and at the worst case 10-20 concurrent users. The LLM would be mostly used for internal company policies/data management… 10 arXiv — Machine Learning research 1mo ago Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark arXiv:2607.11915v1 Announce Type: cross Abstract: Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, individual sensors die, and… 19 arXiv — NLP / Computation & Language research 1mo ago PalmClaw: A Native On-Device Agent Framework for Mobile Phones arXiv:2607.13027v1 Announce Type: new Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or… 25 arXiv — NLP / Computation & Language research 1mo ago On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage arXiv:2607.12257v1 Announce Type: cross Abstract: On-device research agents search a corpus, read sources, and write a cited brief on a personal laptop. Whether their citations are faithful, and at what cost, is unmeasured for a deployable small model. This study fixes one 4B… 13 llama.cpp releases dev-tools 1mo ago b10007 opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable ( #25639 ) opencl: do not fail backend init on devices without cl_khr_integer_dot_product opencl: do not call dp4 kernels when dp is unavailable Co-authored-by: Li He lih@qti.qualcomm.com… 10 r/LocalLLaMA community 1mo ago Good podcasts I'm heading on vacation soon and want to download a few good podcasts about local LLMs, open-weight models, inference, tooling and the broader open-source AI ecosystem. Which podcasts or specific episodes do you genuinely recommend? I'm especially interested in technical… 18 r/LocalLLaMA community 1mo ago Hermes Agent with local LLMs Hey all, I have a 4090 + 5060ti setup with 64gb ram. I have been using opencode with MiMoV2 pro and its great but I don't think spending $20 daily is financially sound decision. I heard that using qwen3.6-27b with hermes goes into loops, have any of you successfully used any… 30 arXiv — NLP / Computation & Language research 1mo ago Workload-Driven Optimization for On-Device Real-Time Subtitle Translation arXiv:2607.09957v1 Announce Type: new Abstract: This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inference, low latency, and privacy constraints. These conditions limit the value of… 23 r/LocalLLaMA community 1mo ago Excel work - best model I’ve been testing various local models for excel tasks related to my job. So far, Deepseek v4 Flash using DS4 has been best. Gets 30-40 tk/s with medium context and generally produces good results. However, looking to see what others may have had good experiences with. Need to… 23 r/LocalLLaMA community 1mo ago This is why we need local models and opensource harnesses   submitted by   /u/Comfortable-Rock-498 [link]   [comments] 19 arXiv — Machine Learning research 1mo ago On-Device Adaptive Battery Power Prediction for Electric Vehicles arXiv:2607.09400v1 Announce Type: new Abstract: Adaptive power management in Electric Vehicles (EVs) requires accurate power prediction. Although deep learning models have emerged as highly effective for time-series forecasting in this domain, their performance is prone to… 17 r/LocalLLaMA community 1mo ago How is Codex as a harness for local models? I was surprised to see that Codex is actually open source, so to my understanding, if you use a local model, it works fully locally. How does it compare to the other popular harnesses like Pi Code and Open Code?   submitted by   /u/A_Wild_Entei [link]   [comments] 4 llama.cpp releases dev-tools 1mo ago b9974 cuda: Don't crash when querying memory on device with no free memory. ( #25157 ) If a Cuda device has no or limited available memory, the actual call to cudaMemGetInfo() itself can cause a fatal crash due to a cuda out of memory error (there is not enough memory to actually… 38 r/MachineLearning community 1mo ago Zer0Fit: I took Google's new TabFM & TimesFM ML foundation models and made them available as an MCP server for zero-shot ML tasks (forecasts / classifications / regressions). 100% local. [P] TL:DR: I’m a grad student in AI, I saw that Google released TabFM and TimesFM last week, I built an MCP wrapper to serve both transformer models in a single Docker container so you can connect their new ML transformer models to a local LLM via Open WebUI, Claude Code, or Codex… 10 r/LocalLLaMA community 1mo ago Working around Qwen3.6-27B's tool-call failures and looping Let's start a discussion about what can be done to make local models more reliable. I've been using Qwen3.6-27B a lot lately, and have noticed the same thing that many others talk about here, which is the tool-call failures and looping that really gets in the way of being able… 33 r/LocalLLaMA community 1mo ago Zer0Fit: I took Google's new TabFM & TimesFM ML foundation models and made them available as an MCP server for zero-shot ML tasks (forecasts / classifications / regressions). 100% local. TL:DR: I’m a grad student in AI. I saw that Google released TabFM and TimesFM last week. I built an MCP wrapper to serve both transformer models in a single Docker container so you can connect their new ML transformer models to a local LLM via Open WebUI, Claude Code, or Codex… 18 r/LocalLLaMA community 1mo ago Opencode Agents vs Claude Code I’ve been playing around with Opencode and realized how 70% of the capability of my model comes from the agents I can use rather than the model size or parameters. So now obviously I have a question… is there a way to use Claude Code but have it pointing at my local model… 7 r/LocalLLaMA community 1mo ago Vellium v1.0.0 released: security hardening, wallpaper-based themes, JSON chat export and a major desktop stability pass Vellium has reached v1.0.0. It is a local-first desktop workspace for writing, roleplay, character creation, lorebooks and knowledge management with local LLMs. This release promotes the previous v1.0.0-beta build to the first stable version. The main focus was security and… 13 Page 3 of 9 · 407 articles ← Newer Older →