My reasons to run local models
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
- I can finetune any model on any dataset I want.
- I can use techniques like speculative decoding and other sota approaches to get the max tps
- The llm provides like anthropic and openai are not getting access to my data
- The hardware is reusable for vision text speech, and I can run any blend of models for free as much as I want
- I can curate any dataset/content that I want without worrying about the costs
- I like watching Dario go up in flames
[link] [comments]
More from r/LocalLLaMA
-
Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
Aug 30
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.