News / #pricing Tag Pricing changes 410 articles archived under #pricing · RSS Sign in to follow r/MachineLearning community 1mo ago Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R] Ran 10 realistic product tasks (classification, RAG QA, multi turn conversation, an agentic plan then execute task, etc) against the live APIs of OpenAI, Anthropic, Gemini and Kimi, using each provider's cost optimized tier. Total cost spread was 10.6x despite published rates… 15 Latent.Space news-outlet 1mo ago [AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro" a quiet day lets us highlight a new neolab win. 22 arXiv — NLP / Computation & Language research 1mo ago PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference arXiv:2607.20327v1 Announce Type: new Abstract: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware… 21 r/LocalLLaMA community 1mo ago Got these baddies in the mail today (2X 3080 20GB) About to plug them in. Currently running a single 3090. I got these for less than the price of a single 3090. 24GB wasn't enough for my use case, so 40 GB should be an upgrade. Going to throw my 3090 on ebay very likely. I feel like it's a perfect time to sell since the prices… 10 arXiv — Machine Learning research 1mo ago Exposure-Based Reinforcement Learning to Rank arXiv:2607.18689v1 Announce Type: new Abstract: Reinforcement learning (RL) methods for learning-to-rank (LTR) can optimize (almost) any ranking goal, e.g., from precision or discounted cumulative gain to fairness-of-exposure or ranking distillation. However, standard RL is… 26 arXiv — NLP / Computation & Language research 1mo ago The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation arXiv:2607.19226v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the… 9 r/LocalLLaMA community 1mo ago Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro Model Size Terminal-Bench 2.1 SWE-bench Multilingual SWE-Bench Pro (Public Dataset) DeepSWE SWE Atlas (Codebase QnA) Toolathlon Verified Laguna S 2.1 118B-A8B 70.2% 78.5% 59.4% 40.4% 46.2% 49.7% Finally the banger we've been waiting from Laguna. probably will be great for 64GB+… 33 Ars Technica — AI news-outlet 1mo ago Google reveals faster and cheaper Gemini 3.6 Flash, says 3.5 Pro is still in testing There are new 3.6 and 3.5 models today, but Google is already training Gemini 4. 35 arXiv — Machine Learning research 1mo ago RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce arXiv:2607.16230v1 Announce Type: new Abstract: Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversion. In practice, shipping cost is shaped not only by distance but also by destination demand… 37 arXiv — NLP / Computation & Language research 1mo ago PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer arXiv:2607.17620v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) makes finetuning large language models cheaper by adding to each weight matrix a trainable low-rank update parameterized as the product of two matrices. These matrices are usually trained with Adam,… 23 Vercel — AI dev-tools 1mo ago Vercel MCP now supports purchases Vercel MCP now supports purchasing Vercel products. You can: Upgrade your team to the Pro plan Add prepaid credits for v0 (requires a paid v0 plan) or AI Gateway Purchase the SIEM add-on (requires an Enterprise plan) Purchase and register a domain Vercel MCP quotes the price,… 32 Simon Willison community 1mo ago Reverse-engineering is cheap now I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes. I think this is an interesting illustration of the impact of the reduced cost of writing code. Prior to agents, it was entirely possible to reverse-engineer home… 33 r/LocalLLaMA community 1mo ago So what happened with OpenClaw? It had an insanely meteoritic rise. It felt like it was the only thing anyone had been talking about for months. Then just, everyone stopped talking about it. Usage based pricing inevitably came and it seems like it was killed over night. Competitors were also rushing to get… 22 r/LocalLLaMA community 1mo ago DeepSeek v4 flash release version appears to have been activated on api. Open weights imminent? https://np.reddit.com/r/DeepSeek/s/skO7urrE2C DS4 sort of came and went from the spotlight. The consensus seemed to be that its most notable feature is its price, and then we got distracted by the next big releases. However, people seem to forget this was only the preview… 32 arXiv — Machine Learning research 1mo ago Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching arXiv:2607.15516v1 Announce Type: new Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware… 33 Hacker News — AI on Front Page community 1mo ago Better and Cheaper Than IPTV Article URL: https://github.com/stupside/castor Comments URL: https://news.ycombinator.com/item?id=48964015 Points: 226 # Comments: 63 25 r/LocalLLaMA community 1mo ago How are y’all stomaching the “AI Boom” prices? I am in the middle of considering an upgrade to my Home Server. I want to get a decent GPU for LocalAI. I mean, I have an RTX 3060 TI 16GB, I know that is much more than most people have, but even though it has a lot of cram the bus width is really slowing it down - I get ~23… 26 r/LocalLLaMA community 1mo ago What kind of dark magic is Deepseek using? I was taking a look at Kimi K3 scores on the Artificial analysis leaderboard and was quite baffled when I saw this chart. Granted, Deepseek has always been the king of price to performance, but this is still incredible. Is it just API subsidization or have they optimized their… 4 TechCrunch — AI news-outlet 1mo ago AI-driven memory crunch jolts India’s smartphone market India's smartphone slowdown highlights how the AI boom is reshaping consumer electronics, from pricing and demand to corporate strategy. 4 r/LocalLLaMA community 1mo ago Kimi K3 one-shotting a racing game Kimi k3 is the goat, especially in UI and web dev. It provided a result matching fable 5, 70% cheaper. Opencode and Kimi just became my default setup now. This is my first time actually coding outside Claude Code and Codex. Any reason to go back, or am I missing something?… 11 r/LocalLLaMA community 1mo ago Any news about MiMo V3? It's truly the most exciting model. Kimi K3 and GLM 5.2 look impressive and will likely be quite helpful for distilling models. But for everyday work, paying for APIs, I would only use something that is an order of magnitude cheaper, like MiMo and DS.   submitted by  … 9 arXiv — NLP / Computation & Language research 1mo ago Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel arXiv:2607.14431v1 Announce Type: new Abstract: We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later… 10 Hugging Face Daily Papers research 1mo ago Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel Abstract We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored, by graft, into a fresh… 17 Latent.Space news-outlet 1mo ago [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing a great week for open models continues. 37 Vercel — AI dev-tools 1mo ago GLM 5.2 is 35% off via Novita on AI Gateway GLM 5.2 is 35% off on AI Gateway through July 24 when routed through Novita. To get the discounted rate, set the model to zai/glm-5.2 in the AI SDK and route requests through Novita: After July 24, the model stays available at standard provider rates with no markup. Try GLM 5.2… 24 Hacker News — AI on Front Page community 1mo ago Kimi K3: Open Frontier Intelligence https://www.kimi.com/en Kimi K3 Intelligence, Performance & Price Analysis: https://artificialanalysis.ai/models/kimi-k3 Comments URL: https://news.ycombinator.com/item?id=48935342 Points: 1341 # Comments: 827 6 TechCrunch — AI news-outlet 1mo ago SpaceX slips below its $135 IPO price ahead of Starship launch The stock has steadily fallen from the euphoric post-IPO high, showing that markets may be sobering up to the promises CEO Elon Musk made before and after SpaceX went public. 25 TechCrunch — AI news-outlet 1mo ago SpaceX falls to $135 IPO price ahead of Starship launch The stock has steadily fallen from the euphoric post-IPO high, showing that markets may be sobering up to the promises CEO Elon Musk made before and after SpaceX went public. 35 Hacker News — AI on Front Page community 1mo ago SpaceX bond worth 10% less than issue price – heading for junk bond status Article URL: https://www.ft.com/content/3a023b95-66c3-41e1-b0ce-df752a499541 Comments URL: https://news.ycombinator.com/item?id=48920181 Points: 202 # Comments: 102 33 arXiv — NLP / Computation & Language research 1mo ago Speculate with Memory: Lossless Acceleration for LLM Agents arXiv:2607.12236v1 Announce Type: cross Abstract: Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle. However, existing speculators are stateless and discard all information between… 19 r/LocalLLaMA community 1mo ago Big Tech who built their empires on web scraping. Crying "existential threat" over model distillation is peak irony I just don't get it. These big tech companies can illegally scrape the entire internet and gatekeep their better models behind higher prices. So it's natural that people look for affordable options, and there will be providers who apparently distill models from themThe hypocrisy… 27 arXiv — Machine Learning research 1mo ago Quota Marketplace: Dynamic Pricing for Efficient Allocation of ML Training Resources arXiv:2607.09802v1 Announce Type: new Abstract: The escalating demand for Machine Learning (ML) training resources in recent years has resulted in a substantial gap between the high demand and the available supply. Efficient allocation of these scarce and expensive resources is… 23 arXiv — Machine Learning research 1mo ago Optimizing ARDL Models for Retail Sales Forecasting and Fair Pricing arXiv:2607.09956v1 Announce Type: new Abstract: Pricing food products to balance profitability with consumer welfare is a central challenge for retailers. Dynamic pricing is widely used to maximize revenue, yet most pricing models optimize business objectives while overlooking… 26 r/LocalLLaMA community 1mo ago I just don't get it. These big tech companies can illegally scrape the entire internet and gatekeep their better models behind higher prices. So it's natural that people look for affordable options, and there will be providers who apparently distill models from them. The irony? They cry existential threat when they were the ones who made us feel that way first.   submitted by   /u/Blue-Sea2255 [link]   [comments] 33 Don't Worry About the Vase community 1mo ago Better Call Sol The Workhorse OpenAI’s GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna. 31 TechCrunch — AI news-outlet 1mo ago Anthropic starts localizing Claude pricing for India, its biggest market after the US Claude users in India are starting to see Indian rupee-denominated subscription plans. 38 Vercel — AI dev-tools 1mo ago Open-weight models surge to 29% of volume, price per token flattens AI Gateway Production Index — July 2026 Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs, giving us a view of what AI usage actually looks like in today’s enterprise. We publish that view here. See the Production Index… 13 arXiv — Machine Learning research 1mo ago DaDaDa: A Dataset for Data Pricing in Data Marketplaces arXiv:2607.08785v1 Announce Type: new Abstract: High-quality data drives machine learning advances across industries. Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplaces, e.g., AWS Marketplace, Databricks, and Datarade.… 36 Vercel — AI dev-tools 1mo ago Web Analytics and Speed Insights are now more cost-efficient Web Analytics and Speed Insights are now almost 10% more cost-efficient for all Pro and Enterprise teams. Infrastructure and event-processing improvements reduced cost across both products, so page views, referrers, top routes, and Core Web Vitals are now processed more… 18 Hacker News — AI on Front Page community 1mo ago Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper Article URL: https://ploy.ai/blog/migrating-a-production-ai-agent-to-gpt-5-6 Comments URL: https://news.ycombinator.com/item?id=48882716 Points: 206 # Comments: 88 17 r/MachineLearning community 1mo ago Obtaining Irregular Learning Curves with HyberBand Tuned ANN model for Price Prediction [P] I have used Hyperband automatic tuning for an ANN model to predict price. After running HyberBand automatic tuning to get the 'best' architecture, I am obtaining a strange Val/Training loss learning curve. I cannot figure out if this is due to an error within the code or just a… 24 r/LocalLLaMA community 1mo ago The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico Politico: Wall Street’s new obsession: Which CEOs have Trump’s ear?: https://www.politico.com/newsletters/politico-influence/2026/07/10/wall-streets-new-obsession-which-ceos-have-trumps-ear-00993324   submitted by   /u/Nunki08 [link]   [comments] 29 r/LocalLLaMA community 1mo ago According to DataBricks, pi-coding-agent is ~2x cheaper than CC/Codex, GLM 5.2 on par with Opus 4.8 high https://preview.redd.it/3p60zyf8afch1.png?width=1840&format=png&auto=webp&s=10dcc90945f0db03352239579fca2132d0c90dfa https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase tl;dr pi-coding-agent (bash for everything/minimum tools) is up… 27 r/LocalLLaMA community 1mo ago Neuralwatt Pricing will Double From 07/16 - Got This Email Welp, the age of cheap tokens coming to an end. Many here use this service to access GLM 5.2. Looks like only sustainable way moving forward to local LLM.   submitted by   /u/BoogerheadCult [link]   [comments] 26 arXiv — Machine Learning research 1mo ago Spectral Analysis of Dueling Q-Learning arXiv:2607.08340v1 Announce Type: new Abstract: Q-learning is a fundamental algorithm in reinforcement learning (RL) for solving discounted Markov decision processes (MDPs) when the transition kernel is unknown. The deep Q-network (DQN) extends Q-learning by using a deep neural… 31 llama.cpp releases dev-tools 1mo ago b9941 Only index by compile times + always multiply/add ( #25445 ) The first one avoids relying on compile to optimize local memory away, and the second is cheaper than issuing control flow statements macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)… 28 Simon Willison community 1mo ago The new GPT-5.6 family: Luna, Terra, Sol OpenAI's latest flagship model hit general availability this morning , and comes in three sizes: Luna, Terra, and Sol (from smallest to largest). The new models are priced per 1M input/output tokens as Luna $1/$6, Terra $2.50/$15, Sol $5/$30. For comparison, the Claude Opus… 35 r/LocalLLaMA community 1mo ago Help for server quotation with RTX 6000 Pro (France) I would like to make a quotation for a server with RTX 6000 Pro (96 GB). Rackable and tower variants. I do prefer reliable hardware than lowest price. I am looking for suggestions: vendor, model, points to check... Thank you for your help!   submitted by  … 18 Smol AI News news-outlet 1mo ago not much happened today **OpenAI** launched the **GPT-5.6** family including **Sol, Terra, and Luna** models, integrated across **ChatGPT, Codex, and API** with immediate rollout. The release emphasized improved **performance-per-dollar** with pricing matching GPT-5.5 but better capabilities,… 21 Smol AI News news-outlet 1mo ago OpenAI launches GPT 5.6 Sol/Terra/Luna **OpenAI** launched the **GPT-5.6** family with three models: **Sol**, **Terra**, and **Luna**, integrated across **ChatGPT**, **Codex**, and the API. Pricing tiers range from **$1 to $5 per million tokens** with new cache-write pricing and a 90% cache-read discount. The launch… 21 Page 4 of 9 · 410 articles ← Newer Older →