News / #gpu Tag Gpu 500 articles archived under #gpu · RSS Sign in to follow The Information — AI news-outlet 3d ago Bill Gates’ AI Warning and Nvidia’s Boffo Quarter You couldn’t miss the disconnect in the AI discussion among major tech figures on Wednesday. Microsoft co-founder Bill Gates warned in a lengthy essay about the dangers posed by the technology , noting that he would support a global slowdown in AI advances if someone had a… 8 Vercel — AI dev-tools 3d ago Ling 3.0 Flash Fin now available on AI Gateway for free Ling 3.0 Flash Fin from Inclusion AI is now available on AI Gateway, free to use through September 25. Ling 3.0 Flash Fin is a finance-focused version of Ling 3.0 Flash . It has a 256K token context window, produces up to 32K output tokens, and supports reasoning and function… 29 Ollama releases dev-tools 3d ago v0.33.1 What's Changed MLX: Qwen3.8 Flash Next support cmake: make external compat patches idempotent MLX and llama.cpp update mlxrunner: add structured output support mlxrunner: avoid Metal GPU timeouts when loading models from slow storage New Contributors @pd95 made their first… 6 TechCrunch — AI news-outlet 3d ago Amazon just tripled its order of Nvidia chips over ‘surging demand’ Amazon is adding another 2 million Nvidia GPU chips to its data centers over the next two years. But this extended partnerships stretches beyond buying more chips. 28 r/MachineLearning community 3d ago A dataset with 52 Text to image model evaluation [P] I created a simple text to image benchmark. I curated 192 prompts that are difficult for T2I models in various ways: text rendering, spatial reasoning, human realism, negations, etc... I then asked a VLM to judge every output against a pre-specified binary question with the… 30 NVIDIA Developer Blog official-blog 3d ago NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,... 33 The Information — AI news-outlet 3d ago Nvidia Says Revenue Rose 106% in July Quarter, But Customers Were Slower to Make Payments Nvidia’s revenue rose 106% to $96.2 billion in the three months that ended in July, or 21 percentage points higher than the growth it reported in the previous quarter. The company said growth would cool a bit, to 89.5%, in the current fiscal quarter. But given its recent… 37 r/LocalLLaMA community 3d ago Anyone else doing eGPUs (OCuLink)? Upgraded to a 5070 Ti so I could run Qwen 3.8 27B, which works perfectly, but didn't want to let the old 4070 Ti go to waste. The cards would touch if I put them both in the PC and I knew the heat would be awful from my crypto mining days. I always was curious about eGPUs so I… 33 r/LocalLLaMA community 3d ago Are models with N-Gram tables going to completely change the AI race? The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they could open the door for 1T+ parameter models to be run on a single server with modest GPUs and a ton of system RAM rather than… 22 NVIDIA Developer Blog official-blog 3d ago Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s... 11 NVIDIA Developer Blog official-blog 3d ago Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s... 36 r/LocalLLaMA community 3d ago A minecraft clone I fully vibecoded with Qwen3.8-27b Q4 I wanted to see just how capable Qwen3.8-27b is locally. I have a RTX 4090 and 96GB of RAM but the Q4 comfortably fits in the GPU with plenty of context, the few times I needed more than 130k context I just loaded it spilled into RAM and it's capable of not degrading even at… 5 llama.cpp releases dev-tools 3d ago b10635 cuda: unblock mmq for MoE on sm_60 ( #26264 ) cuda: unblock mmq for MoE on sm_60 cuda: duplicate mmq-config-pascal for dp4a and older cuda: reduce occupancy on non-dp4a pascal for Q2_K, Q4_K, Q5_K, Q6_K Website: https://llama.app Attestations:… 30 Stratechery (Ben Thompson) community 3d ago Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia. 30 r/LocalLLaMA community 4d ago Underrated Muse Glimmer Benchmarked qwen3.8 xhigh, medium and muse glimmer. Xhigh effort mode with qwen3.8 took almost 30hrs. (And still failed on 16 cases because of the 32K output token limit) Medium effort mode and muse glimmer were 3-4 hours each. But I'm actually surprised by the muse glimmer… 31 The Information — AI news-outlet 4d ago OpenAI Says Its Jalapeño AI Chip Is Better Than Nvidia’s Blackwell OpenAI on Tuesday released an evaluation of its new Jalapeño AI server chip, claiming it is faster and more powerful than Nvidia’s flagship Blackwell chips. The results appear promising for OpenAI, which hopes to reduce its reliance on Nvidia’s hardware someday, but it’s still… 24 r/LocalLLaMA community 4d ago Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs So here's the thing, almost everyone use NVIDIA to run their LLMs, we also do the same, a lot of people we've met use like RTX PRO 6000 or even H100, B300 It seems like everyone eyes is looking at NVIDIA. However we do the math that the raw power alone on AMD GPU MI350X is… 9 arXiv — Machine Learning research 4d ago Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders arXiv:2608.23809v1 Announce Type: new Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We… 31 arXiv — Machine Learning research 4d ago A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU arXiv:2608.24067v1 Announce Type: new Abstract: A self-organising map turns a large corpus into a browsable two-dimensional atlas, but building one at MEDLINE scale has been impractical: the best-matching-unit (BMU) search that dominates training is bound by the bandwidth needed… 26 arXiv — Machine Learning research 4d ago Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning arXiv:2608.24386v1 Announce Type: new Abstract: Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for such outputs remains an open challenge. While E(3)-equivariant neural networks excel at point estimates, they lack rigorous… 36 arXiv — Machine Learning research 4d ago It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces arXiv:2608.24518v1 Announce Type: new Abstract: Uncertainty Quantification (UQ) plays a vital role in enhancing the reliability of deep learning model predictions, especially in scenarios with high-dimensional output spaces. This paper addresses the dual nature of uncertainty --… 8 Hugging Face Daily Papers research 4d ago WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Abstract Off-policy reinforcement learning stabilizers vary with data availability, motivating regime-aware algorithms that adapt normalization and Q-function clipping to improve efficiency across CPU and GPU-parallel training. Generated by thinkingmachines/Inkling-Small… 19 Hugging Face Daily Papers research 4d ago TorchMorph: CUDA-accelerated Morphological Transforms Abstract TorchMorph is a PyTorch extension providing GPU-accelerated morphological and distance-transform operators across up to eight dimensions with a SciPy-compatible API. Generated by thinkingmachines/Inkling-Small Morphological transforms are long-standing tools for shape… 14 r/LocalLLaMA community 4d ago Qwen3.8-27B IQ3_XXS wrote a correct multilayer TMM on a 16 GB Quadro — after 100 minutes, 3 compactions, and 108k output tokens https://preview.redd.it/i8rjx0ar5mlh1.png?width=2160&format=png&auto=webp&s=c2588bc7b2519ea71b176ca73faf566dfc585496 I wanted to see whether a heavily quantized 27B model running entirely on an older 16 GB workstation GPU could do more than the usual coding demos. FFT felt too… 8 The Information — AI news-outlet 4d ago Why OpenAI’s Jalapeño Might Sicken Nvidia Exactly who at OpenAI decided to name the company’s first internally designed chip Jalapeño? That might have seemed like a good idea a few months ago—perhaps OpenAI imagined that name would serve as a subtle indicator that the chip is a hot property. But it’s not ideal right… 36 NVIDIA Developer Blog official-blog 4d ago Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into HBM from storage, compiling kernels,... 10 r/LocalLLaMA community 4d ago CNBC Television: Nvidia partner with Perplexity AI to run locally in DGX Spark.   submitted by   /u/dd32x [link]   [comments] 8 NVIDIA Developer Blog official-blog 4d ago CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and... 31 r/LocalLLaMA community 4d ago 12 abliterated Gemma 4 12B variants, one base, 165 GPU hours - Abliterlitics I ran 11 uncensored variants of Gemma 4 12B that I grabbed from huggingface, sorting by downloads. 10 full abliterations plus 2 LoRA adapters which were requested to be added in the comparison, against the official base. 165 GPU hours over three and a half weeks on a single… 29 Hacker News — AI on Front Page community 4d ago OpenAI Jalapeño: Better than Nvidia Blackwell https://www.bloomberg.com/news/articles/2026-08-25/openai-cl... , https://archive.ph/yCTrr Comments URL: https://news.ycombinator.com/item?id=49434378 Points: 279 # Comments: 191 24 r/LocalLLaMA community 4d ago Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro "A 12-core GPU, also with two more cores than before, now includes Neural Accelerators in each core for the first time on Mac mini, resulting in up to 4x faster AI performance and 2x faster graphics than Mac mini with M4. In addition, the all-new Dual 16-core Neural Engine… 31 arXiv — NLP / Computation & Language research 5d ago On the Role of Citations in Preference Data arXiv:2608.21376v1 Announce Type: new Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model… 15 arXiv — NLP / Computation & Language research 5d ago More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers arXiv:2608.21806v1 Announce Type: new Abstract: Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact remains unclear. We analyze 13,921 ACL, EMNLP, and NAACL main-conference papers published… 13 arXiv — NLP / Computation & Language research 5d ago Grounded Normative Rule Generation with Structured Search arXiv:2608.22229v1 Announce Type: new Abstract: Normative rules like institutional charters and workplace policies must be both human-readable and operationally verifiable against actual environment records. However, current language generation and structured-output benchmarks… 14 arXiv — NLP / Computation & Language research 5d ago Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs arXiv:2608.22367v1 Announce Type: new Abstract: Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies… 4 arXiv — NLP / Computation & Language research 5d ago Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification arXiv:2608.22506v1 Announce Type: new Abstract: Claim-level Uncertainty Quantification (UQ) aims to mitigate the lack of reliability of Large Language Models (LLMs) by evaluating the factuality of each claim in their outputs. We introduce Kernel Token Contradiction (KTC), a… 38 arXiv — NLP / Computation & Language research 5d ago SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning arXiv:2608.22857v1 Announce Type: new Abstract: Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names,… 20 r/LocalLLaMA community 5d ago To all of you who have bought Chinese ASICs, how have they been? as we all know to get anything good and modern for nvidia/amd if you are lucky u can give up your kidneys as a down payment, but the chinese accelerators have a huge value proposition, if ur willing to invest the time and tokens porting frameworks to them. To whoever owns them,… 15 The Information — AI news-outlet 5d ago Nvidia's John Malone Dealmaking Opportunity Here’s a thought about Nvidia’s future: All the equity stakes it’s accumulating in AI companies—including, as we reported on Sunday , an expanded stake in Perplexity—could become more useful in ways other than simply guaranteeing demand for its chips. Over time the stakes could… 25 r/LocalLLaMA community 5d ago Planning to spend ~$100 benchmarking differnet Qwen3.8-27B quants and kv cache and looking for input before I start TL;DR: I'm planning to spend around $100 on cloud GPUs to benchmark Qwen3.8-27B with a focus on questions that actually matter when running it locally: different quant levels/providers, 8-bit vs 16-bit KV cache, GGUF vs EXL3, context length tradeoffs, and token efficiency on… 29 Ars Technica — AI news-outlet 5d ago Nvidia senior manager linked to Supermicro scheme smuggling AI servers to China Nvidia worker indicted after Jensen Huang scolded Supermicro for AI server smuggling. 23 Hacker News — AI on Front Page community 5d ago MS Paint and Photos inivisibly watermark even locally generated output with GUID Article URL: https://xusheng.dev/posts/reversing/mspaint_invisible_watermark/main/ Comments URL: https://news.ycombinator.com/item?id=49421158 Points: 281 # Comments: 130 23 NVIDIA Developer Blog official-blog 5d ago Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,... 34 The Information — AI news-outlet 5d ago Filling In the Gaps in Nvidia’s $500 Billion Financing Pitch Nvidia shocked the market earlier this month when it lined up half a dozen financial giants, including Blackstone, Apollo and Goldman Sachs, to finance $500 billion of AI infrastructure purchases. The eye-popping figure was a signal to the market that Nvidia is moving away from… 10 The Information — AI news-outlet 5d ago Nvidia Announces New Customers For Vera CPU, Groq LPX Racks SpaceX and AI cloud company Nebius are set to be early customers of Nvidia’s Vera central processing units and fast inference-focused Groq LPX racks, two key items intended to expand on Nvidia’s main, graphics processing unit-based products. Analysts and investors have wondered… 32 NVIDIA Developer Blog official-blog 5d ago NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt AI agents have expanded inference from single-turn interactions into multi-step workflows that reason, invoke tools, coordinate subagents, and carry growing... 17 NVIDIA Developer Blog official-blog 5d ago NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agentic AI factories connect diverse users,... 7 NVIDIA Developer Blog official-blog 5d ago Solving Agentic AI Fleet Challenges with NVIDIA Vera CPU AI factories are interconnected systems where fleet economics depend on how efficiently the entire stack converts power and capital into completed agent tasks.... 26 NVIDIA Developer Blog official-blog 5d ago How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin NVIDIA Groq 3 LPX is the interactive AI inference accelerator for the NVIDIA Vera Rubin platform. At the core of the platform is NVIDIA Vera Rubin NVL72, the... 15 The Information — AI news-outlet 5d ago Jensen Huang's Manic Moves Nvidia CEO Jensen Huang continues to not rest on his laurels, particularly as clouds continue forming over the AI buildout. On Sunday, Phoebe, Valida and I reported that Nvidia was planning to lead a multibillion-dollar investment in Perplexity , the search firm-turned-AI agent… 14 Page 2 of 10 · 500 articles ← Newer Older →