News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow r/LocalLLaMA community 2mo ago Buying AI accelerators/GPUs in China... Bit of a long-shot this, but happens I'll be in China next week. Just wondering if there are any Chinese graphics cards/AI accelerators I should be trying to buy when I'm there? :-). I would be looking for something that let me run inference big models (so, lots of (V?)RAM), but… 10 The Information — AI news-outlet 2mo ago Exclusive: Nvidia Server Marketplace Startup Raises $100 Million at $800 Million Valuation Data center software startup and AI-server broker Hydra Host has raised $100 million at a valuation of close to $800 million, led by Kindred Ventures. Nvidia, Cathie Wood’s ARK Invest, early CoreWeave backer Magnetar, and existing investors Founders Fund and Flume Ventures also… 26 arXiv — Machine Learning research 2mo ago Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport arXiv:2606.14157v1 Announce Type: new Abstract: Cities deliver basic services through mixed public-private facility networks, including schools, clinics, transit providers, and subsidized service points. In these systems, planners often observe where households go, but not the… 9 Ars Technica — AI news-outlet 2mo ago $130 billion in data center projects blocked by protests so far this year Winning fight against AI data centers gives people a "taste of political power." 6 Ars Technica — AI news-outlet 2mo ago When it comes to total water use, AI data centers are a drop in the bucket Even moderately sized data centers can have an outsized local impact. 23 The Information — AI news-outlet 2mo ago Meta Bought Rivos to Accelerate Its AI Chip Push. It Isn’t Working. Meta Platforms bought semiconductor startup Rivos last year to accelerate development of in-house chips and reduce its reliance on Nvidia as it pours cash into data centers for its AI ambitions. Now six months since the acquisition closed, Meta is struggling to make it work,… 10 The Information — AI news-outlet 2mo ago Nvidia Pitches Vera CPU to Chinese Customers Nvidia is pitching Chinese customers on its new Vera central processing units for AI data centers, telling them the chips could be available as soon as August and that orders can begin now, Reuters reported, citing three people familiar with the matter. The push gives Nvidia… 7 Hugging Face Daily Papers research 2mo ago Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering Abstract Flash-GMM introduces an efficient fused Triton kernel for Gaussian Mixture Models that achieves significant speedup and enables processing much larger datasets on a single GPU. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present Flash-GMM, a fused Triton kernel for… 18 The Information — AI news-outlet 2mo ago KKR, Nvidia, Others Launch $10 Billion Data Center Company Private equity firm KKR, the Kuwait Investment Authority, Nvidia and power generation company Vistra launched a new company on Thursday to finance and help build AI data centers. Nvidia’s role as an anchor investor in Helix signifies another extension of the AI giant’s growing… 29 The Information — AI news-outlet 2mo ago Anthropic Pursues First Data Center Leases, Seeks Financial Backing From Google Anthropic is moving forward with a plan to control its own servers for developing AI, giving it the ability to cut its computing costs in the long run. The maker of Claude in recent months has signed more than a dozen initial agreements, known as letters of intent, to lease data… 20 r/LocalLLaMA community 2mo ago How I implemented ASR bias for voice transcription models [Open Source] I've been spending the last couple of weeks building a Wispr Flow clone as an open source project. For context, it is a voice dictation app that lets you type faster, by speaking instead of actually typing. I spent the first week building the basic STT capabilities. One of the… 29 r/LocalLLaMA community 2mo ago Tiny Scale Is All I Can Spare To Play With Transformer Hi! I am a student from India, this is my first paper that I published. I was curious whether I can combine both Attention and FFN together to save parameters without sacrificing performance, specifically at parameters <= 10M. Basically my intuition was that Attention is dynamic… 32 arXiv — Machine Learning research 2mo ago Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity arXiv:2606.11431v1 Announce Type: new Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training. This raises a basic… 10 arXiv — Machine Learning research 2mo ago Efficient Time Series Clustering from Multiscale Reservoir Dynamics with Granular-Ball Anchoring Graph Optimization arXiv:2606.12077v1 Announce Type: new Abstract: Time-series clustering remains challenging due to the inherent trade-off between clustering effectiveness and computational efficiency. Similarity-based methods often suffer from quadratic complexity caused by pairwise distance… 15 arXiv — Machine Learning research 2mo ago Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders arXiv:2606.12138v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs. We study this question through \emph{feature… 19 arXiv — NLP / Computation & Language research 2mo ago Small Experiments, Cheaper Decisions: A Case Study in Staged Promotion for Micro-Pretraining arXiv:2606.11387v1 Announce Type: new Abstract: Short pretraining runs can reduce experimental cost, but they can also over-promote configurations that only look strong at tiny budgets. We study an auditable staged-promotion protocol for a fixed micro-pretraining runner on two… 9 r/LocalLLaMA community 2mo ago Tried to benchmark Google’s new on-device dictation models (Eloquent) and basically couldn’t I tried to benchmark Google’s new on-device dictation app (Eloquent) and basically couldn’t. It drops about half of my dictations. tl;dr Full results are 👉 here . Background: Google shipped a new fully‑local dictation app yesterday with proprietary new models , so I was excited… 5 Hacker News — AI on Front Page community 2mo ago Farmer donates land for a park, city sells it for $10M as data center land Article URL: https://www.tomshardware.com/tech-industry/farmer-donates-land-for-a-park-city-sells-it-for-data-center-development-usd10-gift-became-usd10m-for-city-government-with-usd30m-tax-expected-over-next-decade Comments URL: https://news.ycombinator.com/item?id=48481126… 32 NVIDIA Developer Blog official-blog 2mo ago Designing Production-Ready Battery Energy Storage Systems for AI Factories AI factories are changing what data-center infrastructure must do. Unlike traditional data centers, AI factories are built to manufacture intelligence at scale.... 29 TechCrunch — AI news-outlet 2mo ago The three hard-tech moonshots fueling SpaceX’s unbelievable IPO Most of the value in SpaceX's IPO is effectively a call option on the company's ambitious space data center plans. 23 OpenAI official-blog 2mo ago PRC-linked influence operations are targeting AI debates in the US A new report from OpenAI details PRC-linked influence operations using AI to target U.S. tech debates, data center narratives, tariffs, and false claims about ChatGPT. 7 r/MachineLearning community 2mo ago Should I Commit and Publish the Results? [R] Hello Reddit I've been working on QSPR (Quantitative Structure-Property Relationship) analysis for chemical compounds mentioned in the Jean-Claude Bradley Open Melting Point Dataset . Basically the idea is to see how accurate a model can predict melting points of compounds using… 32 TechCrunch — AI news-outlet 2mo ago Meta signs first AI data center deal in India with Reliance The 168-megawatt facility will support Meta's global AI computing needs and can be expanded over time. 17 arXiv — Machine Learning research 2mo ago FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses arXiv:2606.09878v1 Announce Type: new Abstract: Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks. We introduce FailureScope, a behavioral-diagnosis method that clusters evaluation probes by their cross-model… 20 arXiv — Machine Learning research 2mo ago Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters arXiv:2606.09924v1 Announce Type: new Abstract: Deploying deep neural networks on memory-constrained edge accelerators is bottlenecked by per-inference off-chip weight transfer rather than computation: the dense network cannot be retained on-chip, and every parameter must be… 29 arXiv — NLP / Computation & Language research 2mo ago Which LoRA? An Empirical Study on the Effectiveness of LoRA Techniques During Multilingual Instruction Tuning arXiv:2606.10428v1 Announce Type: new Abstract: We investigate whether commonly available LoRA variants have an advantage over basic LoRA in multilingual instruction tuning. Experiments involving LoRA and four other variants on two datasets across diverse target languages show… 9 arXiv — NLP / Computation & Language research 2mo ago Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis arXiv:2606.10381v1 Announce Type: cross Abstract: Muon collider research spans accelerator physics, detector instrumentation, and high-energy phenomenology, with relevant evidence scattered across a rapidly expanding and heterogeneous body of scientific literature. As… 37 arXiv — NLP / Computation & Language research 2mo ago SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference arXiv:2606.10445v1 Announce Type: cross Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity constraint often causes non-negligible accuracy degradation under post-training… 21 MIT News — AI research 2mo ago Startup’s nuclear-inspired cooling system could make data centers more sustainable Founded by two researchers from MIT, Ferveret reduces the amount of energy and water required to cool the chips that power AI. 17 Hugging Face Daily Papers research 2mo ago EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents Abstract EEVEE is a novel test-time prompt learning framework for LLM agents that handles heterogeneous data streams through task clustering and co-evolving router-prompt optimization. Generated by Qwen/Qwen2.5-Coder-32B-Instruct In this paper, we propose EEVEE, the first… 6 The Information — AI news-outlet 2mo ago OpenAI in Talks to Lease 10 Gigawatt Ohio Data Center with Backing From Nvidia OpenAI is in advanced negotiations to lease a proposed 10 gigawatt data center campus on federal land in Ohio as part of a deal that could include financial backing from Nvidia , according to two people with direct knowledge of the discussions. The campus under discussion would… 12 r/LocalLLaMA community 2mo ago Furiosa AI selling inference chip to consumer market will be a game changer to local llm ​ This is south Korean start up all-in on inference chip: https://furiosa.ai/renegade-spec Tsmc 5nm node Hynix HBM3 1.5TB/s 48GB VRAM TDP 180W Already tested on LG LLM. If they opened their programming interface the way NVIDIA opens PTX and Intel opens SPIR-V, and team up… 12 Hugging Face Daily Papers research 2mo ago Agents' Last Exam Abstract Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment. Generated by… 6 The Information — AI news-outlet 2mo ago Broadcom to Help Finance Anthropic, OpenAI Chip Deals With Apollo, Blackstone Broadcom said Tuesday that it is launching a new fund—backed by Apollo and Blackstone—to help finance more than 20 gigawatts of AI data centers through 2028 using chips designed by Broadcom, including projects tied to Anthropic and OpenAI. Apollo will lead an initial $35 billion… 19 Google DeepMind official-blog 2mo ago Powering the future of robotics in Europe Powering the future of robotics in Europe Jun 09, 2026 · Share x.com Facebook LinkedIn Mail Google DeepMind Accelerator selects 15 robotics companies from across Europe to join the program. Providing 3 months of intensive mentorship and technical support, enabling the… 22 r/LocalLLaMA community 2mo ago Apple announced new on device inference engine for Apple Silicon This news seem to have flown under the radar. Apple announced CoreAI on WWDC which is basically a future replacement for CoreML and an alternative to MLX/llama.cpp/torch for on-device optimized inference, especially on phones and tablets. The model weights need to be converted… 25 TechCrunch — AI news-outlet 2mo ago How an e-scooter founder raised $5 million to build space data centers Orbital founder Euwyn Poon built 250,000 scooters at Spin. Now he wants to launch 10,000 space data centers. 27 Hugging Face Daily Papers research 2mo ago Text-to-Image Models Need Less from Text Encoders Than You Think Abstract Text-to-image models primarily utilize basic text representation aspects like word merging and order rather than complex contextual information encoded in full text embeddings. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Text-to-image models rely on text prompts as… 36 arXiv — Machine Learning research 2mo ago The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers arXiv:2606.07587v1 Announce Type: new Abstract: LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by dynamically selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based… 29 arXiv — Machine Learning research 2mo ago EssentialGIN: a new approach for gene essentiality prediction based on graph isomorphism neural networks arXiv:2606.07700v1 Announce Type: new Abstract: Background: Prediction of essential genes (proteins), is a basic and challenging problem but at the same time very costly and time-consuming in wet-lab experiments. Predicting essential genes, only based on computational methods… 5 r/LocalLLaMA community 2mo ago New MLX LM Server From Apple Key Technical Advantages: Performance: The M5 chip's neural accelerators significantly boost prompt processing Concurrency: MLX LM Server utilizes continuous batching to handle multiple sub-agent requests simultaneously without stalling Scaling: For massive models that exceed… 15 NVIDIA Developer Blog official-blog 2mo ago Train Models Faster with JAX and MaxText Using NVFP4 on NVIDIA Blackwell Pre-training frontier LLMs comes down to throughput. When training spans trillions of tokens across thousands of accelerators, every percentage point of step... 34 Hacker News — AI on Front Page community 2mo ago Show HN: Gitdot – A better GitHub. Open-source, written in Rust What works now: user signups, org creations, private/public repos, and importing GitHub repositories (both as read-only mirrors and full migrations). So basically, you can create, push and pull to a repo, but we don't have many features quite yet (issues, PRs, CI). What is a bit… 34 Hacker News — AI on Front Page community 2mo ago A Farmer Donated Land to Turn into a Park. The City Is Building a Data Center Article URL: https://www.404media.co/a-farmer-donated-land-to-turn-into-a-park-the-city-is-building-a-massive-data-center-instead/ Comments URL: https://news.ycombinator.com/item?id=48446439 Points: 252 # Comments: 128 30 The Information — AI news-outlet 2mo ago Developers of OpenAI’s Stargate Data Center Face Higher Costs On a dusty stretch of prairie in Abilene, Texas, anxious hardware engineers from Crusoe, a data center developer for OpenAI and Oracle, have been working overtime to get natural gas turbines to work harmoniously with one of the most expensive AI supercomputers in history. It has… 37 arXiv — Machine Learning research 2mo ago Towards Serverless Semi-Decentralized Federated Learning with Heterogeneous Optimizers arXiv:2606.06687v1 Announce Type: new Abstract: We investigate cluster formation, involving the number and composition of clusters, in decentralized federated learning (FL) with heterogeneous machine learning (ML) optimizers. While clustering in centralized FL has enabled… 21 arXiv — Machine Learning research 2mo ago SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling arXiv:2606.06820v1 Announce Type: new Abstract: Agentic Large Language Model (LLM) systems decompose complex tasks into workflow Directed Acyclic Graphs (DAGs) whose primitives must be scheduled on heterogeneous clusters. Existing deep reinforcement learning (DRL) schedulers are… 26 arXiv — Machine Learning research 2mo ago Explaining Unsupervised Disease Staging in Huntington's Disease: Insights into Model Representations and Clusters arXiv:2606.07135v1 Announce Type: new Abstract: Huntington's disease (HD) is a progressive neurodegenerative disorder that affects motor, cognitive, and behavioral functions, where accurate characterization of disease progression remains essential to improve patient outcome and… 25 arXiv — Machine Learning research 2mo ago Unsupervised Continual Clustering via Forward-Backward Knowledge Distillation arXiv:2606.07474v1 Announce Type: new Abstract: Unsupervised Continual Learning (UCL) aims to enable neural networks to learn sequential tasks without labels or access to past data. A major challenge in this setting is Catastrophic Forgetting, where models forget previously… 23 arXiv — NLP / Computation & Language research 2mo ago Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations arXiv:2606.06740v1 Announce Type: cross Abstract: Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual interference in multilingual multi-speaker speech… 22 Page 8 of 10 · 500 articles ← Newer Older →