News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow arXiv — Machine Learning research 3mo ago M\=oLe-{\Lambda}: Learning the Coupled-Cluster Response State for Energies, Gradients, and Properties arXiv:2605.29622v1 Announce Type: new Abstract: Coupled-cluster (CC) theory is often considered the gold standard of quantum chemistry, but its high computational cost limits routine access to accurate energies, forces and response properties. While the right-hand $T$-amplitudes… 5 r/LocalLLaMA community 3mo ago llama.cpp B9387 Significant AMD/ROCm PP Update https://github.com/ggml-org/llama.cpp/releases/tag/b9387 MFMA is restricted to AMD CDNA architecture that's MI100, MI200, MI300 series datacenter cards. Post your initial results if you try it! wink   submitted by   /u/Bulky-Priority6824 [link]   [comments] 38 r/LocalLLaMA community 3mo ago Mimo 2.5 Pro - 40t/s on 8x Nvidia Spark/GB10 cluster I got Mimo 2.5 Pro running on my 8x Asus Nvidia GB10 cluster using mtp-2, single user request, coding: 40 t/s - 1k context, 32t/s - 30k context, 25t/s - 125k context, 17t/s - 250k context. 2 parallel reached 60t/s and in 4 parallel reached 83t/s, not bad for 1T model… 10 The Information — AI news-outlet 3mo ago SpaceX Committed to Six Month Anthropic Data Center Lease, Musk Says SpaceX is committed to renting its data center capacity to Anthropic for 180 days, though it could extend the deal for longer, CEO Elon Musk said Thursday. The comment provides more detail than language in SpaceX’s recent public S-1 filing. That document, which SpaceX published… 29 r/LocalLLaMA community 3mo ago Zai replaced the network architecture running GLM-5.1 inference and the gains are pretty wild Been following the infrastructure side of AI more lately and stumbled on this from Zai. They upgraded the network architecture on a thousand-GPU cluster running GLM-5.1 coding inference from the standard ROFT setup to something they built called ZCube, developed with Tsinghua… 27 r/LocalLLaMA community 3mo ago Heterogeneous GPU Weighting & Layer Splitting This is what I worked on today. With local LLM of course. So if I didn't write the code, did I really work on it? Who cares. It was my idea and I simply asked it to implement it. I basically downloaded /main/ branch, which is totally broken for Windows by the way (i had to… 21 arXiv — Machine Learning research 3mo ago Robust Contrastive Graph Clustering with Adaptive Local-Global Integration arXiv:2605.28209v1 Announce Type: new Abstract: Graph clustering is essential in graph analysis for revealing structural patterns and node communities. Despite recent advances in self-supervised contrastive learning that have improved clustering via structural and attribute… 23 The Information — AI news-outlet 3mo ago ByteDance Mulls $70 Billion Capex This Year as AI Costs Grow ByteDance is considering more than doubling its capital expenditures this year to as much as $70 billion, as the Chinese tech giant ramps up its investment in data centers and other AI infrastructure, Bloomberg reported. ByteDance’s capex plans reflect the surging costs of… 25 Stratechery (Ben Thompson) community 3mo ago The SpaceX IPO and Data Centers in Space There isn't a financial model that justifies the SpaceX IPO, but data centers in space are plausible, and that might be enough. 24 Hugging Face Daily Papers research 3mo ago DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs Abstract DarkForest is a controlled-communication framework that enhances multi-agent LLM reasoning by clustering semantic candidates and using calibrated belief distributions to reduce error propagation and communication overhead. AI-generated summary Multi-agent LLM systems… 11 The Information — AI news-outlet 3mo ago Qualcomm Strikes AI Chip Deal With ByteDance Qualcomm has reached a deal with ByteDance to supply chips for AI data centers to the Chinese tech giant, Bloomberg reported. The deal comes as Qualcomm, one of the world’s largest suppliers of smartphone processors, is trying to increase its presence in chips for AI computing.… 15 r/MachineLearning community 3mo ago [R]GNN Model For Fraud Detection Isn't Performing Well[R] We're writing a research paper on explainable fraud detection GNN model and in the first step we're creating a basic Graph Neural Network for that. We're using the most famous dataset available on this topic i.e IEEE CIS Fraud Detection Dataset and implemented all necessary… 7 arXiv — Machine Learning research 3mo ago GEM: Geometric Entropy Mixing for Optimal LLM Data Curation arXiv:2605.26121v1 Announce Type: new Abstract: LLM pre-training efficacy increasingly depends on data composition rather than sheer volume. Yet, optimal mixing is hindered by categorization flaws: human taxonomies suffer from ontological misalignment, and Euclidean clustering… 27 r/LocalLLaMA community 3mo ago Small set of local MCP server installers for home Linux users Hi all, I have published a small open-source MCP server bundle called MCP Basic Servers : https://github.com/mchowy-troll/mcp-basic-servers It is a collection of simple Bash installer scripts for running local MCP HTTP servers on Linux . The idea is simple: run one script,… 38 arXiv — Machine Learning research 3mo ago Interdomain Attention: Beyond Token-Level Key-Value Memory arXiv:2605.24330v1 Announce Type: new Abstract: Transformers and deep state space models (SSMs) sit at opposite ends of a basic design choice: attention routes each query through a growing key-value (KV) cache by content-based matching at quadratic cost, while deep SSMs compress… 19 arXiv — Machine Learning research 3mo ago A computational phase transition for learning-to-sample from Ising models arXiv:2605.24752v1 Announce Type: new Abstract: We study \emph{learning-to-sample} -- a basic algorithmic task underlying generative modeling -- for Ising models, a standard testbed for algorithmic ideas in both theoretical computer science and machine learning. Given i.i.d.… 10 arXiv — NLP / Computation & Language research 3mo ago How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly grounded thematic organization. We evaluate whether bibliometric structure improves… 14 arXiv — NLP / Computation & Language research 3mo ago Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation arXiv:2605.24534v1 Announce Type: new Abstract: We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any handcrafted doctrinal framework. Using 4.555 decisions of the German Federal… 26 r/LocalLLaMA community 3mo ago Update on 12x32gb sxm v100 cluster / local AI for legal drafting Update from the lawyer with the V100 server. A few of you asked what I actually ended up running once the dust settled, so here it is. Still just a lawyer, still driving the whole thing through Claude Code, still not fully sure what I'm doing — but it works now, which is more… 15 r/LocalLLaMA community 3mo ago Anyone use QwQ-32B? It's over a year old? Has Qwen 3.6 27b basically replaced it? I seen this one mentioned but it was a source from about 14 months ago. In the age of the Qwen 3.6 and Gemma 4- is there still a use for QwQ 32B? Does anyone still favour it over the new stuff? If so, do you use it for coding? something else? Thanks   submitted by  … 29 r/LocalLLaMA community 3mo ago Embeddings for NVIDIA's Nemotron Personas I extracted embedding vectors for nvidia/Nemotron-Personas dataset. It's an incredible resource consisting of millions of synthetic personas with detailed backgrounds (names, ages, occupations, hobbies, and more), but finding specific personas or clustering them is difficult. To… 5 TechCrunch — AI news-outlet 3mo ago Elon Musk has given up on solar power (on Earth) Elon Muks's xAI has gone all in on natural gas, while SpaceX is obsessed with orbital data centers. What happened to the "solar-electric economy" he promised? 6 r/LocalLLaMA community 3mo ago LLaMa.cpp basic question I'm trying to install LLaMa with PI agent. I ran curl -fsSL https://pi.dev/install.sh | sh export PATH="/home/user/.local/share/pi-node/node-v22.22.3-linux-x64/bin:$PATH pi install npm:pi-llama.cpp These commands installed pi, added them to path and then I lastly installed an… 34 r/MachineLearning community 3mo ago Anthropic posted a profit while xAI burned $4.2B. The AI profitability numbers finally leaked.[D] This week basically forced everyone to stop guessing about AI margins. Three major financial reality checks hit at once: OpenAI confidentially filing their S-1, xAI’s Q1 numbers leaking via SpaceX, and Anthropic somehow posting an actual operating profit. If you are building an… 4 Stratechery (Ben Thompson) community 3mo ago 2026.21: The Data Center Veto The best Stratechery content from the week of May 18, 2026, including data center discontent, agent economics, and slime mold. 26 Dwarkesh Podcast news-outlet 3mo ago Reiner Pope – Chip design from the bottom up Working up from basic logic gates to why GPUs, TPUs, FPGAs, and the human brain each look the way they do. 22 arXiv — Machine Learning research 3mo ago TONIC: Token-Centric Semantic Communication for Task-Oriented Wireless Systems arXiv:2605.21553v1 Announce Type: new Abstract: Tokens are becoming the basic units through which foundation models represent and process information for understanding and inference. However, traditional wireless communication, centered on bit-level fidelity, faces a mismatch… 33 r/LocalLLaMA community 3mo ago When your LLM treats data center GPUs like an optional DLC   submitted by   /u/noprompt [link]   [comments] 10 Hugging Face Daily Papers research 3mo ago Capturing LLM Capabilities via Evidence-Calibrated Query Clustering Abstract Query clustering algorithm ECC improves LLM capability evaluation by aligning semantic embeddings with latent capability demands through posterior model comparisons and Bradley-Terry modeling. AI-generated summary Query clustering organizes queries into groups that… 13 Ars Technica — AI news-outlet 3mo ago As Grok flounders, SpaceX bets future on beating Big Tech at AI SpaceX IPO filing pitches orbital data centers as Grok lags rival AI services. 26 r/LocalLLaMA community 3mo ago Qwen3.6 35Ba3 has changed my workflows and even how I use my computer My workflow has changed basically to ask Codex to do certain tasks and then document how to do them (including errors it found on its way) into a skill. I feed that skill to pi, and suddenly my qwen3.6 gets that hard stuff done: - devops on a VPS - using docling to create epubs… 33 Google DeepMind official-blog 3mo ago We’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risks The Asia-Pacific region is a global engine for economic growth, but it's also highly vulnerable to climate change. While green technologies are gaining momentum, a recent report shows they aren’t scaling fast enough to keep up with the region’s rising environmental risks. To… 22 r/MachineLearning community 3mo ago Can liveness detection models generalise to synthetic media generation techniques they were never trained on? [D] Most liveness detection systems in production today were built around a threat model where the attacker is submitting a static image or a basic replay video. The generation quality of current synthetic media is categorically different from what those training datasets captured.… 32 NVIDIA Developer Blog official-blog 3mo ago Get Real-Time Visibility into GPU Usage Across Kubernetes Clusters Maximizing the value of AI infrastructure demands deep visibility into GPU utilization. Yet many platform teams running AI workloads on Kubernetes operate with... 25 r/MachineLearning community 3mo ago I created an LLM post-training method called RPS. Preliminary results show that it improved Qwen3-8b's program synthesis reliability. [R] RPS is inspired by neuroscience. As humans, we learn basic skills as kids with high neuro-plasticity. We then learn advanced skills as teens and adults with low neuro-plasticity. RPS trains a model in 2 stages. In stage 1, the model is trained on easy data with high learning… 26 Hugging Face Daily Papers research 3mo ago CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing Abstract Current GUI agents show limited effectiveness in professional media post-production tasks despite advances in spatial grounding and multimodal alignment. AI-generated summary While GUI agents have made significant progress in web navigation and basic operating system… 13 arXiv — Machine Learning research 3mo ago Unsupervised clustering and classification of upper limb EMG signals during functional movements: a data-driven arXiv:2605.20599v1 Announce Type: new Abstract: This study presents a comprehensive approach for the clustering and classification of upper-limb surface electromyography (sEMG) signals during functional reach and grasp movements. The methodology was applied to the NINAPRO DB4… 18 arXiv — NLP / Computation & Language research 3mo ago Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy arXiv:2605.21391v1 Announce Type: new Abstract: Metaphor requires a language model to resolve a token whose contextual meaning diverges from its basic literal sense. Understanding how transformer models organize this reinterpretation across depth remains an open problem in… 19 The Information — AI news-outlet 3mo ago Anthropic and SpaceX Detail Compute Deal Worth Up to $40 Billion Anthropic could pay SpaceX up to $40 billion over the next several years to use compute from data centers, but either company has the power to call off the deal early, SpaceX revealed when it filed for an initial public offering on Wednesday. SpaceX is receiving $1.25 billion… 16 Latent.Space news-outlet 3mo ago Railway: The Agent-Native Cloud — Jake Cooper 3M Users, 100K Signups/Week, Own-Metal Data Centers, $200K+ Coding Agent Spend, and the Death of PRs 21 TechCrunch — AI news-outlet 3mo ago Musk’s xAI is being sued over its data center generators. Now, it’s buying $2.8B more. Elon Muks's xAI said it will buy $2.8 billion worth of natural gas turbines over the next three years, according to SpaceX's IPO filing. 6 r/LocalLLaMA community 3mo ago 24GB M4 Mac - is Qwen 9B only option while system is running? I have mac at work that I want to use local model for prototyping and basic prompts that needs to stay on device. What sort of model I can run that I can fit at least 64k context ? Any setups share or guides welcome. I need to have firefox open with one tab at minium. Problem I… 6 The Information — AI news-outlet 3mo ago Sam Altman Offers YC Founders $2 Million in OpenAI Tokens For Equity OpenAI cofounder and CEO Sam Altman late Tuesday offered to invest $2 million in every startup currently in the Y Combinator startup accelerator program—not in cash, but in OpenAI tokens. “I am excited to see what will happen with tokenmaxxing startups, both for how they work… 13 arXiv — Machine Learning research 3mo ago DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training arXiv:2605.18815v1 Announce Type: new Abstract: Modern large language model (LLM) training is inherently dynamic: resource fluctuations, RLHF phase shifts, and cluster elasticity continually reshape the optimal parallelism layout, posing a significant challenge to existing… 22 arXiv — Machine Learning research 3mo ago A Multi-Dimensional Clustering Approach for Identifying Inborn Errors of Immunity arXiv:2605.18880v1 Announce Type: new Abstract: Rare diseases such as inborn errors of immunity (IEI) require early diagnosis to prevent end organ damage and improve quality of life. Hurdles in accessing and curating large scale electronic health record (EHR) data limit routine… 10 arXiv — NLP / Computation & Language research 3mo ago Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering arXiv:2605.19220v1 Announce Type: new Abstract: Uncertainty Quantification (UQ) is widely regarded as the primary safeguard for deploying Large Language Models (LLMs) in high-stakes domains. However, we argue that the field suffers from a category error: mainstream UQ methods… 22 arXiv — NLP / Computation & Language research 3mo ago ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation arXiv:2605.18769v1 Announce Type: cross Abstract: Personalized Retrieval-Augmented Generation (RAG) relies on accurately selecting user-relevant documents. In practice, existing RAG approaches often suffer from high retrieval costs and overlook that collaborative signals from… 36 r/LocalLLaMA community 3mo ago Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s! Hey r/DeepSeek , Who says we need an H100 cluster or the latest expensive GPUs to run frontier MoE models? I wanted to see how far we could push a single node of consumer legacy hardware, so we spent less than $2,500 total to build a budget machine that successfully runs… 29 r/LocalLLaMA community 3mo ago Intel's Crescent Island PCB Leaks, Showing a Massive Xe3P GPU, 16-Pin Connector, 160GB LPDDR5X as Intel Sidesteps the HBM Shortage Upcoming Intel Xe3P data center GPU with 20 8GBLPDDR5X modules for a total of 160GB, bypassing HBM shortages. Assuming a 32-bit interface, that's a 640-bit wide memory interface, or 10 channel memory interface if converted to the 64-bit wide desktop equivalent. At 8800-9500MT,… 35 Ars Technica — AI news-outlet 3mo ago Electrical utility megamerger is all about the data centers NextEra’s blockbuster deal with Dominion likely means higher bills for consumers. 29 Page 10 of 10 · 500 articles ← Newer