News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow r/MachineLearning community 2mo ago Live Continual Learning in Machine Learning [D] My question on live continual learning use cases was removed by moderators here because they think i asked basic level question about live continual learning which i thought is a frontier level research. But anyways. Is anyone interested in talking about continual learning… 30 TechCrunch — AI news-outlet 2mo ago OpenAI’s Jalapeño chip is Big Tech’s spiciest move away from Nvidia Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending.   OpenAI just shared its plans to spice things up with Jalapeño, its custom inference chip built with Broadcom, joining Google, Apple, and SpaceX in a growing list… 25 r/LocalLLaMA community 2mo ago Ornith 1.0 - terminology and concepts explained (basic) I made a quick guide for myself while wanting to try the new models, so I share it with you. It's pretty basic, but it may be useful for new people here. I also published the repo with the open code config and the commands: https://github.com/facuHannoch/AI_Workflows-Ornith-1.0… 34 arXiv — Machine Learning research 2mo ago SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference arXiv:2606.26587v1 Announce Type: new Abstract: Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM activation compression remains challenging: activations contain input-dependent outliers that… 29 arXiv — NLP / Computation & Language research 2mo ago Axon: A Synthesizing Superoptimizer for Tensor Programs arXiv:2606.26344v1 Announce Type: cross Abstract: Writing high performance kernels for AI accelerators requires deep expertise in tiling, instruction selection, data layout, and operator fusion placing a significant burden on programmers. In this paper, we focus on tile based AI… 33 r/LocalLLaMA community 2mo ago When you don't have a data center GPU Please don't tell me someone is going to (yet again) reply with the longest finetune-merge name in eternity...   submitted by   /u/Iwaku_Real [link]   [comments] 4 ThursdAI news-outlet 2mo ago GLM 5.2 total victory: the week open source won and nobody panicked From CoreWeave: A chill week, but a total Open Source victory for GLM 5.2 + Sakana Fugu, Krea Open Sources, OpenAI makes inference chips with broadcom, Karpathy gets heat about the new Claude Tag... 35 arXiv — Machine Learning research 2mo ago Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models arXiv:2606.24898v1 Announce Type: new Abstract: Looped language models turn hidden states into runtime state: each state is decoded for prediction and fed back into future computation. This creates a basic supervision question: which state variables does cross-entropy actually… 37 arXiv — NLP / Computation & Language research 2mo ago Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimization arXiv:2606.25656v1 Announce Type: new Abstract: As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for different RAG scenarios evaluation and comparison on semi-structured knowledge… 21 arXiv — NLP / Computation & Language research 2mo ago Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining arXiv:2606.26050v1 Announce Type: cross Abstract: Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ("Sue cried because"), it resolves the next pronoun to she, generalizing to held-out probes (0.94 by step… 4 r/LocalLLaMA community 2mo ago Locked Dell quote for 6x RTX PRO 6000 Max-Q at $8,960 — expires tonight. What would you do? Building an inference cluster to run GLM 5.2 locally. Got a Dell quote locked at $8,959.99/unit for 6x RTX PRO 6000 Blackwell Max-Q (300W). List price just jumped to $15,999 yesterday. Quote expires in ~3 hours and I can't swing all 6 right now. I have a second quote for 2 units… 36 r/LocalLLaMA community 2mo ago Any chance I could cluster my DGX Spark (128GB unified memory) and my AMD Ryzen AI Max 395 (128GM unified memory) together to run 1 model? Hey all, So I have a Nvidia DGX Spark and an AMD Strix 395, both have 128GB of unified memory. The Spark has 200Gbit network and the AMD Strix has 5Gbit ethernet (but it has a pcie gen 4x4 slot). Is there any chance I can cluster the 2 together to run a larger model that can fit… 30 r/LocalLLaMA community 2mo ago OpenAI and Broadcom unveil LLM-optimized inference chip https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ Quoted from the start of the blog post: Early testing shows that the first-generation accelerator will deliver performance per watt substantially better than current state-of-the-art Built from the ground up for… 11 Hacker News — AI on Front Page community 2mo ago 45°C cooling design cuts data center water use to near zero Article URL: https://blogs.nvidia.com/blog/liquid-cooling-ai-factories/ Comments URL: https://news.ycombinator.com/item?id=48660178 Points: 206 # Comments: 157 22 OpenAI official-blog 2mo ago OpenAI and Broadcom unveil LLM-optimized inference chip OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems. 28 arXiv — NLP / Computation & Language research 2mo ago ModTGCN: Modularity-aware Graph Neural Networks for Text Classification arXiv:2606.23694v1 Announce Type: new Abstract: Graph-based text classification models typically rely on local neighborhood aggregation and overlook global community structure, despite semantic document graphs exhibiting strong class-consistent clustering. Ignoring this can blur… 22 arXiv — NLP / Computation & Language research 2mo ago Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English arXiv:2606.23948v1 Announce Type: new Abstract: Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, and at what level of abstraction they encode it. One underexplored phenomenon is… 6 Ars Technica — AI news-outlet 2mo ago Oracle’s 21,000 layoffs help drive its debt-fueled AI investments Oracle is spending billions on data center infrastructure to support AI. 20 r/LocalLLaMA community 2mo ago Is it possible to run a giant model like GLM5.2 on this cluster (4x servers with 512GB RAM + dual AMD Epyc)? 16 channel memory should hit 409GB/s per node. Hey all, I have a piece of hardware laying around which is pretty fast from a traditional (non-GPU) server viewpoint. The hardware is the following: Dell C6525 Server with Quad Node (4x server blades) with the following: 2x AMD EPYC 7702 64-Core Processors 8 memory channels per… 30 TechCrunch — AI news-outlet 2mo ago Nvidia wants to cut data center water use, but that’s not the same as fixing AI’s water problem Nvidia announced a new cooling system that cuts water use inside the data center. But it does nothing to address AI's biggest water use — fossil fuel power plants. 5 TechCrunch — AI news-outlet 2mo ago SpaceX inks compute deal with Reflection AI, an open-source AI lab Reflection AI will pay $150 million a month beginning July 1, 2026 through 2029 for immediate access to Nvidia's latest GB300 AI chips and supporting hardware across SpaceX's Colossus 2 data center near Memphis, Tennessee. 33 r/MachineLearning community 2mo ago Data-centric debugging for teams training neural nets [P] We just did a big revamp of WeightsLab and wanted to share it here. If you’ve ever spent hours debugging a training run only to discover it was a data problem all along, this is for you. WeightsLab lets you pause training mid-run, inspect your live loss signals, and catch… 29 r/LocalLLaMA community 2mo ago What‘s your local „Haiku“-Replacement? Seriously looking for a reliable and fast local Haiku replacement. Basically it should be able to summarize technical stuff, code documentation, architectural descriptions Any suggestions? Edit: sorry, totally forgot that my local machine is a M4 Max 128GB. But at the same time… 6 r/LocalLLaMA community 2mo ago Deep Neural Network that can turn any Image into a Playable Game! BUT LOCALLY, NOT ON DATACENTER Hi everyone!! I really wanted to share my research what I've been working on. I wanted to build a nn that can simulate games, or at least start doing that Most video generators are too large to run on consumer hardware realtime, so I I designed a model that does this from… 14 r/LocalLLaMA community 2mo ago New Agentic Benchmark Out: Claude Fable and GLM 5.2 Top Their Cohorts You can read about it here: https://artificialanalysis.ai/articles/aa-briefcase This is a solid benchmark from Artificial Analysis. It basically tests an LLMs ability to plan and execute tasks. And more importantly, it is a new benchmark that is not saturated, so no one can… 32 r/LocalLLaMA community 2mo ago EvoTensile: Evolutionary algorithms for AMD Tensile GEMM kernel tuning There has been an effort to tune kernels in hipBLASLt so the most basic matmuls can run faster. It's known that on Strix Halo (gfx1151), GEMM with NN and TN input layouts (used in inference) are already well-tuned, while NT and TT layouts (used in training) are not yet tuned.… 8 arXiv — Machine Learning research 2mo ago Exploring the potential of AlphaEarth and TESSERA embeddings for Fine-scale Local Climate Zone Mapping: A case study across five cities in Switzerland arXiv:2606.20034v1 Announce Type: new Abstract: Understanding urban spatial morphology is critical for climate modeling, risk assessment, and sustainable urban design, and Local Climate Zone (LCZ) mapping provides the basic framework for this. However, many cities still use… 10 arXiv — NLP / Computation & Language research 2mo ago Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability arXiv:2606.19815v1 Announce Type: new Abstract: Pre-trained language models such as BERT achieve strong text classification performance but lack transparency, limiting their use in high-stakes settings. The Tsetlin Machine (TM) offers fully interpretable, clause-based reasoning… 25 arXiv — NLP / Computation & Language research 2mo ago TransLaw: A Large-Scale Dataset and Multi-Agent Benchmark Simulating Professional Translation of Hong Kong Case Law arXiv:2507.00875v3 Announce Type: replace Abstract: Translating Hong Kong Court Judgments from English to Traditional Chinese is mandated by Articles 8-9 of the Basic Law, yet remains constrained by a shortage of parallel resources and rigorous demands on legal terminology,… 38 arXiv — NLP / Computation & Language research 2mo ago ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents arXiv:2508.04266v4 Announce Type: replace Abstract: Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and… 22 Hacker News — AI on Front Page community 2mo ago Show HN: Are You in the Weights? With more traffic moving off-web and into LLMs, I got curious about what traces we leave "in the weights". My design partner and I built a site in the past few weeks that checks recognition across frontier and small models. It queries many of them in parallel, clusters the… 37 TechCrunch — AI news-outlet 2mo ago Amazon hopes to challenge Nvidia more directly by selling its AI chips AWS is in talks to sell its chips to other data centers. CEO Andy Jassy has said this represents a $50 billion opportunity for the company. 37 TechCrunch — AI news-outlet 2mo ago AI data centers just got a government-mandated fast lane to the grid FERC told grid operators to give data centers a fast lane for interconnections, but it failed to address electricity supply shortages. 30 arXiv — Machine Learning research 2mo ago scGTN: Deep Siamese Graph Transformer Network for Single-cell RNA Sequencing Clustering arXiv:2606.18672v1 Announce Type: new Abstract: Single-cell RNA sequencing (scRNA-seq) serves a pivotal role in characterizing gene expression at the cellular level, enabling the identification of cell types and advancing the understanding of cellular heterogeneity. Despite the… 25 arXiv — Machine Learning research 2mo ago Online Distributional Prediction via Latent Cluster Geometry Under Drift and Corruption arXiv:2606.18778v1 Announce Type: new Abstract: Online learning in non-stationary streams is often formulated as tracking a point estimate, but many applications require predicting the full data-generating distribution. We study online distributional prediction under drift and… 7 arXiv — Machine Learning research 2mo ago Seed-Guided Semi-Supervised Clustering by A-Contrario Anomaly Detection arXiv:2606.18833v1 Announce Type: new Abstract: This paper introduces a semi-supervised clustering framework grounded in the statistical duality between grouping principles and anomaly detection. We address the challenge of robust cluster definition in noisy environments -- a… 38 arXiv — Machine Learning research 2mo ago FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs arXiv:2606.19025v1 Announce Type: new Abstract: Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators. While increasing model and dataset scale remains the dominant driver of performance,… 9 r/LocalLLaMA community 2mo ago GLM 5.2 Release Video [Made with GLM 5.2] Everyone's probably seen the remotion thing that went viral a couple months back with CC. Its basically that with GLM 5.2 as the model provider. Close to Fable but still a step below on creativity, top is still Gemini 3.1 pro for vid creation but at least I can see why Design… 21 r/MachineLearning community 2mo ago Contrastive targeted SFT as a mechinterp method - has anyone mapped causal dependency interactions this way? [D] Hi All, I've been running experiments on targeted SFT for specific capability dimensions on a 31B model. After running small training run to prime the model slightly in the direction I want, then ran a judge across 40 domains scoring six independent quality dimensions. One… 21 r/LocalLLaMA community 2mo ago GLM-5.2 is a win for local AI I know GLM 5.2's massive 753B footprint means none of us are running it at home without an enterprise cluster, but having a true frontier-level, MIT-licensed coding agent out in the wild makes me optimistic. The distillation potential here is massive. Once the community starts… 38 TechCrunch — AI news-outlet 2mo ago Canadian pension giant joins race to fund India’s AI-fueled data center boom The Canadian pension giant will acquire an 8.2% stake in CtrlS, a tech giant that operates more than 15 data centers across India. 8 r/LocalLLaMA community 2mo ago Local models went from mostly useless to actually useful really fast. What changed? https://preview.redd.it/knc4ht7bft7h1.png?width=1048&format=png&auto=webp&s=49abdb8b0f358e799ecb06aa49134d9b0fd49336 Mitchell Hashimoto had a good point earlier: local models went from basically useless to actually useful in what feels like one year. I think thats pretty… 5 arXiv — Machine Learning research 2mo ago C2FL: Clustered Continual Federated Learning under Spatial and Temporal Drift arXiv:2606.18003v1 Announce Type: new Abstract: Collective Adaptive Systems (CAS) increasingly rely on machine learning to let each node learn from locally sensed data, aligning its behavior with the surrounding environment. Scaling this intelligence, however, raises fundamental… 8 Ars Technica — AI news-outlet 2mo ago Trump admin tries to block Clean Air Act lawsuit over xAI's gas turbines NAACP lawsuit says xAI uses gas turbines without permits for Grok data center. 19 NVIDIA Developer Blog official-blog 2mo ago How to Optimize Transformer-Based Models for Low-Precision Training Transformer architectures are the backbone of many modern large language and generative AI models. As these models grow in size, training runs consume more GPU... 5 MIT Technology Review — AI news-outlet 2mo ago Want to get a data center online quickly? Give it some flex. At the end of a tense and scoreless first half of a soccer match between the English men’s team and rival Germany, millions of Brits let out a collective sigh and did what they so often do in moments of stress: They made tea. That wave of electric kettles clicking on, however,… 26 arXiv — Machine Learning research 2mo ago Distilling Drifting Transformers with Representation Autoencoders arXiv:2606.15553v1 Announce Type: new Abstract: Representation Autoencoders (RAEs) have improved diffusion and flow models by semantically richer latent space owing to the strongly label-wise clustered DINO features in the pretrained encoders. Yet in the distillation stage, the… 15 r/LocalLLaMA community 2mo ago "My son is a genius coder" - honest Alpha Tester review "It's not slop - it's an art" - Grandma. Introducing you my few weeks brainstorming and writing of the code. I was rewrite everything my AI was creating so basically it's my own creation. Brain Calculator Pro™ — the calculator that made your calculator obsolete. AI-powered… 11 r/MachineLearning community 2mo ago Could AI training be decentralized like Bitcoin mining? [D] I’ve been thinking about whether the same basic concept behind Bitcoin could be applied to AI training. In Bitcoin, miners perform proof-of-work and are rewarded for contributing computational resources to secure the network. The actual computation itself isn’t particularly… 15 The Information — AI news-outlet 2mo ago Nvidia’s Share of AI Inference Chip Market Appears to Be Rising As AI developers and cloud providers have launched server chips to lessen their dependence on Nvidia’s, some analysts and executives at these firms expected the chips to eat into Nvidia’s market share. That doesn’t seem to be happening. Nvidia has actually increased its share of… 4 Page 7 of 10 · 500 articles ← Newer Older →