News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow arXiv — Machine Learning research 1mo ago Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata arXiv:2607.28338v1 Announce Type: new Abstract: Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation,… 15 TechCrunch — AI news-outlet 1mo ago Investors love AI, as long as you’re a cloud host Amazon isn't slowing down on data center spending — but investors don't seem to mind. 29 NVIDIA Developer Blog official-blog 1mo ago NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We... 9 TechCrunch — AI news-outlet 1mo ago Nscale buys Anyscale as it seeks to own more of the AI compute stack British AI neocloud Nscale is buying software startup Anyscale, which helps companies scale their AI workloads across data centers and servers. 12 arXiv — Machine Learning research 1mo ago FloDR: An invertible dimensionality reduction method based on a normalising flow arXiv:2607.26278v1 Announce Type: new Abstract: It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point… 22 arXiv — Machine Learning research 1mo ago PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems arXiv:2607.26710v1 Announce Type: new Abstract: The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. Under stringent grid… 6 arXiv — Machine Learning research 1mo ago Archetypes or ability? Clustering for modelling student mathematical competence arXiv:2607.26063v1 Announce Type: cross Abstract: Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different… 36 arXiv — Machine Learning research 1mo ago Randomizing the Number of Centers in k-means++ arXiv:2607.26202v1 Announce Type: cross Abstract: The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$. We consider the same… 17 r/LocalLLaMA community 1mo ago Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar? I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and was pretty damn happy with the results at first. But, Q8 quantization 130k context was barely squeezing through. Naturally, I bought two RTX 6000… 37 r/LocalLLaMA community 1mo ago "Uncensored" LLMs are measurably more optimistic than their base models Hi. Many people think uncensored models are basically the same model that just doesn't refuse, but... I was recently checking whether uncensored models would give me better answers for stock market predictions (my idea was: the uncensored one will tell you the truth and won't be… 38 arXiv — NLP / Computation & Language research 1mo ago An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data arXiv:2503.07303v3 Announce Type: replace Abstract: Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulaic texts, which are characterized by repetition and constrained expression, tend… 37 r/MachineLearning community 1mo ago Would you take an MLE role with 50% base salary increase with rigid pto policy and no 401K match? [D] I got an offer for an MLE role where basically i will develop ML/GENAI solutions for clients completely and then hand them over and move to other projects. I think its a consulting ML role, although from a salary perspective its amazing and base is 50% more and if i add my new… 23 TechCrunch — AI news-outlet 1mo ago Data centers may face temporary power cuts to prevent blackouts on largest US grid The largest grid operator in the U.S. says it will cut power to large data centers to prevent blackouts starting next year. 7 Hugging Face Daily Papers research 1mo ago Characterizing Warp Divergence from Pascal to Blackwell Abstract Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp divergence in a fixed manner. We test this assumption across Ampere, Hopper, and datacenter and consumer Blackwell GPUs, using pre-ITS Pascal as a baseline.… 23 arXiv — Machine Learning research 1mo ago What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation arXiv:2607.22781v1 Announce Type: new Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph… 34 r/LocalLLaMA community 1mo ago I ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs) Someone in the comments of my 27B post-train bakeoff asked for the 35B version, so I ran it. Same setup as last time: fresh Coder workspaces on my k8s cluster, each driving my own agent (Hermes) headlessly, models on llama.cpp via llama-swap on one 5090, every call traced… 21 Ars Technica — AI news-outlet 1mo ago Verizon seeks AI profits with mini data centers, $1B dark fiber deal with Google Telecom expects AI revenue from dark fiber deals and retrofitted data centers. 23 r/LocalLLaMA community 1mo ago Viable ways to run K3 locally just curious how would people run it cheap if they really want kimi k3. dgx spark / strix halo clusters optane persistent memory platform + some gpus mac studio clusters orange pi 6 clusters ssd streaming + gpus multiple ddr3 + connectx 5 rdma clients two dgx stations power 10… 37 r/LocalLLaMA community 1mo ago You can now fine-tune my 3.96M-parameter TTS on your own voice or language When I released Inflect v2 last week, I thought most people would ask whether a TTS model this small actually sounded decent. Instead, I kept getting two questions: “Can I train it on my own voice?” “Can I move it to another language?” At the time, my answer was basically:… 16 r/LocalLLaMA community 1mo ago Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week. Weights are supposed to hit Hugging Face today… 18 r/LocalLLaMA community 1mo ago I want to run Kimi K3 at home, so I’m trying to make 2.8T-scale experimentation cheaper Hey r/LocalLLaMA , I’m a retired engineer with a background in distributed computing, currently running a 1-person startup. Like many people here, I’d love to experiment with 2T+ MoE models locally. The problem is that I don’t have an H100 cluster in my living room. So I’ve been… 12 arXiv — Machine Learning research 1mo ago RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clusters. The Reduced Interaction Sampling (RIS) inference engine addresses this… 31 arXiv — NLP / Computation & Language research 1mo ago Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA arXiv:2607.21861v1 Announce Type: new Abstract: We study baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so a system can answer questions about a corpus closed-book: no retrieval and no context-window budget. Across roughly 100 training runs… 22 r/LocalLLaMA community 1mo ago Will prices finally go down? I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments into AI data centers and had no use for then, had to rent them, same thing with… 38 r/LocalLLaMA community 1mo ago Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Claude code (with DS in CLIProxyAPI ) takes nearly 4 times longer than the fastest… 38 r/LocalLLaMA community 1mo ago 90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the story: each run was a fresh Coder workspace on my k8s cluster driving my own… 17 r/LocalLLaMA community 1mo ago World's First(?) Underwhelming AMD Ryzen AI Halo Cluster LTT Labs recently received the Linux version of the AMD Ryzen AI Halo for testing, but it turns out that AMD had intended to send the Windows version. Through this stroke of misfortunate, we were fortunate enough to have two Ryzen AI Halos for a short period of time and the… 14 TechCrunch — AI news-outlet 1mo ago One fallen power line exposed a growing AI data center problem. Here’s how to fix it. A close call in Northern Virginia revealed just how poorly data centers respond to grid disruptions. Here's how to fix the problem. 34 Ars Technica — AI news-outlet 1mo ago AI firms want more data centers; Trump's EPA may give neighbors less say Rule would allow states to decide how much—if any—public input there can be. 12 Hacker News — AI on Front Page community 1mo ago IRGC claims it destroyed Amazon's Bahrain data center Article URL: https://houseofsaud.com/irgc-claims-destroyed-amazon-bahrain-data-center/ Comments URL: https://news.ycombinator.com/item?id=49033240 Points: 262 # Comments: 332 8 arXiv — Machine Learning research 1mo ago External Clustering Validation by the Homogeneity-Parsimony Trade-off arXiv:2607.20799v1 Announce Type: new Abstract: Scalar metrics are often used to evaluate clusterings against known classes, but they can obscure a fundamental trade-off: clusterings should be informative about class labels while avoiding unnecessary fragmentation. Here we… 34 arXiv — Machine Learning research 1mo ago Regularized Optimization on Grassmann Manifold: Theory, Algorithm and Applications arXiv:2607.21039v1 Announce Type: new Abstract: Spectral methods are among the most widely used techniques for community detection, clustering, and graph learning. Their performance, however, critically depends on the accurate estimation of the underlying spectral subspace and… 7 arXiv — Machine Learning research 1mo ago CASC: Causal Adversarial Subspace Clustering for Multivariate Spatiotemporal Data arXiv:2607.21088v1 Announce Type: new Abstract: Deep subspace clustering plays a critical role in applications involving multivariate spatiotemporal data, such as sea ice monitoring, disease spread analysis, and tracking neuro-degeneration over time. Despite recent advances,… 29 arXiv — Machine Learning research 1mo ago Semantic-Aware Task Clustering for Constructive and Cooperative Multi-Tasking arXiv:2607.21426v1 Announce Type: new Abstract: Cooperative multi-task semantic communication (CMT-SemCom) improves task execution performance by leveraging shared representations. However, as we demonstrated in [1], cooperative multi-tasking can be either constructive or… 16 r/MachineLearning community 1mo ago NeurIPS E and D, Average rating 3 and average confidence 4, I can rebuttal and address all their concerns? Do I still have a decent shot or unlikely ?[R] NeurIPS E and D track review are out today and the average rating I received is a 3 and confidence is a 4. I can correct and address all their concerns. Do I still have a genuine shot of getting in or is it basically impossible at this point since none of my scores are a 4 or 5?… 37 arXiv — Machine Learning research 1mo ago SCPP: A Unified Python Library for Soft Clustering arXiv:2607.19620v1 Announce Type: new Abstract: In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training,… 8 arXiv — Machine Learning research 1mo ago Efficient Clustering with Provable Guardrails for LLM Inference at Scale arXiv:2607.19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models. A natural fix is to cluster the inputs and call the LLM only on cluster representatives, letting… 31 arXiv — NLP / Computation & Language research 1mo ago BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators arXiv:2607.19438v1 Announce Type: cross Abstract: Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. We show that BaseRT, our native Metal… 8 r/LocalLLaMA community 1mo ago 🇦🇹 Austria is rolling out a government AI-platform using Mistral models and Open WebUI This is a surprisingly large real-world deployment: "GovGPT" is part of Austria’s Public AI initiative, running on sovereign infrastructure (in their BRZ - federal datacenter) with Mistral open-weight models. Trending Topics reports that Open WebUI is used as the interface for… 5 arXiv — Machine Learning research 1mo ago An unsupervised clustering analysis of breast cancer data derived from electronic health records enhanced through UMAP dimensionality reduction arXiv:2607.19089v1 Announce Type: new Abstract: Breast cancer is one of the most widespread types of cancer, affecting approximately 8 million women worldwide. Electronic health records of patients diagnosed with this disease can serve as valuable datasets for computational… 5 arXiv — NLP / Computation & Language research 1mo ago Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression arXiv:2510.02345v4 Announce Type: replace Abstract: Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead. We introduce a unified framework based on dynamic expert clustering and structured… 21 TechCrunch — AI news-outlet 1mo ago Data centers expected to use 4x more electricity by 2035 New data centers built through 2033 could consume as much electricity as India uses today. 27 r/LocalLLaMA community 1mo ago 20B Looping model (paper) matches or beats Qwen3 Coder 30B at 10% of pre-training tokens No weights yet. I feel sad for them, that training run cost maybe 100s of thousands of dollars and they didn't even beat GPT-OSS 20B in every regard But the ability to train a model from scratch on 3.5 trillion tokens instead of 35 trillion sure gives me hope. They only spent… 33 MIT Technology Review — AI news-outlet 1mo ago Advancing next-gen AI with materials science innovation The conversation about AI often centers on algorithms, computing power, or huge investments in new semiconductor fabrication plants and hyperscale data centers. But beneath each of these advances is another layer of innovation that makes them possible: advanced materials. Every… 21 arXiv — Machine Learning research 1mo ago Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies arXiv:2607.17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks. However, their subpar performance on seemingly elementary problems, such as basic… 33 arXiv — NLP / Computation & Language research 1mo ago D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation arXiv:2607.17538v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. However, its dense vector retrieval introduces… 4 Hugging Face Daily Papers research 1mo ago JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Abstract The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an… 13 TechCrunch — AI news-outlet 1mo ago AI’s most important protocol is getting a little bit easier to use The Model Context Protocol (MCP) is one of the basic building blocks of AI interoperability, giving AI models a secure way to access external data sources and services. It’s the plumbing that lets a chatbot reach into your calendar, your database, or your internal tools,… 20 Hugging Face Daily Papers research 1mo ago Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark Abstract Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics fail regularly: acquisition systems start late, individual sensors die, and signal dropouts cluster precisely when a… 13 r/MachineLearning community 1mo ago Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII [P] ASCIITermDraw-Bench: Can a Model Actually Draw in ASCII? Do we really need a image generator to relay our thoughts about - an architecture? a topology? a cluster og N nodes? Is it possible to let our AI assistants, easily absorb and understand and make possible changes easily… 10 Page 4 of 10 · 500 articles ← Newer Older →