News / #hardware Tag Hardware 500 articles archived under #hardware · RSS Sign in to follow arXiv — Machine Learning research 1mo ago Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes arXiv:2607.05464v1 Announce Type: new Abstract: The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categorical… 19 arXiv — Machine Learning research 1mo ago Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy arXiv:2607.05469v1 Announce Type: new Abstract: Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Graph Contrastive Learning has demonstrated promising performance, existing methods often suffer… 30 arXiv — Machine Learning research 1mo ago Modeling Normal Is All You Need: Joint Latent Clustering for Anomaly Detection in Multimodal Cyber-Physical Systems arXiv:2607.06094v1 Announce Type: new Abstract: Faults on a cyber-physical system (CPS) are too rare and unrepresentative to characterise, or even to select a model on, so detection must instead model normal behaviour; the standard point-adjusted evaluation, however, rewards… 26 arXiv — Machine Learning research 1mo ago Performance Optimization and Comparative Analysis of Generative AI Models on Advanced Accelerators arXiv:2607.05400v1 Announce Type: cross Abstract: Generative AI models, such as Large Language Models (LLMs) and diffusion models, have demonstrated impressive performance across a wide range of tasks. Despite these advances, deployment remains challenging due to substantial… 25 Ars Technica — AI news-outlet 1mo ago Data centers’ energy demand threatens Trump’s “Made in America” plan Squeeze on Rust Belt electricity bills threatens Trump’s manufacturing plan. 32 Hugging Face Daily Papers research 1mo ago Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Abstract Temporal aggregation methods for speech-based depression detection show inconsistent performance across different backbones and training runs, highlighting the need for robust benchmarking criteria. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Speech-based depression… 6 arXiv — Machine Learning research 1mo ago Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation arXiv:2607.03007v1 Announce Type: new Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains often come without reliable structural grounding. In particular, existing approaches… 12 arXiv — Machine Learning research 1mo ago Heterogeneous Graph Condensation via Role-Aware Clustering arXiv:2607.03097v1 Announce Type: new Abstract: Heterogeneous Graph Neural Networks (HGNNs) have exhibited remarkable efficacy in modeling complex systems with multiple types of nodes and relations, yet their training on large-scale heterogeneous graphs remains computationally… 37 Hugging Face Daily Papers research 1mo ago Measuring the Gap Between Human and LLM Research Ideas Abstract Large language models generate research ideas that cluster around specific opportunity patterns and paradigms, diverging systematically from the broader and more diverse distributions found in human research papers. Generated by Qwen/Qwen2.5-Coder-32B-Instruct LLMs are… 21 TechCrunch — AI news-outlet 1mo ago Station F ramps up as a launchpad for Europe’s hottest AI startups Station F, a Paris-based startup hub founded by French billionaire Xavier Niel, is gearing up for a new edition of its F/ai accelerator program in a bid to strengthen its positioning as a stepping stone for promising AI startups. 20 r/LocalLLaMA community 1mo ago GLM 5.2 FP8 with FP8 KV - Terminal-Bench 2.1 = 79.8 (with one time-out that I didnt re-run) I wanted to test the official results vs fp8 + fp8 kv. basic sglang setup on H200. If anyone wants one of the official tests do ping me. I didnt rerun the one so it might go up a bit :) TERMINAL-BENCH 2.1 — FINAL RESULTS (mymodel via mini-swe-agent) TOTAL: 89 tasks PASSED: 71… 17 r/LocalLLaMA community 1mo ago Supra Reasoning Summarizer — a tiny model to summarize thinking traces from coding agents Hi, r/LocalLLaMA ! SupraLabs just released a model called: SupraLabs/reasoning-summarizer-800m-pre-gguf It is a thought trace summarizer. Basically, the user/dev can send a reasoning + tool calls (if you want), and then the model will generate a JSON which looks liek this: {… 4 r/LocalLLaMA community 1mo ago Learning to write AI harness old fashioned way. Need help with attention drift and ignoring tool call results! I've been writing a no-compile node.js based AI Harness for llama.cpp as a learning exercise and can really use some help. I'm basing my code off https://github.com/av/mi and https://pi.dev/ with really basic agentic loops. It basically loop until there are no more tool calls… 7 r/LocalLLaMA community 1mo ago [RELEASE] Supra-Router-51M - a tiny prompt routing model/orchestrator Hey r/LocalLLaMA ! SupraLabs is back with a new model. Supra-Router-51M. Basically, this is a model which is made for routing requests to smaller or bigger models - based on the user prompt/request. This is so cool, because it has only 51M parameters and so it can be used in… 34 r/LocalLLaMA community 1mo ago Considering Buying Another RTX 3090 - Benefits? Currently using dual RTX 3090s, and am happy with it. But never satisfied lol :) I know I've basically maxed out my single stream TPS. (140+ on standard benchmarks now). But I only have 48GB VRAM, So I can only do two concurrent requests @ 256k Context Length, anymore and my… 6 r/LocalLLaMA community 1mo ago Best choice of model 40B+ Parameters currently using Qwen3.6 35B as my main assistant model + coding agent but I think sometimes it misses basical general knowledge things, and it is more like executioner that assistant. That's why I though should I go with bigger models, But I don't want to lose speed I am on… 38 Hacker News — AI on Front Page community 1mo ago GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance Article URL: https://github.com/openai/codex/issues/30364 Comments URL: https://news.ycombinator.com/item?id=48789428 Points: 213 # Comments: 75 28 r/LocalLLaMA community 1mo ago Gemma 4 12B - MLX Kernel I've mentioned this kernel project I was working on in a few posts and figured I would just open the project code for anyone curious: MLX Gemma 12B The main constraints for this on my end is an M5 16GB Macbook Pro. I usually do a model development on clusters in the cloud but… 26 arXiv — Machine Learning research 1mo ago Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander arXiv:2607.01736v1 Announce Type: new Abstract: We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choosing the right checkpoint from a world-model training run is difficult: validation loss and… 17 arXiv — NLP / Computation & Language research 1mo ago ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators arXiv:2607.01602v1 Announce Type: new Abstract: SRAM-based FPGAs provide an attractive platform for energy- and latency-constrained CNN inference at the network edge, yet transient faults can lead to silent errors that compromise reliability. Always-on redundancy (e.g., full… 17 arXiv — NLP / Computation & Language research 1mo ago Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts arXiv:2607.01852v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information outside their parameters. We evaluate if cluster-based semantic chunking improves… 29 arXiv — NLP / Computation & Language research 1mo ago Introduction to Transformers: an NLP Perspective arXiv:2311.17633v2 Announce Type: replace Abstract: Transformers have dominated empirical machine learning models of natural language processing. In this paper, we introduce basic concepts of Transformers and present key techniques that form the recent advances of these models.… 22 r/MachineLearning community 1mo ago Improving machine-translated novels via style transfer — looking for advice on the faithfulness/fluency tradeoff [P] Hey all. I recently started working on a project to improve machine-translated webnovels via style transfer. The basic idea is to take the clunky translated prose and rewrite it to something that reads like it was written by a professional author, while remaining as faithful as… 22 r/LocalLLaMA community 1mo ago openlumara, my manually coded super-token-efficient harness, now works across any UI that can connect to an openAI endpoint! koboldlite, openwebui, you name it. basically, openAI bridge. yay! this was a long time coming, but it's finally here! you can now basically supercharge whichever UI you're already using with the power of openlumara . click that link for more information about openlumara itself. TL;DR: super token efficient framework built from the ground up… 25 Ars Technica — AI news-outlet 1mo ago Google’s AI buildout drove 37% increase in electricity use in 2025 Google tries balancing AI data center emissions with clean energy efforts. 34 arXiv — NLP / Computation & Language research 1mo ago Controllable Narrative Rendering for Enhanced Assisted Writing arXiv:2607.00009v1 Announce Type: new Abstract: Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure. This issue manifests as an oscillation… 13 arXiv — NLP / Computation & Language research 1mo ago Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match arXiv:2607.00185v1 Announce Type: new Abstract: Khipus--knotted cord devices--were the primary recording medium of the Inka Empire (c. 1400-1532 CE), yet their system remains undeciphered. We present a reproducible machine-learning pipeline applied to the Open Khipu Repository… 29 Hugging Face Daily Papers research 1mo ago Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models Abstract Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 35 r/LocalLLaMA community 2mo ago LokalBot - fully local macOS app: meetings, autocomplete, and day tracking that all run on your machine with a user friendly UI Been lurking here a while, this sub is basically why LokalBot exists. It's a Mac app that records + summarizes your meetings, autocompletes your typing in any app, and tracks where your day went, with every model running on-device . No cloud, no account, no API keys. Most of the… 15 arXiv — Machine Learning research 2mo ago TabPATE: Differentially Private Tabular In-Context Learning Without Public Data arXiv:2606.31474v1 Announce Type: new Abstract: Tabular foundation models enable accurate in-context learning (ICL) from small labeled datasets, but the private records placed in context can leak through model predictions. We first show that even basic membership inference… 38 Hacker News — AI on Front Page community 2mo ago County with 37 Data Centers Asks Schools to 'Conserve Electricity' Article URL: https://www.404media.co/henrico-virginia-datacenter-energy-cost-email/ Comments URL: https://news.ycombinator.com/item?id=48734699 Points: 209 # Comments: 105 19 arXiv — Machine Learning research 2mo ago scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering arXiv:2606.28459v1 Announce Type: new Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical noise hinder robust expression representation and cell graph construction.… 27 arXiv — Machine Learning research 2mo ago Improving Patient Subtyping on Longitudinal Data using Representations from Mamba-based Architecture arXiv:2606.28623v1 Announce Type: new Abstract: Effective sub-typing (also known as grouping or clustering) of patients using their electronic health record (EHR) data can greatly inform precision medicine efforts. However, subtyping temporal EHR datasets is known to be… 37 arXiv — Machine Learning research 2mo ago Nonlinear mixture model motivated subspace clustering arXiv:2606.29261v1 Announce Type: new Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separation (BSS) to represent a D-dimensional observation vector as an unknown… 7 r/LocalLLaMA community 2mo ago Instead of decentralized training effort we should build the “One dataset” There are many threads here calling for united LLM training run of a new open model. Mainly, after govt. stunt of banning commercial frontier models. And also due to the lack of small-medium open-weight models releases lately. I genuinelly believe at some point we’ll have “SETI… 38 Import AI (Jack Clark) community 2mo ago Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era What eras bookend our interregnum? 36 TechCrunch — AI news-outlet 2mo ago Omen AI’s plan to optimize data centers is all wet Omen AI raised a $31 million Series A to monitor chip coolant and stop bacterial outbreaks in data centers. 8 llama.cpp releases dev-tools 2mo ago b9840 DeepSeek V4 ( #24162 ) convert: add dsv4 conversion add basic setup add llm_graph_input_dsv4 add save-load state add sinkhorn eps - correction by @fairydreaming add rope fix cleanup dead code fix bugs support pro model: added by @fairydreaming remove redundant V cache Chat… 26 arXiv — Machine Learning research 2mo ago Dual-Learning based Penalized Multi-Align Clustering for Multi-View Incomplete and Disorderly Data arXiv:2606.27984v1 Announce Type: new Abstract: Multimodal feature fusion can effectively capture complex patterns in real-world data by integrating complementary information from different modalities. However, in many applications, such as boiler combustion monitoring,… 18 arXiv — NLP / Computation & Language research 2mo ago Mechanism-Driven Monitors for Preemptive Detection of LLM Training Instability arXiv:2606.28116v1 Announce Type: new Abstract: Frontier large language model training consumes massive accelerator fleets and long wall-clock computation, making stability failures costly when they occur. After a numerical or a hyperparameter fault has already destabilized the… 31 arXiv — NLP / Computation & Language research 2mo ago Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving arXiv:2606.27457v1 Announce Type: cross Abstract: Efficient deployment of large language models (LLMs) in production forces a trade-off between accuracy and cost. Operators often default to a single model that is either expensive for easy queries or insufficient for hard ones.… 20 arXiv — NLP / Computation & Language research 2mo ago DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions arXiv:2606.28048v1 Announce Type: cross Abstract: Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at FNOL. While recent fraud detection methods mainly rely on structured data, text, or… 19 r/LocalLLaMA community 2mo ago Success story with MiMo-V2.5-GGUF:UD-Q5_K_XL I don't see many stories about this model, but after several attempts (after I finished finally reconfiguring my cluster) I did something useful with it: it wrote a built-in llama.cpp tool for executing C++ code and using the results. Here's an exercise that MiMo V2.5 gave me to… 27 Hacker News — AI on Front Page community 2mo ago AMD Strix Halo RDMA Cluster Setup Guide Article URL: https://github.com/kyuz0/amd-strix-halo-vllm-toolboxes/blob/main/rdma_cluster/setup_guide.md Comments URL: https://news.ycombinator.com/item?id=48703258 Points: 207 # Comments: 61 22 TechCrunch — AI news-outlet 2mo ago SoftBank’s CEO isn’t the only one with questions about Elon Musk’s orbital data center hype Not everyone is buying Elon Musk’s vision for orbital data centers. 19 r/MachineLearning community 2mo ago Kicking off GPU Mode [D] Hey ! I’m starting a series to document my work on GPU infrastructure, LLMs, and CV. Stop #1 is up: A brief look at why GPUs are the center of the industry, the CPU/GPU divide, and why nvidia-smi is the first place you check when things break. We’ll move past the basics quickly… 27 r/MachineLearning community 2mo ago Roast my 3-year roadmap: Pivoting from Python/BaaS to AI Infrastructure & Go (Graduating 2029) [D] I'm a B.Tech student in India graduating in mid-2029. Currently, I know Python, SQL, Docker, basic prompt engineering, and I've built a few LLM apps using BaaS like Supabase/Firebase. I’m running all this on an Intel i5 13th Gen laptop with an RTX 5050 (8GB VRAM). The Pivot: I… 10 TechCrunch — AI news-outlet 2mo ago Why everyone from OpenAI to SpaceX is building their own chips (and turning up the heat on Nvidia) Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending.   OpenAI just shared its plans to spice things up with Jalapeño, its custom inference chip built with Broadcom, joining Google, Apple, and SpaceX in a growing list… 35 r/LocalLLaMA community 2mo ago Why do people keep investing in Intel for AI? If you get a good deal on some Xeons with a lot of memory bandwidth, or a cheap GPU for home inference, that's cool, no disrespect. But how in the hell are Wall Street types considering Intel part of the "AI picks and shovels" play? Who's buying Intel for their AI data centers?… 17 r/LocalLLaMA community 2mo ago 8 Tesla T4 Cards, what should it do? I have collected 8 Tesla T4 Datacenter Cards from a few retired VDI servers. I have one in a DEG1 and works ok on n its own. What should we do with the rest?   submitted by   /u/imonlysmarterthanyou [link]   [comments] 7 Page 6 of 10 · 500 articles ← Newer Older →