News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow Simon Willison community 9h ago Introducing Hy4 Preview Introducing Hy4 Preview New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face . This is a big size increase from their previous Hy3 in July, which was 295B, 21B… 9 r/LocalLLaMA community 19h ago Did anyone else notice the Ornith 1.5 35B GGUFs got a "silent" update? Hey all, I stumbled onto something funny and figured I'd share before it gets buried. I'd been running the Ornith 1.5 35B A3B GGUF (Q4_K_M) for a while, then recently a new revision showed up in my HF cache and I figured, eh, I'd pruned my old one. Fast forward a bit later when… 6 r/MachineLearning community 1d ago How important is having an internship to get a good job for ML PhD in USA? [D] Hey everyone, I'm an international student studying in the US. I'm on track to graduate late next year. My research is not exactly ML, it is in 3D computer vision but have decent exposure to ML as well. In case you didn't know, the CPT program (which let's internation students… 10 r/LocalLLaMA community 1d ago Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM) I've got Qwen3.8-Flash-next running on RTX 3090, Ryzen 9 3950X, a PCIe 3.0 motherboard, and 64GB DDR RAM from 2020. IQ4_XS weights, full kvarn5 context, vision on GPU, experts in host RAM, n-grams on disk. MTP works but actually slows decode down even with 80% draft acceptance,… 18 arXiv — Machine Learning research 2d ago NeoTriFuse: Reliability-Aware Multimodal Fusion under Missingness Heterogeneity for Neonatal Mortality Risk Prediction arXiv:2608.26436v1 Announce Type: new Abstract: Neonatal mortality risk prediction from bedside monitoring data remains challenging due to extreme class imbalance, heterogeneous clinical risk factors, multi-scale temporal dynamics, and substantial missingness. We propose… 16 arXiv — Machine Learning research 2d ago Chart2SVG: Editable SVG Generation from Raster Chart Images arXiv:2608.26544v1 Announce Type: new Abstract: We present Chart2SVG, a multimodal large language model that converts static raster charts into structurally organized, semantically enriched SVGs that support programmatic editing. By incorporating chart-specific semantic tokens… 4 arXiv — Machine Learning research 2d ago Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs arXiv:2608.26581v1 Announce Type: new Abstract: Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs). Recent hardware support for low-precision formats, ranging from MXFP8 to ultra-low-bit… 25 arXiv — Machine Learning research 2d ago J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data arXiv:2608.26582v1 Announce Type: new Abstract: Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains,… 30 arXiv — Machine Learning research 2d ago Simple Actors and Deep Critics for Scalable Reinforcement Learning arXiv:2608.26659v1 Announce Type: new Abstract: Recent progress in offline reinforcement learning (RL) has been driven by expressive generative actors such as diffusion and flow-matching policies, which capture multimodal behavior in offline datasets. However, these actors… 16 arXiv — Machine Learning research 2d ago SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting arXiv:2608.26829v1 Announce Type: new Abstract: Time series forecasting models operate on raw numerical sequences, lacking the semantic knowledge that domain experts implicitly leverage, such as the physical meaning of each variable, its statistical behavior, and its temporal… 22 arXiv — Machine Learning research 2d ago Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion arXiv:2608.26879v1 Announce Type: new Abstract: Fusing multiple modalities is expected to improve model performance. However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). Pathway isolation of a… 18 arXiv — Machine Learning research 2d ago Graph-Based Pseudo-multimodal Contrastive Learning for 12-Lead ECG Representations arXiv:2608.26964v1 Announce Type: new Abstract: 12-lead electrocardiogram (ECG) is a standard, non-invasive examination widely used for diagnosing coronary artery disease, where clinical interpretation relies on comparing waveform patterns across multiple leads. However, most… 35 arXiv — Machine Learning research 2d ago Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation arXiv:2608.27158v1 Announce Type: new Abstract: Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep,… 10 arXiv — Machine Learning research 2d ago Importance Scoring of Transformer Attention Heads in Learning Tabular Data arXiv:2608.27241v1 Announce Type: new Abstract: Computationally demanding and opaque deep learning models can be better understood and optimized by analyzing how they transform data. While deep transformers have been widely studied in computer vision and natural language… 14 arXiv — Machine Learning research 2d ago MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework arXiv:2608.27286v1 Announce Type: new Abstract: Inferring molecular structures from multimodal spectroscopic measurements requires integrating complementary yet highly heterogeneous signals. However, the common paradigm of directly concatenating multispectral sequences can… 27 arXiv — NLP / Computation & Language research 2d ago AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking arXiv:2608.26141v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models… 13 arXiv — Machine Learning research 2d ago ClassVision: AI-Powered Classroom Attendance System arXiv:2608.26173v1 Announce Type: cross Abstract: Students and working professionals have to go through the attendance process every day. Traditional methods of marking attendance using pen and paper or online platforms are human-intensive and time-consuming. To address the… 31 arXiv — NLP / Computation & Language research 2d ago Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation arXiv:2608.26142v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While recent advances in Multimodal Large Language Models (MLLMs) have substantially boosted RES… 15 arXiv — NLP / Computation & Language research 2d ago Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for… 28 arXiv — NLP / Computation & Language research 2d ago VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation arXiv:2608.26155v1 Announce Type: new Abstract: Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of high-quality non-English image-text… 6 arXiv — NLP / Computation & Language research 2d ago Case2Flow: Bridging Patient Cases and Guideline Flowcharts through Multimodal Retrieval arXiv:2608.26414v1 Announce Type: new Abstract: Medical guidelines encode rich, evidence-based decision logic, yet the specific decision artifact a clinician needs is hard to locate within a guideline, let alone across guidelines covering plausible diseases and treatments. While… 26 arXiv — NLP / Computation & Language research 2d ago SPT: Skills as Pre-Training Data for Agentic Language Models arXiv:2608.26563v1 Announce Type: new Abstract: Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution,… 13 arXiv — NLP / Computation & Language research 2d ago Not Just Reason, Not Just Scan: Reinforcement Learning for Proactive Scientific Error Verification over Academic Paper arXiv:2608.26596v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly capable scientific assistants, yet they remain far from fully autonomous research. This transition requires models to actively inspect academic papers, build global evidence… 4 arXiv — NLP / Computation & Language research 2d ago Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs arXiv:2608.26641v1 Announce Type: new Abstract: Decoding quality in diffusion multimodal language models (dMLLMs) depends heavily on the order in which masked tokens are committed. Existing confidence-based strategies prioritize locally easy tokens, but confidence does not… 19 arXiv — NLP / Computation & Language research 2d ago Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study arXiv:2608.26697v1 Announce Type: new Abstract: Synthetic speech provides scalable supervision for automatic speech recognition (ASR), but its benefit depends on the selected texts, reference speech, and amount of synthesized data. We present a unified phoneme-based TTS-to-ASR… 14 arXiv — NLP / Computation & Language research 2d ago DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali arXiv:2608.27110v1 Announce Type: new Abstract: Reliable medical conversational AI requires authentic expert--patient interaction data, yet such datasets remain scarce, especially for low-resource languages such as Bengali. We present DocTalkBN, a large-scale multimodal dataset… 13 arXiv — NLP / Computation & Language research 2d ago Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models arXiv:2608.27135v1 Announce Type: new Abstract: Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yield… 32 Hugging Face Daily Papers research 2d ago Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning Abstract Aphanta evaluates when image-editing intermediates improve multimodal reasoning by testing direct, editor-generated, and idealized visual states across tasks. Generated by thinkingmachines/Inkling-Small Explicit visual intermediates can help multimodal large language… 26 Hugging Face Daily Papers research 2d ago CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension Abstract CaRGo-T improves multimodal humor understanding by modeling causal relationships as graph-based reasoning structures interpreted by vision-language models. Generated by thinkingmachines/Inkling-Small Large-scale vision-language models (VLMs) have demonstrated remarkable… 5 Hugging Face Daily Papers research 2d ago UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Abstract UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior. Generated by… 37 Hacker News — AI on Front Page community 2d ago We found a division by zero bug in FFmpeg with a vibecoded fuzzer Article URL: https://code.ffmpeg.org/FFmpeg/FFmpeg/issues/24290 Comments URL: https://news.ycombinator.com/item?id=49468642 Points: 200 # Comments: 155 21 r/LocalLLaMA community 2d ago I used local Qwen 27b to build a harness and replace OpenCode Sharing my harness for running local LLMs that I built using Qwen 3.x 27B (> 90% locally built) under my supervision - not vibe-coded. Its free, no telemetry, and open-source (AGPL). Works on Windows, Linux (sorry, no Mac yet). I use it for my own coding + mixed workflows. How… 13 Hugging Face Daily Papers research 3d ago Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans Abstract A real-time framework for online co-speech gesture generation uses a causal multimodal autoregressive model with streaming speech and motion history, supported by synthetic dialogue data and continual user-feedback adaptation. Generated by thinkingmachines/Inkling-Small… 13 Hugging Face Daily Papers research 3d ago A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans Abstract A modular medical imaging agent decomposes spatial relation verification into parsing, anatomical localization, and geometric rules to outperform end-to-end vision-language models on CT spatial reasoning. Generated by thinkingmachines/Inkling-Small Reliable spatial… 27 Hugging Face Daily Papers research 3d ago Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction Abstract A multimodal Turkish dialogue dataset and genetic algorithm-optimized interpretable rules are used to predict turn transitions from visual, acoustic, and linguistic cues. Generated by thinkingmachines/Inkling-Small Turn-taking is a basic organizational feature of human… 10 Hugging Face Daily Papers research 3d ago Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Abstract A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) have shown strong performance in… 14 Hugging Face Daily Papers research 3d ago Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation Abstract RubSE improves UI-to-code generation stability by using rubric-guided self-evolution to prevent visual repair coupling and trajectory collapse. Generated by thinkingmachines/Inkling-Small Large vision-language models have shown strong progress in UI-to-code generation,… 26 arXiv — Machine Learning research 3d ago CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery arXiv:2608.24947v1 Announce Type: new Abstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based… 35 arXiv — Machine Learning research 3d ago MSR-IVA: Masked Structural Residual Independent Vector Analysis for State-Aware Fusion of Structural MRI and Dynamic Functional Network Connectivity arXiv:2608.24978v1 Announce Type: new Abstract: Multimodal fusion of structural MRI (sMRI) and dynamic functional network connectivity (dFNC) can reveal how brain structure relates to changing functional states. When the same structural latent representation is coupled with… 35 arXiv — Machine Learning research 3d ago Multimodal Injury Risk Prediction in Tennis arXiv:2608.25126v1 Announce Type: new Abstract: Machine learning has had a significant positive impact on the prediction of athlete performance and injury risk. Most works in this field rely on subjective observations and expert assessments, which restrict their effectiveness.… 29 arXiv — Machine Learning research 3d ago Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning arXiv:2608.25350v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training. Inspired by the success of preference-based reward… 29 arXiv — Machine Learning research 3d ago PaSta: Noisy Node Classification with Partial Label Learning arXiv:2608.25365v1 Announce Type: new Abstract: Noisy node classification problem is a fundamental yet challenging task for real-world graph-related web services, where node labels are often corrupted or unreliable due to weak supervision or automatic annotation. However,… 16 arXiv — Machine Learning research 3d ago Resolving Multi-Modal Regression by Difference-Quotient-Based Clustering:Fast Coarse Conditional-Label Assignment arXiv:2608.25467v1 Announce Type: new Abstract: Multimodal regression suffers from the mean-collapse pathology: under squared loss, an unconstrained regressor converges to the conditional mean, which for K > 1 lies away from all modes. We attribute this failure to pairwise… 31 arXiv — NLP / Computation & Language research 3d ago Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference arXiv:2608.25542v1 Announce Type: cross Abstract: Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection… 18 arXiv — Machine Learning research 3d ago Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory arXiv:2608.25570v1 Announce Type: new Abstract: Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution… 19 arXiv — Machine Learning research 3d ago Fairness-Aware Test-Time Prompt Tuning arXiv:2608.25707v1 Announce Type: new Abstract: Vision-language models have displayed remarkable capabilities in multi-modal understanding and are increasingly used in critical applications where economic and practical deployment constraints prohibit re-training or fine-tuning.… 25 arXiv — NLP / Computation & Language research 3d ago Why Does Graph Learning Fail to Fully Benefit from a Text Teacher? arXiv:2608.25741v1 Announce Type: cross Abstract: Graph neural networks (GNNs) are widely used to represent complex interactions and relationships among entities. We investigate a multimodal model that combines two complementary ideas: a self-supervised method that enables a GNN… 27 arXiv — Machine Learning research 3d ago Learning from waste: Machine Learning for health risk prediction and computer vision-based sorting in Ghana arXiv:2608.25759v1 Announce Type: new Abstract: The inappropriate disposal of solid waste remains a significant public health and environmental concern worldwide, including in Ghana. Poor sanitation and improper waste management practices contribute to substantial economic costs… 15 arXiv — Machine Learning research 3d ago Drift-Aware Multimodal User Representation Learning via Multi-Scale Temporal Modeling and Sparse Mixture-of-Experts arXiv:2608.25773v1 Announce Type: new Abstract: Understanding user preferences from noisy and temporally evolving social media behaviors is fundamentally challenging due to interest drift, where user preferences shift across time and exhibit both multi-scale temporal patterns… 12 arXiv — Machine Learning research 3d ago One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation arXiv:2608.25936v1 Announce Type: new Abstract: On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it… 32 Page 1 of 10 · 500 articles Older →