News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 27d ago ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification arXiv:2607.28637v1 Announce Type: new Abstract: This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our… 33 arXiv — NLP / Computation & Language research 27d ago TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal… 6 arXiv — NLP / Computation & Language research 27d ago BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning arXiv:2607.28966v1 Announce Type: new Abstract: Large language models often improve task performance by generating long reasoning traces, but the resulting computation is frequently wasted on redundant verification and revision. Existing probe-based early-exit approaches mainly… 4 arXiv — NLP / Computation & Language research 27d ago Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models arXiv:2607.29079v1 Announce Type: new Abstract: Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content. We study this serving-time consistency problem on 300 real images, comparing… 8 arXiv — NLP / Computation & Language research 27d ago Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding arXiv:2607.29196v1 Announce Type: new Abstract: Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, tracking later revisions, identifying intended objects or referents, and withholding… 19 arXiv — NLP / Computation & Language research 27d ago Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks arXiv:2607.29585v1 Announce Type: new Abstract: To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemically vigilant detect when new information conflicts… 8 arXiv — NLP / Computation & Language research 27d ago FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models arXiv:2607.29602v1 Announce Type: new Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic… 5 arXiv — NLP / Computation & Language research 27d ago Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods… 26 arXiv — NLP / Computation & Language research 27d ago WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates… 34 r/LocalLLaMA community 27d ago MiniMax-H3 now on huggingface MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks… 16 Hugging Face Daily Papers research 27d ago N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Abstract We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current… 24 Hugging Face Daily Papers research 28d ago RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models Abstract Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection… 36 Vercel — AI dev-tools 28d ago Qwen 3.8 Max now available on Vercel AI Gateway Qwen 3.8 Max is now available on AI Gateway. Qwen 3.8 Max handles text-only and vision-language work in one model, with 2.4 trillion parameters and a context window of up to 1 million tokens. The model is suited for software engineering and office productivity, along with visual… 13 r/MachineLearning community 29d ago What should we do for EMNLP commitment deadline? [R] We received the reviews, but they don't mention whether we should submit a revised version. Should we prepare one? I also couldn't find anywhere to upload a revision. What exactly is the EMNLP commitment deadline? I had assumed we were supposed to upload an updated version. Do… 22 TechCrunch — AI news-outlet 1mo ago Siri AI could come with a paywall for power users Apple CEO Tim Cook envisions users being able to buy more compute for Siri AI via Apple's existing iCloud+ subscriptions. 28 Hugging Face Daily Papers research 1mo ago Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Abstract Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes… 14 Hugging Face Daily Papers research 1mo ago ReToken: One Token to Improve Vision-Language Models for Visual Retrieval Abstract Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding… 22 Hugging Face Daily Papers research 1mo ago See2Think: Do Multimodal Models Really Use Intermediate Visual States? Abstract Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage… 9 arXiv — Machine Learning research 1mo ago Regularizing modality contribution drift in multimodal continual learning arXiv:2607.27260v1 Announce Type: new Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge. To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic… 36 arXiv — Machine Learning research 1mo ago Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision arXiv:2607.27274v1 Announce Type: new Abstract: EEG-based disease diagnosis requires one prediction per subject, yet common pipelines segment recordings into short instances, inherit the subject label for every instance, and train instance-level classifiers. This assumes that… 10 arXiv — Machine Learning research 1mo ago TIER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification arXiv:2607.27289v1 Announce Type: new Abstract: The promise of multimodal fusion lies in combining complementary sources of evidence, yet more evidence does not always yield a better prediction. Recent multimodal models have advanced fusion through richer cross-modal interaction… 12 arXiv — Machine Learning research 1mo ago Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral… 14 arXiv — Machine Learning research 1mo ago Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective arXiv:2607.27660v1 Announce Type: new Abstract: Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particular, the SCORE framework~\cite{majee2024score} demonstrated that SIMs can serve as… 13 arXiv — Machine Learning research 1mo ago FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning arXiv:2607.27665v1 Announce Type: new Abstract: Federated graph learning enables collaborative training over decentralized graph data without sharing raw graph information. As such risks evolve, clients must learn emerging classes from private multimodal graph streams, retain… 26 arXiv — Machine Learning research 1mo ago Flux-OPD: On-Policy Distillation with Evolving Contexts arXiv:2607.28022v1 Announce Type: new Abstract: Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision… 13 arXiv — Machine Learning research 1mo ago Contrastive Reinforced Policy Optimization via Privileged Self-Distillation arXiv:2607.28026v1 Announce Type: new Abstract: Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it… 12 arXiv — Machine Learning research 1mo ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger arXiv:2607.28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate… 18 arXiv — NLP / Computation & Language research 1mo ago AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes arXiv:2607.27393v1 Announce Type: new Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has… 4 arXiv — NLP / Computation & Language research 1mo ago Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built… 36 arXiv — NLP / Computation & Language research 1mo ago Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or… 29 arXiv — NLP / Computation & Language research 1mo ago Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis arXiv:2607.27790v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be… 23 arXiv — NLP / Computation & Language research 1mo ago AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying… 28 arXiv — NLP / Computation & Language research 1mo ago RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning arXiv:2607.28156v1 Announce Type: new Abstract: Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When… 18 arXiv — NLP / Computation & Language research 1mo ago Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models arXiv:2607.28166v1 Announce Type: new Abstract: Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide… 28 arXiv — NLP / Computation & Language research 1mo ago Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory arXiv:2607.28263v1 Announce Type: new Abstract: Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for prediction. We turn this division of labor into CoMem (Comprehension Memory),… 13 arXiv — NLP / Computation & Language research 1mo ago Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the… 35 arXiv — NLP / Computation & Language research 1mo ago Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a… 26 arXiv — NLP / Computation & Language research 1mo ago VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation arXiv:2607.28590v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual… 4 Hugging Face Daily Papers research 1mo ago LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Abstract Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct… 37 Hugging Face Daily Papers research 1mo ago Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Abstract GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI… 31 Hugging Face Daily Papers research 1mo ago Beacon: Knowing When and How to Perform Agentic Visual Reasoning Abstract The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic… 14 Hugging Face Daily Papers research 1mo ago Flux-OPD: On-Policy Distillation with Evolving Contexts Abstract Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student,… 24 r/LocalLLaMA community 1mo ago Minimax-H3 video model released, open weights coming in the next few days https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound,… 12 Hugging Face Daily Papers research 1mo ago Metis: Memory Foundation Model Abstract Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external… 7 Hugging Face Daily Papers research 1mo ago SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them Abstract Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason… 14 r/LocalLLaMA community 1mo ago Smallest model (& tips) for intelligent computer use via Hermes? Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive an actual machine via hermes' computer_use tool and cua_driver to click through… 7 r/LocalLLaMA community 1mo ago GLM 5.2 with vision on Hugging Face Hi all, I have not seen this model talked about here but it seems like baseten (inference provider on OpenRouter) merged the vision encoder from Kimi k2.6 into GLM 5.2. I think the lack of vision was one of the big complaint when GLM 5.2 came out, I have not tested this model… 34 Hugging Face Daily Papers research 1mo ago CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition Abstract Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical… 14 arXiv — NLP / Computation & Language research 1mo ago Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement arXiv:2607.26473v1 Announce Type: cross Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic… 14 arXiv — Machine Learning research 1mo ago What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations arXiv:2607.27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this?… 12 Page 10 of 10 · 500 articles ← Newer