News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — Machine Learning research 10d ago Rethinking Privileged Information in On-Policy Self-Distillation arXiv:2608.18271v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a student on its own responses using token-level supervision from the same model conditioned on privileged reference information. We investigate whether performance gains from OPSD show… 30 arXiv — Machine Learning research 10d ago Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies arXiv:2608.18410v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Existing efficiency methods mainly reduce visual tokens, but aggressive token pruning becomes… 10 arXiv — Machine Learning research 10d ago MARCUS: Missing-Aware Region Representation with Contextual Urban Signals for Rent Prediction arXiv:2608.18546v1 Announce Type: new Abstract: Multimodal urban data has expanded the applications of urban region representation learning, such as functional zone identification and real estate appraisal, but also introduces challenges caused by data incompleteness. Existing… 13 arXiv — Machine Learning research 10d ago FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs arXiv:2608.18736v1 Announce Type: new Abstract: Federated training enables language models to learn from distributed private text, but the server cannot directly verify the local supervision or optimization process that produces each client update. A malicious client can… 25 arXiv — NLP / Computation & Language research 10d ago Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation arXiv:2608.19098v1 Announce Type: cross Abstract: Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision.… 6 arXiv — NLP / Computation & Language research 10d ago MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators arXiv:2608.18096v1 Announce Type: new Abstract: Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies,… 38 arXiv — NLP / Computation & Language research 10d ago You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language Models arXiv:2608.18116v1 Announce Type: new Abstract: Vision-language models enable zero-shot classification through natural language prompts, but performance is sensitive to prompt formulation, especially in specialized domains. Zero-shot Prompt Ensembling (ZPE) addresses this by… 29 arXiv — NLP / Computation & Language research 10d ago Backdoor Learning in Language Models and Vision-Language Models arXiv:2608.18095v1 Announce Type: new Abstract: Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through… 22 arXiv — NLP / Computation & Language research 10d ago Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models arXiv:2608.18132v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are typically built through a multi-stage pipeline consisting of cross-modal alignment, supervised fine-tuning (SFT), and preference optimization. This pipeline assumes that adapting an LLM… 30 arXiv — NLP / Computation & Language research 10d ago Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage arXiv:2608.18438v1 Announce Type: new Abstract: Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a "supervision gap" where novice therapists manage high-stakes risks with delayed professional feedback. This paper proposes a new… 5 arXiv — NLP / Computation & Language research 10d ago MedUAG: Unified Understanding and Generation for Medical Multimodal Models arXiv:2608.18937v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of… 13 arXiv — NLP / Computation & Language research 10d ago From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model arXiv:2608.18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. Although significant efforts are devoted to adapting VLMs at test time,… 7 arXiv — NLP / Computation & Language research 10d ago Multimodal Rapport Estimation in Real-World HRI arXiv:2608.18401v1 Announce Type: cross Abstract: Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their… 28 arXiv — NLP / Computation & Language research 10d ago Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference arXiv:2608.18591v1 Announce Type: cross Abstract: Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first… 35 arXiv — NLP / Computation & Language research 10d ago When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models arXiv:2608.18628v1 Announce Type: cross Abstract: Aligned vision-language models (VLMs) are designed to balance grounded visual reasoning with safe generation behavior. However, we observe a striking phenomenon: under safety-constrained instruction, models frequently abstain… 30 arXiv — NLP / Computation & Language research 10d ago ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models arXiv:2608.19075v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content during decoding calls for a candidate-specific measure of how strongly the image supports… 10 Hugging Face Daily Papers research 10d ago SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Abstract Semantic task completion video generation evaluates whether generated videos achieve intended outcomes with semantic grounding, supported by a curated dataset and vision-language model-based benchmark. Generated by thinkingmachines/Inkling-Small We introduce Semantic… 22 r/LocalLLaMA community 10d ago We quantized the new Ornith 1.5 9B and 35B-A3B ornith lab dropped new ornith 1.5 today, a 9B dense with vision and a 35B-A3B MoE, both MIT, trained on a loop that generates its own tasks. in addition there was giant 397b model, but we didn't quantize it (but if you want to try - we will do it) we made our AD (Atomic Dynamic)… 21 Don't Worry About the Vase community 10d ago OpenAI Takes Initial Steps To Address Its Alignment Problems OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision. 13 r/LocalLLaMA community 10d ago Reverse-Engineering the RK3588 NPU: Building an Open Compiler to Run GPT-2 at 36 tok/s Last year I posted about hacking the RK3588 NPU to run one vision encoder ( previous post ). This year I opened the whole thing up: reverse-engineered the register format, built an open compiler + runtime, and now GPT-2 and SigLIP run from PyTorch, ONNX, and JAX, no vendor SDK.… 30 NVIDIA Developer Blog official-blog 10d ago Building Federated Multimodal AI Workflows with NVIDIA FLARE Modern vision-language models (VLMs) can support tasks such as visual question answering, captioning, and image-text reasoning. In practice, however, the data... 36 Hugging Face Daily Papers research 10d ago CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation Abstract CardioState-JEPA learns a unified cardiac representation across ECG, PPG, and PCG by predicting masked latent physiological states with cross-modal delay alignment, improving downstream classification across all three modalities. Generated by… 6 Hugging Face Daily Papers research 11d ago V-RAE: Rethinking Video Latent Spaces for Generation Abstract V-RAE constructs semantically organized video latents from frozen vision representations to improve generation quality, convergence speed, and predictive modeling. Generated by thinkingmachines/Inkling-Small Latent video generation relies on autoencoders to define a… 32 Hugging Face Daily Papers research 11d ago MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding Abstract Mixture-of-Experts vision encoders with fine-grained topologies, auxiliary-loss-free balancing, and specialized kernels scale efficiently while outperforming larger dense models on image and video tasks. Generated by thinkingmachines/Inkling-Small Vision encoders are a… 21 arXiv — Machine Learning research 11d ago Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction arXiv:2608.16929v1 Announce Type: new Abstract: Predicting 30-day hospital readmission is essential for assessing patient stability and optimizing healthcare resources. As clinical risk evolves with the accumulation of evidence during hospitalization, capturing these dynamic… 14 arXiv — Machine Learning research 11d ago MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology arXiv:2608.16972v1 Announce Type: new Abstract: Machine learning has become an essential component of modern healthcare, where the integration of heterogeneous data sources offers unprecedented opportunities to improve clinical decision-making. Electronic Health Records (EHR)… 38 arXiv — Machine Learning research 11d ago Q-Learning With World Models arXiv:2608.17163v1 Announce Type: new Abstract: Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further… 9 arXiv — Machine Learning research 11d ago SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version arXiv:2608.17164v1 Announce Type: new Abstract: Textual context such as news, reports, and logs can provide valuable signals for time series forecasting, especially when future dynamics are driven by external events that are not yet visible in historical values. Existing… 20 arXiv — Machine Learning research 11d ago Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL arXiv:2608.17253v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward).… 37 arXiv — Machine Learning research 11d ago VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation arXiv:2608.16978v1 Announce Type: cross Abstract: Turning a frontier vision-language model into a robot policy usually means fine-tuning it to emit an action representation it never saw in pretraining, which throws away much of the reasoning that made the model worth reaching… 13 arXiv — NLP / Computation & Language research 11d ago Uncertainty-Aware Decision Making in Multimodal Large Language Models arXiv:2608.17084v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A… 26 arXiv — NLP / Computation & Language research 11d ago Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models arXiv:2608.17102v1 Announce Type: new Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize… 15 arXiv — NLP / Computation & Language research 11d ago Which Source Wins? Task-Dependent Reliance in Vision-Language Models arXiv:2608.17205v1 Announce Type: new Abstract: Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled… 29 arXiv — NLP / Computation & Language research 11d ago Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study arXiv:2608.17583v1 Announce Type: new Abstract: Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and… 23 arXiv — NLP / Computation & Language research 11d ago Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges arXiv:2608.17605v1 Announce Type: new Abstract: Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while… 30 arXiv — NLP / Computation & Language research 11d ago TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification arXiv:2608.17795v1 Announce Type: new Abstract: Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification… 29 arXiv — NLP / Computation & Language research 11d ago BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models arXiv:2608.17895v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize… 11 arXiv — NLP / Computation & Language research 11d ago Code as Representation: A Compilable Parsing Paradigm for Academic Documents arXiv:2608.17550v1 Announce Type: cross Abstract: Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core… 19 Vercel — AI dev-tools 11d ago Algolia joins the Vercel Marketplace You can now provision and manage Algolia directly from the Vercel Marketplace. Algolia is a hosted search platform. You send it your content, it builds an index, and your app queries that index instead of your database. Results come back in milliseconds, with typo tolerance,… 36 Hugging Face Daily Papers research 11d ago Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation Abstract TAMP-Nav improves embodied navigation by aligning vision-language models with 2D visual prompting, selective reasoning with compressed memory, and dense policy optimization. Generated by thinkingmachines/Inkling-Small Although Large Vision-Language Models (VLMs) have… 26 Hugging Face Daily Papers research 11d ago From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation Abstract A capability-driven data infrastructure with curriculum scheduling and specialized data engines trains large multimodal diffusion models on curated heterogeneous supervision for diverse generative tasks. Generated by thinkingmachines/Inkling-Small Large-scale image… 13 Hugging Face Daily Papers research 12d ago MOSS-VL Technical Report Abstract MOSS-VL is an open vision-language model family enabling real-time interaction by attending to vision via gated cross-attention during generation, using a synthesized interaction corpus and staged curriculum to achieve strong streaming performance with reduced… 25 Hugging Face Daily Papers research 12d ago TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation Abstract This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations. Generated by thinkingmachines/Inkling-Small Despite recent advances in unified multimodal models for multi-reference… 22 arXiv — Machine Learning research 12d ago LUNG-KGMM: Knowledge-Guided Multimodal Learning for Lung Cancer Incidence Prediction arXiv:2608.14657v1 Announce Type: new Abstract: Early identification of lung cancer risk is critical for timely intervention, yet existing prediction models are limited by their reliance on single data modalities and their inability to leverage structured clinical knowledge. We… 12 arXiv — Machine Learning research 12d ago PathFinder: Joint Decompositions of Linked Multimodal Datasets arXiv:2608.14951v1 Announce Type: new Abstract: Low-rank matrix decompositions can uncover patterns and structure in data and have a number of different applications across many disciplines. Extensions to "joint" low-rank decompositions have been proposed to link datasets from… 27 arXiv — Machine Learning research 12d ago GATTA: Graph Active Learning with Test-Time Augmentation arXiv:2608.15084v1 Announce Type: new Abstract: Test-time augmentation (TTA) has proven effective for improving model robustness and uncertainty estimation in computer vision, yet its application to graph-structured data remains largely unexplored. We introduce GATTA (Graph… 17 arXiv — Machine Learning research 12d ago UniFed-VLM: Federated Instruction Tuning for Vision-Language Models with Multiple Heterogeneity arXiv:2608.15516v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated strong performance in multimodal understanding and generation. However, fine-tuning of VLMs typically relies on centralized data, which raises privacy concerns in certain domains… 32 arXiv — Machine Learning research 12d ago Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group arXiv:2608.15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost. With adaptive acquisition, the policy determines which inputs are ultimately observed, so we state the guarantee… 13 arXiv — NLP / Computation & Language research 12d ago Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework arXiv:2608.14584v1 Announce Type: new Abstract: In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision… 29 arXiv — NLP / Computation & Language research 12d ago Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models arXiv:2608.14797v1 Announce Type: new Abstract: Large language models (LLMs) and large vision-language models (LVLMs) have demonstrated impressive generative capabilities, yet ensuring their outputs align with user intent is still challenging. While most existing approaches… 23 Page 4 of 10 · 500 articles ← Newer Older →