News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — Machine Learning research 3d ago DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation arXiv:2608.26019v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. OPSD keeps this privileged teacher fixed, even though the student distribution and output… 36 arXiv — NLP / Computation & Language research 3d ago SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation arXiv:2608.25123v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves large language models by incorporating external knowledge without retraining, but existing methods often underuse the relational structure encoded in knowledge graphs. Graph-based RAG… 24 arXiv — NLP / Computation & Language research 3d ago OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora arXiv:2608.25398v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of a… 37 arXiv — NLP / Computation & Language research 3d ago Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models arXiv:2608.25662v1 Announce Type: new Abstract: In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in… 35 arXiv — NLP / Computation & Language research 3d ago VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following arXiv:2608.26013v1 Announce Type: new Abstract: Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feedback from… 7 arXiv — NLP / Computation & Language research 3d ago Beyond Local Surprise: Grounded Dialogue as Selective Belief Revision under Referential Uncertainty arXiv:2608.26035v1 Announce Type: new Abstract: When a speaker refers to a scene that the listener cannot directly see, the listener must decide whether to preserve its current understanding or revise it as new utterances arrive. Many language systems treat local mismatch as a… 16 arXiv — NLP / Computation & Language research 3d ago GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models arXiv:2608.25375v1 Announce Type: cross Abstract: Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or… 7 Hugging Face Daily Papers research 3d ago MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization Abstract MA-VLA enables multi-arm collaboration by assigning atomic actions to individual arms and using training-time permutations to generalize to unseen coordination patterns. Generated by thinkingmachines/Inkling-Small Multi-arm collaboration is becoming a core capability in… 31 Hugging Face Daily Papers research 3d ago StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models Abstract StreamPI enhances single-frame vision-language-action models with streaming temporal reasoning via instruction-anchored attention and randomized interval training, improving robot manipulation without extra parameters. Generated by thinkingmachines/Inkling-Small… 6 Hugging Face Daily Papers research 3d ago V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning Abstract Visual Rubrics-Based Reinforcement Learning improves vision-language model grounding by scoring answers on visual faithfulness, reasoning consistency, and instruction following using structured partial credit. Generated by thinkingmachines/Inkling-Small Vision-language… 22 Simon Willison community 3d ago Qwen3.8-Flash-Next Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost. I've been… 22 r/LocalLLaMA community 3d ago Any news about DeepSeek V4 Flash Vision weights? I'd be curious to try it locally since I use 0731 daily but still no news on the weights   submitted by   /u/LegacyRemaster [link]   [comments] 12 r/LocalLLaMA community 3d ago Forget the Pelican, it's Weevil-Time! / Benchmaxxing-Proof SVG and Vision Benchmark The Artist: Qwen3.8-27B-UD-Q3\ K_XL, q8_0 caches, xhigh, temp 1.0, image-min-tokens 1024, froggeric template) I was screwing around with different Qwen3.8-27B quants and thought of this very simplistic but seemingly bechmaxxing resistant combined SVG and vision test. Just let… 4 Hugging Face Daily Papers research 4d ago LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Abstract LAION-BVD is a large-scale open video dataset enabling multimodal pre-training across video, audio, and image modalities with synthetic captions and strong benchmark performance. Generated by thinkingmachines/Inkling-Small We present LAION-BVD, a large-scale open video… 22 r/LocalLLaMA community 4d ago First serious confirmation. Ox Alpha is GLM-5.3-Flash https://x.com/romanchernin/status/2092488160680751437?s=20 - Multimodal (Vision) - 1M Tokens Context Window - DeepSWE ~63%   submitted by   /u/MrWidmoreHK [link]   [comments] 14 Smol AI News news-outlet 4d ago not much happened today **Z.ai** launched **GLM-5.3-Flash**, a natively multimodal model with a **1M-token context window**, **320B total parameters / 18B active parameters**, under the **MIT License**. It is positioned as a price-competitive successor to GLM-5.2 and claims performance on par with… 29 Hugging Face Daily Papers research 4d ago Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Abstract OraRL improves reinforcement learning post-training for video multimodal language models by integrating oracle rollouts with decoupled advantage estimation and sign-balanced pruning, achieving higher sample efficiency and scalability without chain-of-thought generation.… 17 arXiv — Machine Learning research 4d ago Joint Distribution Alignment for Universal Domain Adaptation arXiv:2608.24429v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) has been widely concerned in the fields of machine learning, pattern recognition, and computer vision. Traditional UDA learning usually assumes that the label spaces of the source and target… 17 arXiv — Machine Learning research 4d ago A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology arXiv:2608.24688v1 Announce Type: new Abstract: Precision oncology necessitates a longitudinal model of patient state that captures cancer evolution and treatment over time, integrating multimodal observations. We introduce the oFM, a foundation model developed on a real-world… 4 arXiv — Machine Learning research 4d ago LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning arXiv:2608.24795v1 Announce Type: new Abstract: Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multimodal-attributed graphs. This advancement significantly enhances data… 14 arXiv — Machine Learning research 4d ago BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval arXiv:2608.24823v1 Announce Type: new Abstract: Spatially resolved biology requires representations that preserve biological neighborhood structure rather than only exact cross-modal correspondences. Existing histology--transcriptomics objectives can emphasize instance-level… 10 arXiv — Machine Learning research 4d ago Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control arXiv:2608.07265v2 Announce Type: cross Abstract: We establish a finite-sample learning-to-control theory for geometrically supervised latent models of nonlinear deterministic systems. Geometric supervision is used only during training: simulator state, proprioception, or state… 15 arXiv — Machine Learning research 4d ago DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection arXiv:2608.22368v1 Announce Type: cross Abstract: While linear attention is a compelling mechanism for high-resolution object detection due to its reduced cost for global token mixing, converting the Softmax-attention ViT backbone of a trained detector into a linear-attention… 28 arXiv — Machine Learning research 4d ago InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis arXiv:2608.23574v1 Announce Type: cross Abstract: Each WSI slide contains thousands of candidate tissue patches, while supervision is usually available only at slide level. Existing bag-construction strategies like Uniform extraction and handcrafted heuristics do not control… 7 arXiv — Machine Learning research 4d ago The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models arXiv:2608.23634v1 Announce Type: cross Abstract: Many few-shot adaptation methods for vision-language models classify with a convex combination of the zero-shot text prototype and the mean of the K labelled image features, with a single blending ratio routinely tuned on… 17 arXiv — Machine Learning research 4d ago MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models arXiv:2608.23646v1 Announce Type: cross Abstract: Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug discovery, where reusable vector representations support property prediction, virtual screening, and retrieval. Most… 6 Hugging Face Daily Papers research 4d ago WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Abstract WeMM-Embedding is a family of universal multimodal embedding models that align text, images, videos, and interleaved inputs in a shared space, achieving state-of-the-art retrieval and recommendation performance across public benchmarks and large-scale WeChat… 15 Vercel — AI dev-tools 4d ago GLM 5.3 Flash now available on AI Gateway GLM 5.3 Flash from Z.ai is now available on AI Gateway. The model is a faster, cheaper sibling of GLM 5.3 built for coding and agent tasks that run across many steps. GLM-5.3 Flash is a multimodal model that supports text and vision input, with a 1M token context window and a… 4 Hugging Face Daily Papers research 4d ago GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture Abstract GigaBrain-0.7 is a vision-language-action model that improves embodied generalization via a three-system architecture, large-scale heterogeneous pretraining, and joint alignment training. Generated by thinkingmachines/Inkling-Small Vision-language-action (VLA) models… 31 r/LocalLLaMA community 4d ago CNBC Television: Nvidia partner with Perplexity AI to run locally in DGX Spark.   submitted by   /u/dd32x [link]   [comments] 8 llama.cpp releases dev-tools 4d ago v0.3.0 Overview llama.cpp 0.3.0 introduces the dots3-note multimodal model (with a new DSA-ISWA KV cache), MTP support for GLM-4.5-Air, and tensor-split ( -sm tensor ) plus multi-sequence rollback fixes for DeepSeek 4. ggml is bumped to v0.22.0 (meta-backend tensor split, per-op Metal… 23 r/LocalLLaMA community 5d ago tencent/WeMM-Embedding 9B/4B/2B WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.… 10 arXiv — NLP / Computation & Language research 5d ago Distinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation arXiv:2608.21364v1 Announce Type: new Abstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations must be updated accordingly. Incremental interpretation, therefore, depends not… 20 arXiv — NLP / Computation & Language research 5d ago Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding arXiv:2608.21415v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from… 9 arXiv — NLP / Computation & Language research 5d ago L\"etzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents arXiv:2608.21714v1 Announce Type: new Abstract: Recent page-image retrievers such as ColPali have improved retrieval over visually rich documents, yet little is known about how they behave in cross-lingual, low-resource settings. We introduce L\"etzCross, a benchmark for… 9 arXiv — NLP / Computation & Language research 5d ago MCite-RL: Towards Reliable Multimodal RAG via Citation-enhanced Agentic Reinforcement Learning arXiv:2608.21808v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (RAG) with visual citation is crucial for ensuring the traceability and verifiability of MLLMs. However, current RAG and SFT-based methods struggle to achieve robust cross-modal reasoning,… 18 arXiv — NLP / Computation & Language research 5d ago GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding arXiv:2608.21832v1 Announce Type: new Abstract: Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce… 37 arXiv — NLP / Computation & Language research 5d ago PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding arXiv:2608.21853v1 Announce Type: new Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing… 12 arXiv — NLP / Computation & Language research 5d ago Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction arXiv:2608.22071v1 Announce Type: new Abstract: Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language models for… 30 arXiv — NLP / Computation & Language research 5d ago Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models arXiv:2608.22312v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the… 13 arXiv — NLP / Computation & Language research 5d ago Semantics or Structure? Auditing Text Sensitivity in Multimodal Time-Series Forecasting arXiv:2608.22321v1 Announce Type: new Abstract: Multimodal time-series forecasting has emerged as a promising paradigm in which natural-language context is expected to improve predictive performance. Recent multimodal foundation models, including Aurora, as well as early- and… 35 arXiv — NLP / Computation & Language research 5d ago Context-Aware Cluster Decoding: Semantic Anchor-Driven Coherence in dMLLMs arXiv:2608.22367v1 Announce Type: new Abstract: Diffusion multimodal large language models (dMLLMs) frequently produce long-form outputs marred by semantic drift and repetition, with quality generally degrading as output length increases. We identify two structural deficiencies… 4 arXiv — NLP / Computation & Language research 5d ago ProBel: Propaganda Detection with Techniques, Spans, and Explanations arXiv:2608.22388v1 Announce Type: new Abstract: Propaganda detection includes several related prediction levels, ranging from sentence-level decisions to technique classification and span identification. However, it remains unclear how supervision at these levels interacts when… 10 arXiv — NLP / Computation & Language research 5d ago A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports arXiv:2608.22713v1 Announce Type: new Abstract: Clinical diagnosis requires progressive integration of patient history, physical examination, laboratory findings, medical images, and diagnostic-informative tests. However, most multimodal medical benchmarks evaluate fixed inputs… 38 arXiv — NLP / Computation & Language research 5d ago SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning arXiv:2608.22857v1 Announce Type: new Abstract: Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names,… 20 arXiv — NLP / Computation & Language research 5d ago Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models? arXiv:2608.22916v1 Announce Type: new Abstract: Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this encoded information reaches the answer. We address this… 33 Vercel — AI dev-tools 5d ago The end of credential sprawl for agents Every useful agent reaches beyond your codebase. It posts to Slack, opens pull requests, queries Snowflake, or calls an internal API. That reach is what makes it valuable, and it's also where the risk lives, because for years, granting it meant provisioning a long-lived token… 5 Hugging Face Daily Papers research 5d ago Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision Abstract A hierarchical taxonomy and dense supervision strategy improve diffusion-based image editing through fine-grained concepts, large-scale paired data, and granular evaluation. Generated by thinkingmachines/Inkling-Small Existing image editing frameworks predominantly… 17 Hugging Face Daily Papers research 6d ago UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Abstract A reparameterized pretrained vision transformer unifies semantic understanding, high-fidelity reconstruction, and image generation within a single visual space without requiring a separate VAE. Generated by thinkingmachines/Inkling-Small Semantic vision encoders have… 26 Hugging Face Daily Papers research 6d ago EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking Abstract EviRank reformulates multimodal image re-ranking as semantic constraint satisfaction by parsing queries into structured evidence packages and verifying candidates via rubric scoring and listwise comparison without training. Generated by thinkingmachines/Inkling-Small… 35 Page 2 of 10 · 500 articles ← Newer Older →