News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow arXiv — NLP / Computation & Language research 12d ago Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints arXiv:2608.15032v1 Announce Type: new Abstract: Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the… 30 arXiv — NLP / Computation & Language research 12d ago Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs arXiv:2608.15085v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) exhibit substantial performance degradation in non-English visual reasoning, despite the strong multilingual competence of their text-only backbones. While mechanistic evidence from… 38 arXiv — NLP / Computation & Language research 12d ago Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval arXiv:2608.16071v1 Announce Type: new Abstract: Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples… 31 Hugging Face Daily Papers research 12d ago PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments Abstract PACE-Bench evaluates self-evolving agents on physics adaptation tasks requiring iterative code redesign after environmental mutations, revealing that simulator-grounded reflection outperforms unverified self-revision but mechanism redesign remains a major bottleneck.… 30 Hugging Face Daily Papers research 12d ago VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding Abstract VideoGAIA introduces a multi-turn, tool-augmented benchmark that evaluates agentic video understanding for advanced multimodal models through complex real-world tasks. Generated by thinkingmachines/Inkling-Small Video understanding is a fundamental task for evaluating… 4 Hugging Face Daily Papers research 12d ago VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Abstract A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers. Generated by… 14 Hugging Face Daily Papers research 12d ago NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents Abstract NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning. Generated by thinkingmachines/Inkling-Small… 17 r/LocalLLaMA community 12d ago Qwen3.8-27B Uncensored Aggressive is out with K_P quants and HauhauCS FastMTP (up to 3.02x TG)! The dense Qwen release is back! Qwen3.8-27B Uncensored Aggressive is out with the complete K_P quant range, Vision, native NextN, and HauhauCS FastMTP. Aggressive here means no refusals, no personality alterations, and very little preamble on difficult prompts. It keeps… 5 Hacker News — AI on Front Page community 13d ago GPT 5.6 Sol is the best "vision" model OpenAI ever released Article URL: https://blog.roboflow.com/openai-gpt-5-6/ Comments URL: https://news.ycombinator.com/item?id=49329575 Points: 216 # Comments: 108 16 Hugging Face Daily Papers research 13d ago A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images Abstract The ALD/E-ImageMiner benchmark and ICDAR 2026 competition advance machine interpretation of scientific figures through tasks spanning visual reading, domain reasoning, and evidential justification, proposing long-term goals for verifiable multimodal scientific AI.… 30 Hugging Face Daily Papers research 13d ago Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Abstract GAS improves multimodal understanding by using generation as auxiliary supervision via next embedding prediction and a decoupled mixture-of-transformers architecture, with no inference overhead. Generated by thinkingmachines/Inkling-Small While Multimodal Large Language… 19 Hugging Face Daily Papers research 13d ago SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation Abstract SPARGen unifies 3D reconstruction, dense correspondence, and spatial reasoning into a single instruction-conditioned multimodal generative model that jointly learns shared spatial representations. Generated by thinkingmachines/Inkling-Small Spatial perception and… 31 Hugging Face Daily Papers research 13d ago Self-Supervised Visual On-Policy Distillation Abstract Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers. Generated by thinkingmachines/Inkling-Small Visual on-policy… 5 Hugging Face Daily Papers research 13d ago UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations Abstract UniProbe is a lightweight learnable detector that uses a directed graph and alternating GNN, ViT, and GRU modules to identify hallucinated tokens in frozen large vision-language models, enabling real-time resampling during generation. Generated by… 22 Hugging Face Daily Papers research 13d ago Latent On-Policy Self-Distillation Abstract Latent On-Policy Self-Distillation learns privileged teaching context end-to-end from experience to provide dense token-level supervision, improving agent performance and efficiency. Generated by thinkingmachines/Inkling-Small Enabling agents to learn from experience… 10 Hugging Face Daily Papers research 13d ago MobileMem: Learning from a Year of Mobile Experiences Abstract MobileMem is a benchmark and framework for evaluating on-device long-term memory through year-scale, multimodal mobile experience trajectories that require temporal reasoning, knowledge updating, and preference inference. Generated by thinkingmachines/Inkling-Small The… 5 Hugging Face Daily Papers research 13d ago Multimodal Model Diffing for Feature Discovery and Control Abstract MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors. Generated by thinkingmachines/Inkling-Small Multimodal Large… 20 arXiv — Machine Learning research 13d ago MedMix: Specialization-Consistent Federated Sparse MoEs under Modality Heterogeneity arXiv:2608.13911v1 Announce Type: new Abstract: Federated multimodal medical AI faces modality heterogeneity at both the client and sample levels: clients may systematically lack access to specific modality types, while individual records within the same client may contain… 34 arXiv — Machine Learning research 13d ago Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training arXiv:2608.14498v1 Announce Type: new Abstract: Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current… 8 arXiv — NLP / Computation & Language research 13d ago GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis arXiv:2608.13741v1 Announce Type: new Abstract: Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf… 8 arXiv — NLP / Computation & Language research 13d ago CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA arXiv:2608.13706v1 Announce Type: new Abstract: Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such… 8 arXiv — NLP / Computation & Language research 13d ago Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision arXiv:2608.13854v1 Announce Type: new Abstract: Code translation must preserve executable behavior across many programming languages, yet neural code translation has largely focused on a few popular languages such as C++, Java, and Python. This leaves a niche, many-to-many… 34 arXiv — NLP / Computation & Language research 13d ago S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling arXiv:2608.14029v1 Announce Type: new Abstract: Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level retrieval is… 6 arXiv — NLP / Computation & Language research 13d ago Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge arXiv:2608.14150v1 Announce Type: new Abstract: The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversations: speaker diarization and recognition (Task 1) and conversational speech… 16 arXiv — NLP / Computation & Language research 13d ago Seeing Red, Thinking Bad: Color Bias in Vision Language Models arXiv:2608.14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In… 30 arXiv — NLP / Computation & Language research 13d ago Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding arXiv:2602.01785v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success in source code understanding, yet as software systems grow in scale, computational efficiency has become a critical bottleneck. Currently, these models rely on a… 23 Simon Willison community 13d ago Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things Friday's big release was Qwen 3.8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor Qwen 3.6 27B… 17 TechCrunch — AI news-outlet 13d ago Why people aren’t buying Mark Zuckerberg’s AI future On the latest episode of Equity podcast, we discuss why not everyone is buying Zuckerberg’s vision. 32 r/LocalLLaMA community 13d ago Let's play : guess the model Single prompt "make me a 3d scene of an anime girl" (stolen from someone on discord, I like it a lot since it seems like every model struggles a lot with it) Everything is one shot. No vision, no iterations, no corrections, no going in the browser. Essentially depends on… 33 r/LocalLLaMA community 14d ago The perfect way for Google to screw over OAI and Anthropic is by releasing a 120B dense multimodal Gemma model The two leading labs are already feeling extremely threatened by Qwen & friends, however I think there are tons of enterprises and organizations in the West that don't feel comfortable using Chinese models. I believe these orgs would be all over a near-frontier open-weight model… 35 r/LocalLLaMA community 14d ago A nice local vision test What is the meter reading? It should be 37461. What does your favourite vision model give? (Typo fixed)   submitted by   /u/MrMrsPotts [link]   [comments] 8 llama.cpp releases dev-tools 15d ago b10438: mtmd: fix Granite4 Vision image sequence assembly (#26653) mtmd: fix granite 4v grid assembly (cherry picked from commit 91f82eb ) mtmd: fix truncation for scaled image height and width before unpad Signed-off-by: Hemanth Battu hbattu@ibm.com mtmd: remove MTMD_DUMP_EMBD debug scaffolding Signed-off-by: Hemanth Battu hbattu@ibm.com clean… 23 Hugging Face Daily Papers research 15d ago Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation Abstract Context-Matched Distillation aligns teacher supervision with causal generation context for few-step autoregressive video models, improving control adherence and long-video quality. Generated by thinkingmachines/Inkling-Small Interactive autoregressive video generation… 12 Hugging Face Daily Papers research 16d ago Intern-S2-Preview: Scientific Agentic Foundation Model Abstract Intern-S2-Preview is a scientific agentic foundation model series that integrates multimodal pre-training, multi-task reinforcement learning, and memory-augmented extensions to support long-horizon scientific reasoning and forecasting. Generated by… 33 Hugging Face Daily Papers research 16d ago An AI4AI Framework for Visual Token Pruning Abstract AutoPrune uses large language models to automatically design visual-token pruning policies for multimodal models via a domain-specific language and residual search formulation, achieving high efficiency with minimal performance loss. Generated by… 27 Smol AI News news-outlet 16d ago not much happened today **Z.ai launched GLM-5.3**, a coding- and cyber-focused model with significant gains on agentic and security benchmarks, achieved through scaled post-training rather than a larger base model. **Alibaba released Qwen3.8-27B**, a native multimodal dense model under Apache 2.0 with… 5 Hugging Face Daily Papers research 16d ago Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Abstract Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency. Generated by thinkingmachines/Inkling-Small Interactive world models… 34 arXiv — Machine Learning research 16d ago Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness arXiv:2608.12592v1 Announce Type: new Abstract: Continuous physiological time series underpin modern clinical monitoring, yet many of the most informative signals are invasive, expensive, or simply unavailable for a given patient. Conditional generation offers a remedy: an… 6 arXiv — Machine Learning research 16d ago A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings arXiv:2608.12745v1 Announce Type: new Abstract: Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making… 19 arXiv — Machine Learning research 16d ago Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance arXiv:2608.12929v1 Announce Type: new Abstract: The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. We compared the predictive power of network, environmental, device, and vision… 38 arXiv — Machine Learning research 16d ago CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation arXiv:2608.12944v1 Announce Type: new Abstract: Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the… 23 arXiv — NLP / Computation & Language research 16d ago Vision-Language Models are Fragile Multilingual Associators arXiv:2608.12333v1 Announce Type: new Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark… 37 arXiv — NLP / Computation & Language research 16d ago CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model arXiv:2608.13101v1 Announce Type: new Abstract: Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and… 36 arXiv — NLP / Computation & Language research 16d ago How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures arXiv:2608.13267v1 Announce Type: new Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty… 20 arXiv — NLP / Computation & Language research 16d ago Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety arXiv:2608.13304v1 Announce Type: new Abstract: Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refused. We propose Wrapper-Based Intent-Form… 38 arXiv — NLP / Computation & Language research 16d ago CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation arXiv:2608.13387v1 Announce Type: new Abstract: On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation… 33 arXiv — NLP / Computation & Language research 16d ago Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of Object-Level Anchors arXiv:2608.12746v1 Announce Type: cross Abstract: Object hallucination in multimodal large language models arises when language priors and corpus co-occurrence bias outweigh the visual evidence, with nothing tying an individual object mention to what the image shows. Most… 4 arXiv — NLP / Computation & Language research 16d ago TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint arXiv:2608.13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce… 30 Hugging Face Daily Papers research 16d ago Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Abstract A frozen vision-language model improves spatial reasoning by self-evolving through verified experience, reflection, and reusable memory retrieval without parameter updates or external tools. Generated by thinkingmachines/Inkling-Small Spatial intelligence is becoming a… 30 Hugging Face Daily Papers research 16d ago AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Abstract AutoDesign uses a meta-harness optimizer to recursively improve a code agent for structured media generation, achieving state-of-the-art results on paper-to-poster synthesis. Generated by thinkingmachines/Inkling-Small Transforming multimodal sources into condensed and… 11 Page 5 of 10 · 500 articles ← Newer Older →