News / #multimodal Tag Multimodal 500 articles archived under #multimodal · RSS Sign in to follow Hugging Face official-blog 16d ago State of Open Models: Summer 2026 Observations Back to Articles a]:hidden"> State of Open Models: Summer 2026 Observations Published August 14, 2026 Update on GitHub Upvote 2 Adina Yakefu AdinaY Apolinário from multimodal AI art multimodalart Irene Solaiman irenesolaiman In the AI world, time feels compressed. A few months… 28 r/LocalLLaMA community 16d ago SenseNova-Vision: a 7B open model that does segmentation, depth, detection, OCR, and 3D reconstruction with no task-specific heads Stumbled across this new vision model, SenseNova-Vision. It's a 7B MoT model, Apache 2.0 license, which is cool. The main idea is it treats pretty much all computer vision stuff as just one generation problem. Like, instead of needing a bunch of different models for detection,… 17 r/LocalLLaMA community 17d ago I asked DeepSeek-V4-Flash to work with Muse-Glimmer for Vision ability in PI agent and it produced this Same old prompt, just appended a TIP in the end: "Write a single HTML file with a full-page canvas and no libraries. Simulate a realistic side-view of a moving car as the main subject. Keep the car visible in the foreground while the background landscape scrolls continuously to… 33 Hugging Face Daily Papers research 17d ago AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models Abstract AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera. Generated by thinkingmachines/Inkling-Small While Vision-Language-Action (VLA)… 36 arXiv — Machine Learning research 17d ago FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting arXiv:2608.11623v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational… 23 arXiv — NLP / Computation & Language research 17d ago LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection arXiv:2608.11691v1 Announce Type: cross Abstract: Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a… 18 arXiv — Machine Learning research 17d ago REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation arXiv:2608.11698v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond… 8 arXiv — Machine Learning research 17d ago Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision arXiv:2608.12027v1 Announce Type: new Abstract: Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing… 13 arXiv — NLP / Computation & Language research 17d ago CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that… 10 arXiv — NLP / Computation & Language research 17d ago AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention arXiv:2608.11758v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong generalization and reasoning abilities due to large-scale multimodal pre-training. However, fine-tuning these models on downstream tasks often leads to catastrophic… 34 arXiv — NLP / Computation & Language research 17d ago GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation arXiv:2608.11787v1 Announce Type: new Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision… 6 arXiv — NLP / Computation & Language research 17d ago BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model arXiv:2608.11244v1 Announce Type: cross Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support… 9 arXiv — NLP / Computation & Language research 17d ago How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment arXiv:2608.11816v1 Announce Type: cross Abstract: State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced… 25 arXiv — NLP / Computation & Language research 17d ago LookBack: Where and How to Score LVLM Responses via Visual Reference Usage arXiv:2608.11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations;… 14 arXiv — NLP / Computation & Language research 17d ago Investigating Learner-Aware Design of LLM-Generated Educational Feedback arXiv:2602.11650v2 Announce Type: replace Abstract: Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information coverage) to support answer revision and learner acceptance… 12 Hugging Face Daily Papers research 17d ago Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models Abstract Self-Geometry improves vision foundation model predictions by enforcing explicit multi-view geometric constraints via test-time adaptation with LoRA, disentangled losses, and angular neighbor sampling. Generated by thinkingmachines/Inkling-Small Recent Vision Foundation… 21 Hugging Face Daily Papers research 17d ago The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Abstract Visual tool-use in multimodal LLMs often lacks causal effectiveness, with returned observations frequently failing to influence answers or being used incoherently despite aggregate accuracy improvements. Generated by thinkingmachines/Inkling-Small The… 6 Hugging Face Daily Papers research 17d ago NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs Abstract NeuPAT selectively constrains updates to language-sensitive neurons during multimodal tuning to preserve LLM language capabilities while enabling perceptual adaptation. Generated by thinkingmachines/Inkling-Small Multimodal expansion of large language models (LLMs)… 18 Hugging Face Daily Papers research 17d ago MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only… 21 Hugging Face Daily Papers research 17d ago Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Abstract Spark-to-Paper is a lightweight, composable workflow inside coding assistants that generates research papers by separating planning from reporting, enforcing evidence-based claim revision, and using integrity checks to reduce fabrication. Generated by… 9 r/LocalLLaMA community 17d ago LFM2.5-VL-3B recognizes Steve from Minecraft running locally on an iPhone 17 Liquid AI put out LFM2.5-VL-3B today, which is a 3.1B vision model that weighs roughly 2GB and fits well on a phone Benchmarks are benchmarks so I tried something sillier. Took a photo of a little Steve toy I have, gave it to the model and asked it what it was looking at It… 29 r/LocalLLaMA community 17d ago CohereLabs/North-Micro-Vision-Instruct · Hugging Face North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation for prototyping, task-specific fine-tuning, and specialized multimodal… 8 r/LocalLLaMA community 17d ago LiquidAI/LFM2.5-VL-3B · Hugging Face LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment . It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both text and images, and uses the LFM2.5-2.6B language model as its backbone,… 15 Hugging Face official-blog 18d ago LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge Back to Articles a]:hidden"> LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge Team Article Published August 12, 2026 Upvote - Samuel Stevens samuelstevens LiquidAI Ryan Shubert shubeydoo LiquidAI Sina s-jse LiquidAI Tianshu Yu tianshu-yu LiquidAI Brandon… 24 Hugging Face Daily Papers research 18d ago Articulated Object Reconstruction from Rest-State Observation Abstract A rest-state framework reconstructs articulated objects from a single closed configuration by fusing vision-language outputs into consistent part meshes and validating synthesized motion hypotheses via geometric consistency. Generated by thinkingmachines/Inkling-Small… 29 Hugging Face Daily Papers research 18d ago DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Abstract DistilVDR is a compact 524M vision-document retriever distilled from an 8B teacher using cosine alignment without relevance labels, achieving near-teacher accuracy with far smaller indexes and faster indexing. Generated by thinkingmachines/Inkling-Small Visual document… 33 arXiv — Machine Learning research 18d ago Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory arXiv:2608.09997v1 Announce Type: new Abstract: Transformers have had a profound impact on the world of language processing and computer vision. As efforts to answer the million-dollar question of ``How does a Transformer learn?" have been increasing, existing interpretability… 6 arXiv — Machine Learning research 18d ago ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation arXiv:2608.10905v1 Announce Type: new Abstract: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight,… 7 arXiv — Machine Learning research 18d ago MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis arXiv:2608.09986v1 Announce Type: cross Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although… 29 arXiv — NLP / Computation & Language research 18d ago Multimodal Item Parameter Estimation using Simulated Response Probabilitie arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to… 19 arXiv — NLP / Computation & Language research 18d ago VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise… 14 arXiv — NLP / Computation & Language research 18d ago Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models arXiv:2608.10484v1 Announce Type: cross Abstract: Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2… 21 arXiv — NLP / Computation & Language research 18d ago The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces arXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision… 28 arXiv — NLP / Computation & Language research 18d ago Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence arXiv:2608.10720v1 Announce Type: cross Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a… 22 arXiv — NLP / Computation & Language research 18d ago StreamFlow: Dynamic Memory Flows for Streaming Video Understanding arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded memory. Yet existing paradigms remain limited:… 18 arXiv — NLP / Computation & Language research 18d ago MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment… 16 arXiv — NLP / Computation & Language research 18d ago HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models arXiv:2506.03922v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical… 15 arXiv — NLP / Computation & Language research 18d ago Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pretraining priors. For example, models will confidently assert a five-legged dog has four legs; consequently, on the VLMBias… 13 Hugging Face Daily Papers research 18d ago JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles Abstract A new jigsaw benchmark with interlocking pieces reveals that vision-language models fail at geometric reasoning and suffer a sharp performance drop as puzzle size increases. Generated by thinkingmachines/Inkling-Small Jigsaw puzzle solving requires jointly reasoning… 7 Vercel — AI dev-tools 18d ago Set up coding agents in one command with AI Gateway Using coding agents means setting up multiple accounts, provisioning API keys, and scattering observability and billing. Now, you can route them through AI Gateway to centralize all of this and add controls, with set up in one command: Any of 200+ models in any agent , including… 29 Hugging Face Daily Papers research 18d ago On-Policy Self-Distillation without Any Supervision Abstract Unsupervised on-policy self-distillation improves large language models by using internal consistency and majority-vote pseudo-solutions to correct confident errors without external supervision. Generated by thinkingmachines/Inkling-Small On-policy (Self-)Distillation… 34 TechCrunch — AI news-outlet 18d ago General Catalyst leads $1.1B round into 2-month-old River AI River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate. 5 r/LocalLLaMA community 19d ago Revision Prompting: Trades slow (decoded) output tokens for cheap (prefilled) input tokens. TL;DR: If you re-run the same prompt whenever the input changes, try sending the old input/output plus a diff of the input, and ask the model for a patch to the output. You generate ~2-10x fewer output tokens, and the untouched parts of the output stay byte-identical. This… 22 Hugging Face Daily Papers research 19d ago MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Abstract MMOOC is a large-scale benchmark assessing whether multimodal language models can correctly refuse out-of-context questions while answering shifted in-context questions, revealing that current models struggle to balance these abilities. Generated by… 37 Hugging Face Daily Papers research 19d ago Vision-Language Grounding as Bidirectional Concept Correspondence Abstract ConCor-1 treats vision-language grounding as bidirectional concept correspondence, jointly predicting text spans, image segments, and cross-modal matches without prespecified phrases. Generated by thinkingmachines/Inkling-Small Vision-language grounding connects… 30 Hugging Face Daily Papers research 19d ago What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems Abstract A three-stage multimodal framework improves follow-up edit recommendations in image-creation conversations by combining supervised fine-tuning, multi-objective reinforcement learning, and visual verification. Generated by thinkingmachines/Inkling-Small Conversational… 22 arXiv — Machine Learning research 19d ago CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion Recognition arXiv:2608.07867v1 Announce Type: new Abstract: Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal conflict. We propose \textbf{CONFER}, a graph-based conflict-aware evidence… 9 arXiv — Machine Learning research 19d ago Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees arXiv:2608.08002v1 Announce Type: new Abstract: Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of… 25 arXiv — Machine Learning research 19d ago The Neural Division of Labor: Biologically-Inspired Modular Architectures for Robust Neuromorphic Computing arXiv:2608.08317v1 Announce Type: new Abstract: Biological neural systems achieve high efficiency and robustness through compartmentalized architectures. In contrast, modern artificial neural networks rely on globally entangled structures, which obscure decision logic and suffer… 7 arXiv — Machine Learning research 19d ago Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles arXiv:2608.08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable adversarial attacks, including shadow perturbations, natural-light… 15 Page 6 of 10 · 500 articles ← Newer Older →