We introduce MMDiff, a multimodal model-diffing pipeline to discover task-specific features in MLLMs and enable targeted feature-level control.</p>\n<p>🌟 MMDiff diffs a base-LM SAE against a multimodal SAE to isolate vision-adapted features, making it easier to distinguish visually responsive features from representations inherited from the base language model.</p>\n<p>🌟 We then identify task-specific features within these vision-adapted features, enabling targeted feature-level steering and ablation of specific model behaviours.</p>\n<p>🌟 We apply MMDiff across three MLLM families with different language backbones, vision encoders, and SAE objectives, showing that the pipeline generalizes across LLaVA-MORE, PaliGemma 2, and InternVL3.5.</p>\n<p>🌟 We demonstrate targeted control across spatial reasoning, multimodal safety, and OCR, where targeted ablations selectively affect the corresponding behaviours while largely preserving general VQA performance.</p>\n<p>🌟 MMDiff provides a practical route to discovering, understanding, and controlling task-specific representations in multimodal models.</p>\n<p>Project Page: <a href=\"https://pixl.cs.ox.ac.uk/mmdiff/\" rel=\"nofollow\">https://pixl.cs.ox.ac.uk/mmdiff/</a><br>arXiv: <a href=\"https://arxiv.org/abs/2608.09928\" rel=\"nofollow\">https://arxiv.org/abs/2608.09928</a></p>\n","updatedAt":"2026-08-17T04:04:50.986Z","author":{"_id":"62f5c24eea5bd6b1abc8e151","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1660273191881-noauth.jpeg","fullname":"Hunar Batra","name":"hunarbatra","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8482587337493896},"editors":["hunarbatra"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1660273191881-noauth.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09928","authors":[{"_id":"6a7ad372019ce76dc7b3abca","user":{"_id":"62f5c24eea5bd6b1abc8e151","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1660273191881-noauth.jpeg","isPro":true,"fullname":"Hunar Batra","user":"hunarbatra","type":"user","name":"hunarbatra"},"name":"Hunar Batra","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.582Z","hidden":false},{"_id":"6a7ad372019ce76dc7b3abcb","name":"Lachin Naghashyar","hidden":false},{"_id":"6a7ad372019ce76dc7b3abcc","name":"Ashkan Khakzar","hidden":false},{"_id":"6a7ad372019ce76dc7b3abcd","name":"Philip Torr","hidden":false},{"_id":"6a7ad372019ce76dc7b3abce","name":"Christian Schroeder de Witt","hidden":false},{"_id":"6a7ad372019ce76dc7b3abcf","name":"Constantin Venhoff","hidden":false},{"_id":"6a7ad372019ce76dc7b3abd0","name":"Ronald Clark","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/62f5c24eea5bd6b1abc8e151/U9Slp3_JLGDYI2cxN1CKG.gif"],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"Multimodal Model Diffing for Feature Discovery and Control","submittedOnDailyBy":{"_id":"62f5c24eea5bd6b1abc8e151","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1660273191881-noauth.jpeg","isPro":true,"fullname":"Hunar Batra","user":"hunarbatra","type":"user","name":"hunarbatra"},"summary":"Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.","upvotes":3,"discussionId":"6a7ad372019ce76dc7b3abd1","projectPage":"https://pixl.cs.ox.ac.uk/mmdiff","githubRepo":"https://github.com/hunarbatra/MMDiff","githubRepoAddedBy":"user","ai_summary":"MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.","ai_keywords":["multimodal large language models","sparse autoencoders","MMDiff","feature isolation","contrastive firing analysis","feature-level control","multimodal SAEs","visual-spatial understanding","multimodal safety","steering"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"627bbc28fbab61b048eba8b6","name":"Oxford","fullname":"University of Oxford","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/u0ey2LfYu6uG6iu8m_kH7.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62f5c24eea5bd6b1abc8e151","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1660273191881-noauth.jpeg","isPro":true,"fullname":"Hunar Batra","user":"hunarbatra","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"662edf5f04f9341b56fa8a81","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/AucZp4zy0w7c7EjtVYujE.jpeg","isPro":false,"fullname":"Ananay arora","user":"ananayarora","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"627bbc28fbab61b048eba8b6","name":"Oxford","fullname":"University of Oxford","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/u0ey2LfYu6uG6iu8m_kH7.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09928.md","query":{}}">
Multimodal Model Diffing for Feature Discovery and Control
Abstract
MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
Community
We introduce MMDiff, a multimodal model-diffing pipeline to discover task-specific features in MLLMs and enable targeted feature-level control.
🌟 MMDiff diffs a base-LM SAE against a multimodal SAE to isolate vision-adapted features, making it easier to distinguish visually responsive features from representations inherited from the base language model.
🌟 We then identify task-specific features within these vision-adapted features, enabling targeted feature-level steering and ablation of specific model behaviours.
🌟 We apply MMDiff across three MLLM families with different language backbones, vision encoders, and SAE objectives, showing that the pipeline generalizes across LLaVA-MORE, PaliGemma 2, and InternVL3.5.
🌟 We demonstrate targeted control across spatial reasoning, multimodal safety, and OCR, where targeted ablations selectively affect the corresponding behaviours while largely preserving general VQA performance.
🌟 MMDiff provides a practical route to discovering, understanding, and controlling task-specific representations in multimodal models.
Project Page: https://pixl.cs.ox.ac.uk/mmdiff/
arXiv: https://arxiv.org/abs/2608.09928
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.09928 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.09928 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.09928 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.