PatternEval reveals widespread response-pattern misalignment between thinking and non-thinking modes in hybrid-thinking MLLMs, with non-thinking inference exhibiting substantially more failures such as chain-of-thought leakage, repetition, contradiction, and performative reasoning. PatternRL mitigates this cross-mode misalignment by incorporating pattern-specific penalties during reinforcement learning, with only a marginal trade-off in task performance.</p>\n","updatedAt":"2026-08-24T01:57:27.143Z","author":{"_id":"630716d11801ecc7d2595021","avatarUrl":"/avatars/2d36a880ce4a3cf7efc5ff3987dbeaf3.svg","fullname":"Songyang Zhang","name":"zsytony","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":29,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9326317310333252},"editors":["zsytony"],"editorAvatarUrls":["/avatars/2d36a880ce4a3cf7efc5ff3987dbeaf3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.12781","authors":[{"_id":"6a8ba4c53d26296ea30918d3","name":"Xinming Wang","hidden":false},{"_id":"6a8ba4c53d26296ea30918d4","name":"Weinong Wang","hidden":false},{"_id":"6a8ba4c53d26296ea30918d5","name":"Hongming Yang","hidden":false},{"_id":"6a8ba4c53d26296ea30918d6","name":"Yansong Lin","hidden":false},{"_id":"6a8ba4c53d26296ea30918d7","name":"Zheng Ruan","hidden":false},{"_id":"6a8ba4c53d26296ea30918d8","name":"Shangpin Peng","hidden":false},{"_id":"6a8ba4c53d26296ea30918d9","name":"Qiming Peng","hidden":false},{"_id":"6a8ba4c53d26296ea30918da","name":"Nan Qiao","hidden":false},{"_id":"6a8ba4c53d26296ea30918db","name":"Fengyuan Lu","hidden":false},{"_id":"6a8ba4c53d26296ea30918dc","name":"Guoqing Ma","hidden":false},{"_id":"6a8ba4c53d26296ea30918dd","name":"Marito Li","hidden":false},{"_id":"6a8ba4c53d26296ea30918de","name":"Songyang Zhang","hidden":false},{"_id":"6a8ba4c53d26296ea30918df","name":"Saiyong Yang","hidden":false},{"_id":"6a8ba4c53d26296ea30918e0","name":"Han Hu","hidden":false},{"_id":"6a8ba4c53d26296ea30918e1","name":"Yonglong Tian","hidden":false},{"_id":"6a8ba4c53d26296ea30918e2","name":"Xu-Yao Zhang","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-24T00:00:00.000Z","title":"Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs","submittedOnDailyBy":{"_id":"630716d11801ecc7d2595021","avatarUrl":"/avatars/2d36a880ce4a3cf7efc5ff3987dbeaf3.svg","isPro":false,"fullname":"Songyang Zhang","user":"zsytony","type":"user","name":"zsytony"},"summary":"Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through response-pattern alignment: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce PatternEval, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop PatternRM, a response-level reward model, and PatternRL, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.","upvotes":5,"discussionId":"6a8ba4c63d26296ea30918e3","ai_summary":"Hybrid-thinking multimodal language models suffer from response-pattern misalignment between thinking and non-thinking modes, which is addressed by a diagnostic benchmark and pattern-specific reinforcement learning penalties.","ai_keywords":["multimodal large language models","hybrid-thinking","response-pattern alignment","PatternEval","chain-of-thought leakage","logical contradiction","performative reasoning","PatternRM","PatternRL","reinforcement learning"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"630716d11801ecc7d2595021","avatarUrl":"/avatars/2d36a880ce4a3cf7efc5ff3987dbeaf3.svg","isPro":false,"fullname":"Songyang Zhang","user":"zsytony","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"63c1699e40a26dd2db32400d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c1699e40a26dd2db32400d/3N0-Zp8igv8-52mXAdiiq.jpeg","isPro":false,"fullname":"Chroma","user":"Chroma111","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"69e5a48215e887b7b6624512","avatarUrl":"/avatars/99183659f6927cde86826c9d0380a6d5.svg","isPro":false,"fullname":"zi","user":"TAOTAO777","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.12781.md","query":{}}">
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Abstract
Hybrid-thinking multimodal language models suffer from response-pattern misalignment between thinking and non-thinking modes, which is addressed by a diagnostic benchmark and pattern-specific reinforcement learning penalties.
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through response-pattern alignment: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce PatternEval, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop PatternRM, a response-level reward model, and PatternRL, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.
Community
PatternEval reveals widespread response-pattern misalignment between thinking and non-thinking modes in hybrid-thinking MLLMs, with non-thinking inference exhibiting substantially more failures such as chain-of-thought leakage, repetition, contradiction, and performative reasoning. PatternRL mitigates this cross-mode misalignment by incorporating pattern-specific penalties during reinforcement learning, with only a marginal trade-off in task performance.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.12781 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.12781 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.12781 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.