Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S2VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S2VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S2VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information.</p>\n","updatedAt":"2026-08-17T07:52:16.890Z","author":{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","fullname":"William Li","name":"williamium","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8745443820953369},"editors":["williamium"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.14144","authors":[{"_id":"6a82bd85b601d59c652815a0","name":"Yijiang Li","hidden":false},{"_id":"6a82bd85b601d59c652815a1","name":"Yijun Liang","hidden":false},{"_id":"6a82bd85b601d59c652815a2","name":"Yunjie Tian","hidden":false},{"_id":"6a82bd85b601d59c652815a3","name":"Bingyang Wang","hidden":false},{"_id":"6a82bd85b601d59c652815a4","name":"Ke Zhang","hidden":false},{"_id":"6a82bd85b601d59c652815a5","name":"Zhenfei Yin","hidden":false},{"_id":"6a82bd85b601d59c652815a6","name":"Di Fu","hidden":false},{"_id":"6a82bd85b601d59c652815a7","name":"Philip Torr","hidden":false},{"_id":"6a82bd85b601d59c652815a8","name":"Nuno Vasconcelos","hidden":false}],"publishedAt":"2026-08-14T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"Self-Supervised Visual On-Policy Distillation","submittedOnDailyBy":{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","isPro":true,"fullname":"William Li","user":"williamium","type":"user","name":"williamium"},"summary":"Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S^2VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S^2VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S^2VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at https://williamium3000.github.io/s2vopd","upvotes":31,"discussionId":"6a82bd85b601d59c652815a9","projectPage":"https://williamium3000.github.io/s2vopd/","ai_summary":"Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers.","ai_keywords":["visual on-policy distillation","teacher-student asymmetry","privileged supervision","Self-Supervised Visual On-Policy Distillation","S²VOPD","asymmetric augmented views","self-distillation","fine-grained perception"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","isPro":true,"fullname":"William Li","user":"williamium","type":"user"},{"_id":"670f632944d497ffc693129f","avatarUrl":"/avatars/371406b9f53c2edc6271e38a753a889c.svg","isPro":false,"fullname":"Zongyang QIU","user":"Zane-QIU","type":"user"},{"_id":"657858d0dddc2360b01643a1","avatarUrl":"/avatars/00fb2410aaa74e220d69b15084dfebd6.svg","isPro":false,"fullname":"l","user":"gouerrrr","type":"user"},{"_id":"67199be1d1962b005dc59086","avatarUrl":"/avatars/5a6571326dfbbcf86e1863d6732f8aa4.svg","isPro":false,"fullname":"DC","user":"Passenger555","type":"user"},{"_id":"6a6b9f11f2a4d13387747a02","avatarUrl":"/avatars/4b86c0b1d5f90922c4b2ab9c2ac74fc6.svg","isPro":false,"fullname":"ruiyangwang","user":"rywang-thu","type":"user"},{"_id":"634cfebc350bcee9bed20a4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634cfebc350bcee9bed20a4d/fN47nN5rhw-HJaFLBZWQy.png","isPro":false,"fullname":"Xingyi Yang","user":"adamdad","type":"user"},{"_id":"67312401d433c6b122c38202","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67312401d433c6b122c38202/3VxusGGchOJTk9bEPfxSl.jpeg","isPro":false,"fullname":"Yipeng Gao","user":"YipengGao","type":"user"},{"_id":"641e04c28f9507b613ed1186","avatarUrl":"/avatars/8ca9272e5c829bac7d60f710b01296f3.svg","isPro":false,"fullname":"Chengjun Zhang","user":"Raina23278","type":"user"},{"_id":"68db9b2dadc11a1b70408123","avatarUrl":"/avatars/428b4ba13ee57c598a786d9421e68a10.svg","isPro":true,"fullname":"Ruoliu Yang","user":"RuoliuYang","type":"user"},{"_id":"694e6186c717ebb1689a274e","avatarUrl":"/avatars/53c8f9d4d4ac5068c466469c24095787.svg","isPro":true,"fullname":"Logan Yang","user":"q1716523669","type":"user"},{"_id":"66572d970c9058052fbc98b0","avatarUrl":"/avatars/0da4a91ea7c36b7109afe19f193961bf.svg","isPro":false,"fullname":"ZX","user":"Yezixiao","type":"user"},{"_id":"69f90c93649ddb48e25d7d4b","avatarUrl":"/avatars/71e6f2b60059aac3700154cf8dc3fd38.svg","isPro":false,"fullname":"anony-review-prism-2026","user":"PRISM-Bench","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.14144.md","query":{}}">
Self-Supervised Visual On-Policy Distillation
Abstract
Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers.
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S^2VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S^2VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S^2VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at https://williamium3000.github.io/s2vopd
Community
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S2VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S2VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S2VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.14144 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.14144 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.14144 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.