✨ Highlights<br><strong>Timestamped Per-vGrid layout</strong> — Audio and video tokens are organized into timestamped grid with explicit boundaries, keeping audio segments adjacent to their corresponding visual content for fine-grained temporal alignment over long streams.<br><strong>Three-stage SFT recipe</strong> — Progressive training from audio-language alignment to full multimodal SFT, developing live-commerce understanding from omni-modal perception to instruction-following responses.<br><strong>Faithful-RFT</strong> — A reinforcement fine-tuning stage for faithful and real-time live-stream demands, suppressing explicit reasoning traces and directly optimizing answer quality for live-commerce tasks.<br><strong>Rich atomic capabilities</strong> — A scenario-oriented taxonomy covering speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc, supported by a compact data production engine.<br><strong>Strong live-commerce performance with competitive generalization</strong> — 4B and 9B variants demonstrate strong results across live-commerce audio, image, and video tasks, together with excellent generalization on general benchmarks.</p>\n<p>📦 Resources & Code: <a href=\"https://github.com/TaoLiveAIGC/TLive-Omni\" rel=\"nofollow\">https://github.com/TaoLiveAIGC/TLive-Omni</a></p>\n","updatedAt":"2026-08-25T02:15:11.982Z","author":{"_id":"65d46326492611d68f33b88b","avatarUrl":"/avatars/19f18d675d2eff03980e78aeca9cacf7.svg","fullname":"Leon","name":"Leon1207","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8189127445220947},"editors":["Leon1207"],"editorAvatarUrls":["/avatars/19f18d675d2eff03980e78aeca9cacf7.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.20958","authors":[{"_id":"6a8baab23d26296ea30918fc","user":{"_id":"6445e0d57a7b94ddc2d6f8e6","avatarUrl":"/avatars/0607c2be23a143fa1afcb07251b26756.svg","isPro":false,"fullname":"Aber Hu","user":"Aber-r","type":"user","name":"Aber-r"},"name":"Yibo Hu","status":"claimed_verified","statusLastChangedAt":"2026-08-24T08:14:22.379Z","hidden":false},{"_id":"6a8baab23d26296ea30918fd","name":"Yu Qian","hidden":false},{"_id":"6a8baab23d26296ea30918fe","name":"Mao Gu","hidden":false},{"_id":"6a8baab23d26296ea30918ff","name":"Yingfan Tao","hidden":false},{"_id":"6a8baab23d26296ea3091900","name":"Yuhao Chen","hidden":false},{"_id":"6a8baab23d26296ea3091901","user":{"_id":"65d46326492611d68f33b88b","avatarUrl":"/avatars/19f18d675d2eff03980e78aeca9cacf7.svg","isPro":false,"fullname":"Leon","user":"Leon1207","type":"user","name":"Leon1207"},"name":"Yongdong Luo","status":"claimed_verified","statusLastChangedAt":"2026-08-24T08:14:25.127Z","hidden":false},{"_id":"6a8baab23d26296ea3091902","name":"Zhuoqun Liu","hidden":false},{"_id":"6a8baab23d26296ea3091903","name":"Meiguang Jin","hidden":false},{"_id":"6a8baab23d26296ea3091904","name":"Junfeng Ma","hidden":false}],"publishedAt":"2026-08-21T00:00:00.000Z","submittedOnDailyAt":"2026-08-25T00:00:00.000Z","title":"TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming","submittedOnDailyBy":{"_id":"65d46326492611d68f33b88b","avatarUrl":"/avatars/19f18d675d2eff03980e78aeca9cacf7.svg","isPro":false,"fullname":"Leon","user":"Leon1207","type":"user","name":"Leon1207"},"summary":"E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.","upvotes":28,"discussionId":"6a8baab23d26296ea3091905","githubRepo":"https://github.com/TaoLiveAIGC/TLive-Omni","githubRepoAddedBy":"user","ai_summary":"TLive-Omni is an omni-modal model for live-commerce that unifies image, video, audio, and text via timestamped token grouping, staged supervised training, and reinforcement fine-tuning with verifiable feedback to enable accurate real-time understanding.","ai_keywords":["omni-modal understanding","Per-vGrid","timestamped token organization","cross-modal alignment","three-stage supervised training","Faithful-RFT","reinforcement fine-tuning","GRPO","dynamic sampling","length-grouped sampler","live-commerce","temporal grounding","product visual grounding"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":15,"organization":{"_id":"6a6205cbc80c71fffa0d91c0","name":"TaoLiveAIGC","fullname":"TaoLive AIGC","avatar":"https://www.gravatar.com/avatar/6ad8c78ff1fe4f98ab8ae0c6fe81f1e4?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65d46326492611d68f33b88b","avatarUrl":"/avatars/19f18d675d2eff03980e78aeca9cacf7.svg","isPro":false,"fullname":"Leon","user":"Leon1207","type":"user"},{"_id":"64f5884336a565f2319059ce","avatarUrl":"/avatars/8fe800ce098b5b133c16ca7641ce190c.svg","isPro":false,"fullname":"GMadeus","user":"AmadeusL","type":"user"},{"_id":"64e57b65ac2899e39ad59acc","avatarUrl":"/avatars/e73db207003c9be2041bc6f14ef514fa.svg","isPro":false,"fullname":"star_rising","user":"starrising","type":"user"},{"_id":"6406d8f4d684369027164e41","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406d8f4d684369027164e41/vHRepecrzKtxdm0q52lVd.jpeg","isPro":false,"fullname":"sunyuhan","user":"yuuhan","type":"user"},{"_id":"65c0cee5bdfafce750bc92d8","avatarUrl":"/avatars/cca0752a7dc8958f120be93b14bf7c13.svg","isPro":false,"fullname":"Akane Senri","user":"SenriAkane","type":"user"},{"_id":"65d4ebdbde187cd7fceb9669","avatarUrl":"/avatars/05111bc64ff13fb3b137026e57e6815a.svg","isPro":false,"fullname":"WanqingCui","user":"VickiCui","type":"user"},{"_id":"6445e0d57a7b94ddc2d6f8e6","avatarUrl":"/avatars/0607c2be23a143fa1afcb07251b26756.svg","isPro":false,"fullname":"Aber Hu","user":"Aber-r","type":"user"},{"_id":"67e294e37210beea5ee1f6b0","avatarUrl":"/avatars/cddf8bb0aa89777d027d8dcc410b6bc9.svg","isPro":false,"fullname":"Chen","user":"dlyldxwl","type":"user"},{"_id":"64af8f33c313684399b90fa2","avatarUrl":"/avatars/b20165cfa600579071f604a362d9c57d.svg","isPro":false,"fullname":"Taoyingfan","user":"helloworldtvt","type":"user"},{"_id":"65b74a8fe2510e886da95b1e","avatarUrl":"/avatars/f95cd55083a23fa8d2c1fb98709087d8.svg","isPro":false,"fullname":"Liu Zhuoqun","user":"lzq-971128","type":"user"},{"_id":"65571120c7c44ce359045a10","avatarUrl":"/avatars/226040c15c506317234817a87c461b04.svg","isPro":false,"fullname":"zivkidd","user":"zivkidd","type":"user"},{"_id":"6798383802ff123f680538ff","avatarUrl":"/avatars/682e16fe5b976d3b1c624ad95fa717c4.svg","isPro":false,"fullname":"Wang Weishu","user":"lohr","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a6205cbc80c71fffa0d91c0","name":"TaoLiveAIGC","fullname":"TaoLive AIGC","avatar":"https://www.gravatar.com/avatar/6ad8c78ff1fe4f98ab8ae0c6fe81f1e4?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.20958.md","query":{}}">
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Published on Aug 21
· Submitted by Leon on Aug 25 Abstract
TLive-Omni is an omni-modal model for live-commerce that unifies image, video, audio, and text via timestamped token grouping, staged supervised training, and reinforcement fine-tuning with verifiable feedback to enable accurate real-time understanding.
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.
Community
✨ Highlights
Timestamped Per-vGrid layout — Audio and video tokens are organized into timestamped grid with explicit boundaries, keeping audio segments adjacent to their corresponding visual content for fine-grained temporal alignment over long streams.
Three-stage SFT recipe — Progressive training from audio-language alignment to full multimodal SFT, developing live-commerce understanding from omni-modal perception to instruction-following responses.
Faithful-RFT — A reinforcement fine-tuning stage for faithful and real-time live-stream demands, suppressing explicit reasoning traces and directly optimizing answer quality for live-commerce tasks.
Rich atomic capabilities — A scenario-oriented taxonomy covering speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc, supported by a compact data production engine.
Strong live-commerce performance with competitive generalization — 4B and 9B variants demonstrate strong results across live-commerce audio, image, and video tasks, together with excellent generalization on general benchmarks.
📦 Resources & Code: https://github.com/TaoLiveAIGC/TLive-Omni
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.20958 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.