| FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation Overview FireRedAudio is a general-purpose audio language model built on a shared 9B-parameter LLM with decoupled continuous representations: an Audio Encoder handles understanding, while a RedAE pathway handles generation. A single model supports ASR, audio understanding, zero-shot TTS, instruct TTS, semantic/acoustic speech editing, and accurate temporal grounding over recordings up to one hour long. Highlights ✨ - 🧩 Purpose-built representations, one shared backbone — The Audio Encoder pathway serves understanding, while the RedAE-Patch pathway serves speech generation. Their representations remain decoupled but share the same language and reasoning backbone. To the best of our knowledge, this is the first publicly disclosed design of its kind in a unified audio-language model.
- 📊 One model, a full audio stack — FireRedAudio spans ASR, broad and fine-grained audio understanding, zero-shot TTS, Instruct TTS, and free-form speech editing, achieving competitive or leading results across MMAU, MMSU, Seed-TTS-Eval, InstructTTSEval, and Ming-Freeform-Audio-Edit.
- 🎙️ Create and edit speech with natural language — Clone a voice from a reference clip, design a voice from a description, or edit what was said and how it sounds through one continuous-latent generation pathway.
- ⏱️ Go from minutes to hour-long recordings — Understand recordings up to one hour with precise time-to-content alignment. Organize audio into timestamped structures, produce grounded summaries, retrieve content by time (or time by content), and reason over evidence distributed across the recording.
FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations Overview FireRedTTS3 is a unified speech generation and editing system built on semantically enriched continuous speech representations. It comes in two variants: - FireRedTTS3-Base — zero-shot voice cloning across 24 languages and 21 Chinese dialects
- FireRedTTS3-Instruct — natural-language voice design and speech editing (semantic + acoustic) in one unified model
Highlights ✨ - 🌍 Multilingual — 24 Languages — Best average WER/CER (avg 3.754%) and best average speaker similarity on MiniMax-MLS-Test (avg 84.8%), plus best-in-class cloning WER/CER (avg 3.04%) and similarity on Seed-TTS-eval (avg 78.8%). Supported languages:
Arabic · Cantonese · Chinese · Czech · Dutch · English · Finnish · French · German · Greek · Hindi · Indonesian · Italian · Japanese · Korean · Polish · Portuguese · Romanian · Russian · Spanish · Thai · Turkish · Ukrainian · Vietnamese - 🗣️ Multi-Dialect — 21 Chinese Dialects — Zero-shot voice cloning across major Chinese dialect groups. Supported dialects:
Anhui · Fujian · Gansu · Guizhou · Hebei · Henan · Hubei · Hunan · Jiangxi · Liaoning · Minnan · Ningxia · Shaanxi · Shandong · Shanghai · Shanxi · Sichuan · Tianjin · Wenzhou · Wu · Yunnan - 🎨 Instruction-Controlled Voice Design — Generate a brand-new voice from a natural-language description (gender, age, timbre, emotion, pace, accent…) with no reference audio, guided by an explicit textual plainning step before synthesis.
- ✂️ Free-Form Speech Editing — Semantic editing (insertion / deletion / substitution) and acoustic editing (speed / pitch / volume) driven by free-form instructions.
Project : https://fireredteam.github.io/ Their Opensource Projects: Really worth to check the project page. Nice Ecosystem. They also published some papers. - OpenStoryline: An Agentic Framework for Autonomous, Human-Aligned Video Creation
- FireRedChat: A Fully Self-Hosted Solution for Full-Duplex Voice Interaction
- IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
- FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
- InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
- InstantID: Zero-shot Identity-Preserving Generation in Seconds
- DynamicPose: A Robust Image-to-Video Framework for Portrait Animation Driven by Pose Sequences
- PhotoPoster: A High-Fidelity Two-Stage Pose-Driven Image Generation Framework
- CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
- FireRedASR: Open-Source Industrial-Grade Automatic Speech Recognition Models
- FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
- The Xiaohongshu Speech Synthesis System for Blizzard Challenge 2023
- StoryMaker: Towards Consistent Characters in Text-to-Image Generation
- LayerDiffuse-Flux
- InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
submitted by /u/pmttyji [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.