News / #image-gen Tag Image Gen 176 articles archived under #image-gen · RSS Sign in to follow Smol AI News news-outlet 3mo ago not much happened today **RAEv2** advances representation-first tokenization with **>10x faster convergence** and improved generation, tested on **text-to-image** and **world models**. **NVIDIA's Gated DeltaNet-2** innovates linear attention with channel-wise gates, outperforming **KDA** and… 23 Hugging Face Daily Papers research 3mo ago OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation Abstract OcclusionFormer addresses inter-object occlusion challenges in layout-to-image generation by modeling explicit Z-order priority through diffusion transformers and volume rendering techniques. AI-generated summary Recent layout-to-image models have achieved remarkable… 36 Hugging Face Daily Papers research 3mo ago PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset Abstract A large-scale UHR image-text dataset and evaluation benchmark are introduced to advance ultra-high-resolution text-to-image generation capabilities. AI-generated summary Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the… 29 r/LocalLLaMA community 3mo ago bytedance released an open source model that attempts to do just about anything with only 3b parameters Lance is a lightweight native unified multimodal model that supports image and video understanding, generation, and editing within a single framework. Efficient at 3B scale. With only 3B active parameters , Lance delivers strong performance across image generation, image… 32 arXiv — Machine Learning research 3mo ago Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra arXiv:2605.16259v1 Announce Type: new Abstract: While real-time image generation using diffusion models has advanced rapidly on NVIDIA GPUs, systematic optimization research on non-CUDA platforms such as Apple Silicon remains extremely limited. In this study, we conducted… 32 Hugging Face Daily Papers research 3mo ago Efficient Image Synthesis with Sphere Latent Encoder Abstract A decoupled framework for few-step image generation that improves efficiency and performance by separating pixel-space operations from latent denoising training. AI-generated summary Few-step image generation has seen rapid progress, with consistency and meanflow-based… 15 Hugging Face Daily Papers research 3mo ago InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Abstract InsightTok improves discrete visual tokenization for better text and face reconstruction through content-aware perceptual losses, enhancing autoregressive image generation quality. AI-generated summary Text and faces are among the most perceptually salient and… 12 Hugging Face Daily Papers research 3mo ago Aligning Latent Geometry for Spherical Flow Matching in Image Generation Abstract Geodesic flow matching improves image generation by projecting latents onto fixed radius spheres and using spherical linear interpolation instead of linear paths, preserving semantic content through angular components. AI-generated summary Latent flow matching for image… 26 Hugging Face Daily Papers research 3mo ago Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning Abstract Realiz3D addresses the domain gap between synthetic renders and real images in 3D-consistent image generation by decoupling visual domain from control signals through residual adapters and layer-specific denoising strategies. AI-generated summary We often aim to… 19 Hugging Face Daily Papers research 3mo ago Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Abstract A closed-loop visual reasoning framework integrates visual-language planning with diffusion generation to improve complex image synthesis while addressing latency and optimization challenges. AI-generated summary Despite rapid advancements, current text-to-image (T2I)… 14 Hugging Face Daily Papers research 3mo ago Does Synthetic Layered Design Data Benefit Layered Design Decomposition? Abstract Synthetic layered image data improves graphic design decomposition by enabling scalable training and better layer distribution control compared to traditional methods. AI-generated summary Recent advances in image generation have made it easy to produce high-quality… 35 r/LocalLLaMA community 3mo ago Built an open-source one-prompt-to-cinematic-reel pipeline on a single GPU — FLUX.2 [klein] for character keyframes, Wan2.2-I2V for animation, vision critic with auto-retry, music + 9-language narration in the same pipeline Shipped this for the AMD x lablab hackathon. Attached video is one of the actual reels the pipeline produced - one English sentence in, finished mp4 with characters, story, music, and voice-over out (fast demo video, not the best quality). ~45 minutes end-to-end on a single AMD… 13 Hugging Face Daily Papers research 3mo ago Asymmetric Flow Models Abstract Asymmetric Flow Modeling enables efficient high-dimensional flow-based generation by restricting noise prediction to low-rank subspaces while maintaining full-dimensional data prediction, achieving superior performance in pixel-space text-to-image generation through… 12 Hugging Face Daily Papers research 3mo ago Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Abstract INSET is a unified multimodal model that embeds images as native vocabulary within textual instructions, enabling better handling of complex interleaved inputs through transformer-based contextual locality and supporting both image generation and editing tasks.… 34 r/MachineLearning community 3mo ago Image generation models running locally on limited resources [P] I have a project consisting of generating high quality free ebook covers out of its content. On my 16GB of ram machine with no gpu, i have tested the opensourced stable diffusion models without any success. All return bad quality covers with blurred faces and scenes that do not… 6 arXiv — Machine Learning research 3mo ago Efficient Adjoint Matching for Fine-tuning Diffusion Models arXiv:2605.11480v1 Announce Type: new Abstract: Reward fine-tuning has become a common approach for aligning pretrained diffusion and flow models with human preferences in text-to-image generation. Among reward-gradient-based methods, Adjoint Matching (AM) provides a principled… 30 OpenAI news 4mo ago Introducing ChatGPT Images 2.0 ChatGPT Images 2.0 introduces a state-of-the-art image generation model with improved text rendering, multilingual support, and advanced visual reasoning. 9 Smol AI News news-outlet 4mo ago GPT-Image-2 **OpenAI** launched **GPT-Image-2**, enhancing image generation with improved text rendering, layout fidelity, editing, multilingual support, and "thinking" capabilities. It supports generating slides, infographics, diagrams, UI mockups, and QR codes, and integrates with tools… 36 OpenAI news 4mo ago Codex for (almost) everything The updated Codex app for macOS and Windows adds computer use, in-app browsing, image generation, memory, and plugins to accelerate developer workflows. 5 Hugging Face official-blog 5mo ago PRX Part 3 — Training a Text-to-Image Model in 24h! Back to Articles PRX Part 3 — Training a Text-to-Image Model in 24h! Team Article Published March 3, 2026 Upvote 64 David Bertoin Bertoin Photoroom Roman Frigg photoroman Photoroom Jon Almazán jon-almazan Photoroom Introduction Welcome back 👋 In the last two posts ( Part 1 and… 23 Smol AI News news-outlet 6mo ago Nano Banana 2 aka Gemini 3.1 Flash Image Preview: the new SOTA Imagegen model **Google and DeepMind** launched **Nano Banana 2** (aka **Gemini 3.1 Flash Image Preview**), a leading image generation and editing model integrated across multiple Google products with features like **4K upscaling**, **multi-subject consistency**, and **real-time… 29 Hugging Face official-blog 6mo ago Training Design for Text-to-Image Models: Lessons from Ablations Back to Articles Training Design for Text-to-Image Models: Lessons from Ablations Team Article Published February 3, 2026 Upvote 73 David Bertoin Bertoin Photoroom Roman Frigg photoroman Photoroom Jon Almazán jon-almazan Photoroom Welcome back! This is the second part of our… 13 Hugging Face official-blog 9mo ago Diffusers welcomes FLUX-2 Back to Articles Welcome FLUX.2 - BFL’s new open image generation model 🤗 Published November 25, 2025 Update on GitHub Upvote 190 YiYi Xu YiYiXu Daniel Gu dg845 Sayak Paul sayakpaul Alvaro Somoza OzzyGT Dhruv Nair dn6 Aritra Roy Gosthipaty ariG23498 Linoy Tsaban linoyts… 12 Google DeepMind official-blog 9mo ago Build with Nano Banana Pro, our Gemini 3 Pro Image model Build with Nano Banana Pro, our Gemini 3 Pro Image model Share x.com Facebook LinkedIn Mail Here’s how developers can use Nano Banana Pro (Gemini 3 Pro Image), a powerful new image generation and editing model with advanced features and creative control. Alisa Fortin Product… 10 Google DeepMind official-blog 17mo ago Experiment with Gemini 2.0 Flash native image generation Native image output is available in Gemini 2.0 Flash for developers to experiment with in Google AI Studio and the Gemini API. 5 Eugene Yan research 45mo ago Text-to-Image: Diffusion, Text Conditioning, Guidance, Latent Space The fundamentals of text-to-image generation, relevant papers, and experimenting with DDPM. 35 Page 4 of 4 · 176 articles ← Newer