News / #image-gen Tag Image Gen 176 articles archived under #image-gen · RSS Sign in to follow Hugging Face Daily Papers research 1mo ago Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Abstract Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two… 23 arXiv — Machine Learning research 1mo ago Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation arXiv:2607.14962v1 Announce Type: new Abstract: Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for… 4 arXiv — Machine Learning research 1mo ago Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models arXiv:2607.14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must… 5 Hugging Face Daily Papers research 1mo ago Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation Abstract We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference,… 10 arXiv — Machine Learning research 1mo ago SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning arXiv:2607.12042v1 Announce Type: cross Abstract: Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn… 32 Hugging Face Daily Papers research 1mo ago Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Abstract In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification… 27 Hugging Face Daily Papers research 1mo ago Latent-Identity Tuning in Text-to-Image Personalization Models Abstract Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the… 36 Hugging Face Daily Papers research 1mo ago From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Abstract Pretrained diffusion transformers can be adapted for dense prediction tasks by mapping tokens to task-native outputs instead of generating RGB images, achieving state-of-the-art results with minimal additional parameters. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 6 r/LocalLLaMA community 1mo ago I built Flaxeo Image a local desktop ui for stable diffusion cpp Built around a recent sd.cpp release, aims to expose most of what the backend can do (generate, edit, video paths, models, hardware options), Windows + Linux builds GitHub: https://github.com/fabricio3g/FlaxeoUI   submitted by   /u/fabricio3g [link]   [comments] 6 Vercel — AI dev-tools 1mo ago Seedream 5.0 Pro is now available on AI Gateway Seedream 5.0 Pro is now available on AI Gateway . Seedream 5.0 Pro is an image generation and editing model. It generates images from text, rendering text without spelling errors and following typographic rules, and produces dense infographics with charts, timelines, and layouts… 20 arXiv — Machine Learning research 1mo ago AutoAnchor: Stable Diffusion Unlearning Using Cross-Attention as a Manifold Surrogate arXiv:2607.08337v1 Announce Type: new Abstract: Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of… 21 arXiv — NLP / Computation & Language research 1mo ago Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared… 23 Hugging Face Daily Papers research 1mo ago Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Abstract Flash-BoN improves text-to-image generation efficiency by using inexpensive draft candidates generated through timestep truncation, layer skipping, and activation proxies, followed by multi-stage verification that outperforms existing methods under fixed wall-clock… 9 arXiv — Machine Learning research 1mo ago An Hybrid Quantum-Classical Diffusion Model for Image Generation arXiv:2607.07072v1 Announce Type: new Abstract: Quantum diffusion models provide a physics-consistent route to generative learning by formulating noising and denoising directly on quantum states. However, applying such models to classical high-dimensional data is constrained by… 9 arXiv — NLP / Computation & Language research 1mo ago Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing… 5 arXiv — Machine Learning research 1mo ago TILDE: TILt-based Distributional Erasure for Concept Unlearning arXiv:2607.06432v1 Announce Type: new Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to… 22 Hugging Face Daily Papers research 1mo ago CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training Abstract A 3D latent diffusion model for chest CT generation that enables controlled synthesis of medical images with clinical attributes through adaptive layer normalization and reinforcement learning post-training. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Controllable… 35 r/LocalLLaMA community 1mo ago [Paper] Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system-level optimization. Among them, multi-resolution generation strategies have recently received… 25 TechCrunch — AI news-outlet 1mo ago Midjourney wants Hollywood studios to reveal the details of their AI usage As part of an ongoing legal dispute with three Hollywood studios, Midjourney is seeking to compel those studios to reveal how they use AI themselves. 24 r/MachineLearning community 1mo ago I built my 'first' flow matching image generator, here's what I learned [P] Today I put out my first flow matching image generation model! This is a toy example trained on a 2024 MPS Macbook Pro using a small sample of images—specifically, the Apple emoji library and their text labels. Because of this, it’s not a massive model (clocking in at ~4.7… 20 Hugging Face Daily Papers research 1mo ago Representation Distribution Matching for One-Step Visual Generation Abstract Representation Distribution Matching enables high-quality image generation by matching feature distributions under pretrained encoders, with improved performance through optimized batch sizes and multi-encoder evaluation metrics. Generated by… 33 Hugging Face Daily Papers research 1mo ago Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Abstract MrFlow accelerates text-to-image diffusion by combining low-resolution generation with pixel-space super-resolution and noise injection, achieving up to 25x speedup without training or runtime modifications. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Hardware-agnostic… 6 Hugging Face Daily Papers research 1mo ago SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Abstract Scientific image generation faces challenges in semantic alignment and logical reasoning, prompting the creation of SciIR-82k dataset and SciIR-Bench evaluation framework to improve scientific reasoning capabilities in text-to-image models. Generated by… 23 Hugging Face Daily Papers research 2mo ago DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Abstract DataEvolver is a self-evolving multi-agent framework that improves text-rich image generation by leveraging feedback from rejected samples to iteratively enhance data quality. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Text-rich image generation is one of the most… 11 Hugging Face Daily Papers research 2mo ago InstanceControl: Controllable Complex Image Generation without Instance Labeling Abstract InstanceControl enables multi-instance image generation by using vision-language models to establish instance-level correspondences between text prompts and visual conditions, while employing adaptive mask refinement for improved accuracy. Generated by… 29 arXiv — Machine Learning research 2mo ago Quality-Aware Modulation for Diffusion Transformers arXiv:2606.30934v1 Announce Type: new Abstract: Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep. While this modulation communicates the current… 31 Hugging Face Daily Papers research 2mo ago Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation Abstract ILLUME-X is a unified multimodal paradigm that enhances text-image generation through improved data efficiency, stable training processes, and comprehensive evaluation metrics. Generated by Qwen/Qwen2.5-Coder-32B-Instruct The advancement of generative AI models capable… 17 arXiv — Machine Learning research 2mo ago Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models arXiv:2606.28406v1 Announce Type: new Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts. Yet existing… 36 Hugging Face Daily Papers research 2mo ago Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting Abstract Flux-GS enables real-time high-fidelity 3D Gaussian Splatting on mobile platforms through efficient lighting representation, attribute-conditioned enhancement, and multi-view densification strategies. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent advances in 3D… 10 Hugging Face Daily Papers research 2mo ago Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Abstract A masked discrete diffusion model for text-to-image synthesis that addresses limitations in token refinement and training efficiency through novel mechanisms and optimizations. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We propose Nemotron-Labs-Diffusion-Image, a… 25 Hugging Face Daily Papers research 2mo ago MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Abstract MIMFlow combines Normalizing Flows with Masked Image Modeling to improve generative modeling by decoupling semantic representation from pixel-level details, achieving better performance with fewer tokens. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Normalizing Flows… 37 TechCrunch — AI news-outlet 2mo ago Gemini’s personalized AI image generation is now free for US users Google is expanding Gemini’s personalized AI image generation to eligible free users in the U.S., allowing the chatbot to create images based on your interests and data from connected Google apps. 29 r/LocalLLaMA community 2mo ago clark-labs/clark-air-sana-1.6b-1.58bit · Hugging Face A Sana 1.6B text-to-image transformer compressed to ternary (~1.85 bits/weight): 8.6× smaller than FP16, near-FP16 quality. Footprint (measured) Artifact Size vs FP16 What it is FP16 transformer 3.21 GB 1× (100%) reference Clark Air (packed) 374 MB 8.6× (≈12%) packed ternary (… 36 Hugging Face Daily Papers research 2mo ago Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Abstract A unified agentic framework called Qwen-Image-Agent is proposed to address the context gap in text-to-image generation by progressively constructing complete generation context through planning, reasoning, searching, and memory mechanisms. Generated by… 22 arXiv — NLP / Computation & Language research 2mo ago DanceOPD: On-Policy Generative Field Distillation arXiv:2606.27377v1 Announce Type: cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For… 18 Hugging Face Daily Papers research 2mo ago DanceOPD: On-Policy Generative Field Distillation Abstract A novel on-policy generative field distillation framework called DanceOPD is proposed to unify text-to-image generation, local editing, and global editing capabilities in flow-matching models through capability-specific routing and velocity-based training. Generated by… 10 Hugging Face Daily Papers research 2mo ago IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Abstract Implicit Visual Chain-of-Thought decomposes visual conditioning into structural and semantic cascades for improved structure-aware image generation with sketch supervision. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Unified multi-modal large language models (MLLMs)… 7 r/LocalLLaMA community 2mo ago SDXL running locally in the browser on WebGPU, open-source I needed simple local image generation without the usual setup. No virtual environments, no ComfyUI with a complex graph and installation as an exe. So i tried to push the whole thing into the browser and run it on WebGPU. It's a browser extension. You install it, then it loads… 13 Hugging Face Daily Papers research 2mo ago Semantic Browsing: Controllable Diversity for Image Generation Abstract Text-to-image models are enhanced with controlled diversity through semantic browsing capabilities that enable structured navigation of image variations based on meaningful semantic decisions. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Modern text-to-image models… 4 Hugging Face Daily Papers research 2mo ago FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation Abstract FLUX3D addresses limitations in image-to-3D Gaussian Splatting generation by improving representation learning and cross-modal alignment through specialized architectures and attention mechanisms. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Sparse voxel representation… 34 arXiv — Machine Learning research 2mo ago Information-Theoretic Classifier-Free Guidance with Adaptive Schedule Optimization arXiv:2606.24025v1 Announce Type: new Abstract: Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-free guidance (CFG). CFG improves condition consistency by increasing a… 35 Hugging Face Daily Papers research 2mo ago Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning Abstract Text-to-image models fail to generate counterfactual scenes because they rely on tightly coupled visual-textual patterns rather than causal reasoning, demonstrating limited understanding beyond pattern matching. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Text-to-image… 26 Hugging Face Daily Papers research 2mo ago Safe Few-Step Generation via Velocity Editing Abstract VESFlow is a training-free safety method for flow matching-based text-to-image generation that edits velocity fields to ensure safe output while maintaining prompt integrity. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Flow matching has recently emerged as a strong… 16 r/LocalLLaMA community 2mo ago Boogu Base, Turbo, Edit - open-source unified image generation and editing model series Boogu-Image-0.1 is a competitive Apache-2.0 open-source unified image generation and editing model family , including Base , Turbo , Edit , and other variants that provide stable, practical capabilities for high-quality text-to-image generation, fast generation, image editing,… 22 Hugging Face Daily Papers research 2mo ago Exploring the Design Space of Reward Backpropagation for Flow Matching Abstract FlowBP addresses limitations in flow matching model alignment by using a surrogate trajectory framework that reduces memory usage and gradient chaining while maintaining performance across multiple text-to-image models. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 23 Hugging Face Daily Papers research 2mo ago BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation Abstract A 3D brain MRI generative model uses a masked-autoencoder tokenizer to create clinically informative embeddings that support both medical task performance and controlled image generation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Three-dimensional (3D) brain MRI is… 6 r/LocalLLaMA community 2mo ago Local text to image model comparaison: The ultimate test. I selected 192 prompts to evaluate text-to-image model various capabilities and generated images for all the local models I was able to make work on my GX10 Spark. For instance: Is the model good at text? At faces? At human anatomy? At respecting spatial composition, etc...? You… 4 r/MachineLearning community 2mo ago Studying FLUX in diffusers library was hard, so I built a smaller open-source version [P] If you've tried to study modern diffusion models by digging through the official diffusers library, you know it can be overwhelming with its complexity and abstractions. I wanted to simplify FLUX diffusion models, so I built minFLUX : a PyTorch implementation focused on its core… 38 Hugging Face Daily Papers research 2mo ago The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation Abstract Analysis of FID variance across different training and sampling seeds reveals significant reproducibility issues in image generation evaluation, with retraining causing larger fluctuations than resampling, and recommends updated evaluation protocols with error bars and… 21 arXiv — NLP / Computation & Language research 2mo ago NAMESAKES: Probing Identity Memorization in Text-to-Image Models arXiv:2606.20155v1 Announce Type: cross Abstract: Text-to-image (T2I) models generate realistic likenesses of some individuals when prompted with their names, raising privacy concerns. However, distinguishing whether a generated face is memorized or fabricated currently requires… 10 Page 2 of 4 · 176 articles ← Newer Older →