News / #image-gen Tag Image Gen 176 articles archived under #image-gen · RSS Sign in to follow Latent.Space news-outlet 2mo ago [AINews] Midjourney Medical: scan your organs like you step on a scale The only bootstrapped frontier lab announces its second product and second 12 Hacker News — AI on Front Page community 2mo ago Midjourney Medical https://www.midjourney.com/medical Video: https://x.com/midjourney/status/2067422898407837797 Comments URL: https://news.ycombinator.com/item?id=48579650 Points: 228 # Comments: 203 10 Smol AI News news-outlet 2mo ago Midjourney Medical: scan your organs like you step on a scale **Midjourney** unveiled a new **medical imaging/scanning system** called the **Midjourney Scanner**, described as **radiation-free, magnet-free, fast, and low-cost**, but requiring a **water immersion tank** and having **coarser resolution than CT/MRI**. The announcement… 12 Hugging Face Daily Papers research 2mo ago Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification Abstract UniAR presents a unified autoregressive framework that uses a single discrete visual tokenizer to bridge visual understanding and generation, achieving state-of-the-art results in image generation and editing through multi-level feature fusion, bitwise quantization, and… 19 r/LocalLLaMA community 2mo ago Open Dungeon: local roleplay with Gemma 4 QAT + inline Uncen-FLUX images, running at full 256K context under 8GB RAM (OS) I wanted AI Dungeon but fully local and actually private, so I built it. The narrator is Gemma 4 (QAT Q4) through Ollama, and when a scene is worth showing it draws the picture too, locally, with FLUX. No API keys, no cloud, nothing leaves your machine. The part that surprised… 26 Hugging Face Daily Papers research 2mo ago Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback Abstract Structured Defect Grounding (SDG) addresses limitations in text-to-image model diagnosis by modeling defects as structured sets and using vision-language models for detection and reward-based alignment. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Despite generating… 22 Hugging Face Daily Papers research 2mo ago High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation Abstract A 2-step image generation model is developed through distillation from an 8-step teacher using distribution-aligned adversarial learning, step-decoupled parameterization, and end-to-end training with iterative regularization. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 33 Hugging Face Daily Papers research 2mo ago Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents Abstract Evoflux enables compact language models to execute tool workflows more reliably by using evolutionary search to repair failed plans during inference, significantly improving execution feasibility compared to traditional fine-tuning methods. Generated by… 20 arXiv — Machine Learning research 2mo ago Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs arXiv:2606.12280v1 Announce Type: new Abstract: Post-training quantization lets large text-to-image diffusion transformers run on consumer GPUs, yet the hardware-specific trade-offs are seldom measured directly. We quantize Ideogram 4.0 - a 9.3B flow-matching diffusion… 17 Hugging Face Daily Papers research 2mo ago Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Abstract A teacher-student framework decouples complex reasoning from efficient reward deployment in text-to-image training, achieving superior preference accuracy and optimization performance. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Reward models are central to… 22 Hugging Face Daily Papers research 2mo ago i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models Abstract A comprehensive experimental study of text-to-image diffusion models reveals key design choices and training insights leading to the development of i1, a 3B-parameter model that matches leading performance while maintaining full openness. Generated by… 21 Ars Technica — AI news-outlet 2mo ago Google DeepMind releases DiffusionGemma, a model that runs local AI 4x faster Diffusion AI is most common in image generation, but it can make text outputs much faster. 29 Hacker News — AI on Front Page community 2mo ago Mercedes‑Benz starts large‑scale production of electric axial flux motor Article URL: https://media.mercedes-benz.com/en/article/bebac2af-acdc-465a-9538-adb0bf3d8ccf Comments URL: https://news.ycombinator.com/item?id=48472877 Points: 262 # Comments: 139 21 Hugging Face Daily Papers research 2mo ago Text-to-Image Models Need Less from Text Encoders Than You Think Abstract Text-to-image models primarily utilize basic text representation aspects like word merging and order rather than complex contextual information encoded in full text embeddings. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Text-to-image models rely on text prompts as… 36 r/MachineLearning community 2mo ago Open image generation models are closer to closed-source quality than this sub thinks [D] I run evaluations on generative image models as part of my workflow, mostly comparing coherence, prompt adherence, and compositional accuracy across different architectures. The consensus here seems to be that open models are still a generation behind closed APIs. Based on my… 25 Hacker News — AI on Front Page community 2mo ago Ask HN: What was your "oh shit" moment with GenAI? Most of us were amused when DALL-E and its peers went mainstream, and we were quick to point out the obvious flaws. Then ChatGPT hit the scene and again, many of us dismissed it as a parlor trick that would never amount to much. Using LLMs for coding initially was a only small… 26 Hugging Face Daily Papers research 2mo ago Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting Abstract Multi-concept customization in text-to-image generation is improved through prompt-aware weighting strategies that reduce interference between learned visual concepts. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Low-Rank Adaptation (LoRA) successfully enables… 5 Hugging Face Daily Papers research 2mo ago Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs Abstract Research reveals significant disparities between text and image generation capabilities in multimodal models, with effective textual knowledge editing not transferring reliably to visual output, necessitating modality-aware editing approaches. Generated by… 9 r/MachineLearning community 2mo ago Research in Image/Video Gen AI models [D] I've been going down a rabbit hole with image/video generation/editing models for a few months now, started with playing around with Stable Diffusion and ComfyUI, then got genuinely hooked on understanding why things work, not just that they do. I have an Engineering background… 20 The Information — AI news-outlet 2mo ago Cybersecurity’s AI Paradox It's no secret that criminals are using AI to streamline computer hacks in hopes of emptying out people’s bank accounts (never has it looked more appealing to stash cash under the mattress!). Cybersecurity executives, meanwhile, are rubbing their hands with glee at the influx of… 32 Hugging Face Daily Papers research 2mo ago Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation Abstract Decoupled Residual Denoising Diffusion models (DRDD) improve unified image-to-image translation by separating noise diffusion for domain harmonization from residual diffusion for semantic mapping, enhancing data efficiency and performance. Generated by… 32 r/LocalLLaMA community 2mo ago 1-bit Bonsai Image 4B and Ternary Bonsai Image 4B Image Generation for Local Devices with just 0.93 GB and 1.21 GB respectively of Diffusion Transformer Footprint. So tiny! https://prismml.com/news/bonsai-image-4b   submitted by   /u/Addyad [link]   [comments] 6 Hacker News — AI on Front Page community 2mo ago Adafruit Receives Demand Letter from Fenwick Legal Counsel on Behalf of Flux.ai Article URL: https://blog.adafruit.com/ Comments URL: https://news.ycombinator.com/item?id=48368121 Points: 255 # Comments: 87 11 arXiv — Machine Learning research 2mo ago CHAM-net: A Contrastive Hierarchical Adaptive Meta-network for Robust Global Methane Flux Prediction arXiv:2606.00338v1 Announce Type: new Abstract: Methane is a potent greenhouse gas that significantly contributes to global warming. However, accurately estimating global methane emissions and consumption remains challenging due to the complex interactions among environmental… 28 Hugging Face Daily Papers research 2mo ago Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization Abstract BiDPO enhances text-to-image models for complex compositional prompts through preference-based fine-tuning and region-level guidance. AI-generated summary Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex… 18 Hugging Face Daily Papers research 2mo ago Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization Abstract GCPO enables per-token credit assignment in reinforcement learning by contrasting model predictions under positive and negative prompts, improving performance in text-to-image generation and chain-of-thought reasoning tasks. AI-generated summary Group-advantage-based… 27 Hugging Face Daily Papers research 3mo ago Representation Forcing for Bottleneck-Free Unified Multimodal Models Abstract Representation Forcing enables unified multimodal models to perform both perception and generation tasks end-to-end without relying on external latent spaces, matching state-of-the-art performance in image generation while improving understanding capabilities.… 27 Hacker News — AI on Front Page community 3mo ago 1-Bit Bonsai Image 4B Image Generation for Local Devices Article URL: https://prismml.com/news/bonsai-image-4b Comments URL: https://news.ycombinator.com/item?id=48346257 Points: 228 # Comments: 81 36 r/LocalLLaMA community 3mo ago Should I buy this RTX 2060 12GB graphics card at around $260 for AI purpose ? I’m interested in running Gemma 4 model/s for text only . It runs smooth even on my laptop but gets crazy hot. Initially wanted to buy an 8 GB card. But I find this price for 12 GB good. (Maybe I can run some image generation models too. But its not important.) It has 6 Month… 10 r/LocalLLaMA community 3mo ago Could someone make some ggufs for Qwen-Image-Bench? I'd like to try it out for automating image generation quality output, I haven't had great luck with that using 27b base or gemma. If this can reliably detect 6 fingered generations and other undesirable outputs it would be a great boon. I took a swing and quantizing it myself… 30 arXiv — Machine Learning research 3mo ago Moment Matching Q-Learning arXiv:2605.29033v1 Announce Type: new Abstract: Score-based and flow-based generative models exhibit remarkable expressive capacity in capturing complex distributions, and have been extensively deployed in tasks ranging from image generation to reinforcement learning.… 31 Hugging Face Daily Papers research 3mo ago GenClaw: Code-Driven Agentic Image Generation Abstract GenClaw presents a code-driven agentic image generation framework that enables precise visual construction through conceptualization, sketching, and coloring stages, integrating programmatic logic with generative models. AI-generated summary Image generation models have… 8 r/LocalLLaMA community 3mo ago Qwen/Qwen-Image-Bench · Hugging Face Model Description Q-Judger is a vision-language model fine-tuned specifically for automated evaluation of text-to-image generated images. Given a text prompt and a generated image, the model evaluates the image on fine-grained quality criteria organized in a 3-level hierarchy… 8 arXiv — NLP / Computation & Language research 3mo ago ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment arXiv:2605.27374v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, personalized cover image generation remains underexplored, despite its critical… 26 arXiv — NLP / Computation & Language research 3mo ago PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI arXiv:2605.27545v1 Announce Type: new Abstract: Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text and current defenses are relatively immature. We introduce PAST2HARM, a simple… 38 Hugging Face Daily Papers research 3mo ago MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale Abstract A 20B-parameter masked region diffusion model enables scalable multi-layer transparent image generation and editing through unified task handling and efficient canvas management. AI-generated summary Layered image generation and editing is a fundamental capability that… 21 arXiv — Machine Learning research 3mo ago Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models arXiv:2605.26491v1 Announce Type: new Abstract: Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image diffusion models. However, existing methods largely reduce supervision to binary… 10 arXiv — Machine Learning research 3mo ago On the Error-Correcting Effects of Stochasticity in Discrete Diffusion arXiv:2605.26582v1 Announce Type: new Abstract: Discrete diffusion models achieve strong performance in text and image generation, but their inference remains slow and must inherently balance sampling efficiency and sample quality. In this work, we present a systematic study of… 8 arXiv — Machine Learning research 3mo ago RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models arXiv:2605.26632v1 Announce Type: new Abstract: Diffusion Transformers (DiT) achieve strong performance in image generation but incur substantial inference costs. While prior work has reduced this cost via quantization and distillation, semi-structured sparsity, which can nearly… 27 Hugging Face Daily Papers research 3mo ago Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation Abstract A novel approach conditions diffusion models on multimodal large language models for subject-driven image generation, combining text and reference image encoding with VAE-based identity conditioning to improve both semantic understanding and identity preservation.… 7 Hugging Face Daily Papers research 3mo ago RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models Abstract Diffusion Transformers achieve strong image generation performance but face high inference costs; this work proposes RT-Lynx, which uses activation sparsification and optimized CUDA kernels to accelerate inference while maintaining generation quality. AI-generated… 27 r/LocalLLaMA community 3mo ago PrismML just released Binary and Ternary Bonsai Image 4B: 1-bit/ternary text-to-image diffusion transformers that can even run 100% locally in your browser on WebGPU. The PrismML team really cooked with these models. They're only ~3GB in size (compared to FLUX.2 Klein 4B, which is ~16GB). Apache-2.0! Official collection on HF: https://huggingface.co/collections/prism-ml/bonsai-image Link to demo:… 11 Hugging Face Daily Papers research 3mo ago Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference Abstract Visual Concept Fusion enables dual text and image conditioning in diffusion models through feature alignment and fusion strategies without requiring retraining. AI-generated summary Text-to-image diffusion models like Stable Diffusion generate high-quality images from… 35 Hugging Face Daily Papers research 3mo ago Reinforcing Few-step Generators via Reward-Tilted Distribution Matching Abstract RTDMD is a two-stage framework that combines distribution matching distillation with reward-guided reinforcement learning to improve few-step image generation alignment with human preferences. AI-generated summary Recent advances in few-step diffusion distillation have… 30 Hugging Face Daily Papers research 3mo ago Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Abstract Lens is a compact 3.8B-parameter text-to-image model achieving superior performance with reduced training compute through dense caption datasets, multi-resolution batching, efficient architecture, and optimization techniques. AI-generated summary We introduce Lens, a… 19 Hugging Face Daily Papers research 3mo ago ETCHR: Editing To Clarify and Harness Reasoning Abstract A novel image editing approach called ETCHR is introduced that decouples visual reasoning from image generation, improving multimodal language model performance across multiple visual reasoning tasks through a two-stage training process. AI-generated summary Multimodal… 7 Hugging Face Daily Papers research 3mo ago AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Abstract AutoRubric-T2I automatically generates and selects explicit rubrics to guide Vision-Language Model judges for text-to-image generation, achieving high-quality reward signals with minimal human annotation while improving generation quality in downstream tasks.… 36 Hugging Face Daily Papers research 3mo ago RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution Abstract Discrete autoregressive text-to-image models suffer from latent covariate shift during policy optimization, which RankE addresses through end-to-end co-evolution of policy and decoder components. AI-generated summary Discrete autoregressive (AR) text-to-image (T2I)… 9 Hugging Face Daily Papers research 3mo ago SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers Abstract SEGA improves high-resolution text-to-image generation by adaptively scaling attention across RoPE components based on spatial-frequency structure during denoising steps. AI-generated summary Diffusion transformers (DiTs) have emerged as a dominant architecture for… 33 Hugging Face Daily Papers research 3mo ago GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation Abstract A self-evolving image generation framework uses tool-orchestrated trajectories and visual experience distillation to improve generative capabilities through iterative learning and reference-based prompting. AI-generated summary Open-ended image generation is no longer a… 19 Page 3 of 4 · 176 articles ← Newer Older →