News / #image-gen Tag Image Gen 176 articles archived under #image-gen · RSS Sign in to follow r/MachineLearning community 1d ago I implemented a very tiny image generation model (latent flow transformer) on a RP2350 microcontroller - it can generate 128x128 images of faces [P] Its a 2.4-4 million parameter model, quantized to int8, that can be fully executed on the microcontroller in ~20s with the longest generation. The generated image will then be displayed on a monitor or transferred via usb. Its a latent flow transformer with 12 layers using… 26 arXiv — Machine Learning research 3d ago Joint Initialization of Flux Networks and Effective Multiplication Factor for Physics-Informed Neural Networks Solving Neutron Diffusion Problems arXiv:2608.25443v1 Announce Type: new Abstract: Efficient determination of the effective multiplication factor (keff) is an important computational task in reactor core neutronics analysis. Physics-informed neural networks (PINNs) incorporate neutron diffusion equations and… 10 TechCrunch — AI news-outlet 4d ago Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding The company's new fundraising total now stands at $232 million. 29 Hugging Face Daily Papers research 6d ago UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Abstract A reparameterized pretrained vision transformer unifies semantic understanding, high-fidelity reconstruction, and image generation within a single visual space without requiring a separate VAE. Generated by thinkingmachines/Inkling-Small Semantic vision encoders have… 26 r/MachineLearning community 6d ago OpenAI advertises "unlimited" image generation in Pro. It is not unlimited. [D] Signed up for ChatGPT Pro for an image generation project. The pricing page says Pro comes with "unlimited and faster image creation" (chatgpt.com/pricing). A few hundred images in, I got rate limited. Reached out to support, and the response was: "Pro access remains subject to… 7 Hugging Face Daily Papers research 9d ago WithEveryone: Unified Planning and Identity Grounding for Group Image Generation Abstract WithEveryone enables reliable identity-preserving group image generation for up to ten people by grounding identities to layout plans and using region-based identity losses. Generated by thinkingmachines/Inkling-Small Identity-preserving image generation becomes… 17 Hugging Face Daily Papers research 11d ago DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization Abstract DiSCO is a black-box, zero-shot prompt-level defense that uses distribution-guided suffix expansion and contrastive scoring to reduce harmful image generation without altering the model. Generated by thinkingmachines/Inkling-Small As text-to-image generative models… 27 Hugging Face Daily Papers research 11d ago Abra: Scaling Diffusion Image Training Abstract Scaling laws for text-to-image diffusion models reveal predictable compute-optimal training requiring far more data per parameter than language models, with robust overtraining behavior and universal curve shapes. Generated by thinkingmachines/Inkling-Small… 22 arXiv — Machine Learning research 11d ago Abra: Scaling Diffusion Image Training arXiv:2608.17286v1 Announce Type: new Abstract: Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled… 7 Hugging Face Daily Papers research 11d ago From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation Abstract A capability-driven data infrastructure with curriculum scheduling and specialized data engines trains large multimodal diffusion models on curated heterogeneous supervision for diverse generative tasks. Generated by thinkingmachines/Inkling-Small Large-scale image… 13 r/MachineLearning community 12d ago Trained an diffusion model that runs on 264KB of RAM [P] I recently bought a Shrike lite which has got 264KB of SRAM. I decided to train an image generation model that generates 32*32 pixel images. The microcontroller also has an FPGA onboard which I used to create two parallel INT8 MAC engines with 16 bit accumulation to speed up… 36 Hugging Face Daily Papers research 12d ago TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation Abstract This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations. Generated by thinkingmachines/Inkling-Small Despite recent advances in unified multimodal models for multi-reference… 22 Hugging Face Daily Papers research 12d ago An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Abstract Researchers propose a latent-to-pixel training strategy that accelerates convergence and improves inference speed for large-scale pixel-space diffusion models. Generated by thinkingmachines/Inkling-Small This paper investigates an increasingly important topic in… 16 Hugging Face Daily Papers research 12d ago GenRouter: Unified Workflow Routing for Agentic Image Generation Abstract GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution. Generated by thinkingmachines/Inkling-Small The… 34 arXiv — Machine Learning research 13d ago Adversarial Learning of Classifier-Free Guidance Schedules arXiv:2608.14038v1 Announce Type: new Abstract: Modern text-to-image diffusion models rely on classifier-free guidance (CFG) to achieve high image fidelity and text alignment. However, CFG typically applies a static, global scale across all timesteps, samples, and conditions --… 26 r/LocalLLaMA community 15d ago [Megathread] Qwen 3.8 27B Release Day Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comparisons Official:… 10 r/MachineLearning community 16d ago Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D] I may have stumbled onto something interesting while trying to figure out a recurring artifact in ChatGPT image generation and editing (maybe applicable to other models as well?). It started with a very practical problem: After several rounds of generative editing on portraits,… 15 arXiv — NLP / Computation & Language research 18d ago On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation arXiv:2608.11002v1 Announce Type: new Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects… 37 Hugging Face Daily Papers research 18d ago Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation Abstract Atelier improves artist-grounded image generation by translating vague artistic intent into explicit control states that separate scene content from style, reducing reliance on stereotypical shortcuts. Generated by thinkingmachines/Inkling-Small Artist-grounded image… 20 arXiv — NLP / Computation & Language research 20d ago Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar… 36 Simon Willison community 22d ago Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to… 33 Simon Willison community 22d ago Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to… 20 arXiv — NLP / Computation & Language research 23d ago GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers arXiv:2608.05478v1 Announce Type: cross Abstract: Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. Recently, advancements in vision-language models and image generation… 38 Hugging Face Daily Papers research 24d ago ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Abstract Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent… 36 arXiv — NLP / Computation & Language research 24d ago Simile Understanding in Text-to-Image Models: An Evaluation Framework arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models… 12 Hugging Face Daily Papers research 24d ago Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models Abstract Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to… 23 Simon Willison community 24d ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 29 Simon Willison community 24d ago One-shotting a Raccoon Heist game using Claude Fable 5 Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content… 30 Hugging Face Daily Papers research 25d ago PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs Abstract Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generation is not element-editable, while coding-agent workflows are costly.… 35 arXiv — Machine Learning research 25d ago A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics arXiv:2608.02965v1 Announce Type: new Abstract: Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. Under these… 26 Hugging Face Daily Papers research 25d ago Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Abstract Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and… 8 Hugging Face Daily Papers research 25d ago CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Abstract Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information… 22 Hugging Face Daily Papers research 25d ago UniWorld-Design: From Pixel Generation to Layer-Native Design Abstract We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an… 17 arXiv — NLP / Computation & Language research 26d ago Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable to malicious jailbreak prompts. Transfer-based attacks construct adversarial… 6 arXiv — Machine Learning research 27d ago WaiT for the Signal: Simple Frequency-Aware Flow-Matching arXiv:2607.28760v1 Announce Type: cross Abstract: As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies… 37 arXiv — Machine Learning research 1mo ago Flux-OPD: On-Policy Distillation with Evolving Contexts arXiv:2607.28022v1 Announce Type: new Abstract: Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision… 13 Hugging Face Daily Papers research 1mo ago MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Abstract Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities,… 12 Hugging Face Daily Papers research 1mo ago Flux-OPD: On-Policy Distillation with Evolving Contexts Abstract Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student,… 24 Hugging Face Daily Papers research 1mo ago TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Abstract Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework… 27 arXiv — Machine Learning research 1mo ago Learning Sampling Parameters for Diffusion Models arXiv:2607.23488v1 Announce Type: new Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then… 29 TechCrunch — AI news-outlet 1mo ago Midjourney acquired the astrology app Co-Star The AI lab Midjourney continues to expand its purview beyond image and video generation. 11 r/LocalLLaMA community 1mo ago FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3   submitted by   /u/pmttyji [link]   [comments] 23 Hacker News — AI on Front Page community 1mo ago Flux 3 X Mimic: The Next Generation of Video-Action Models Article URL: https://bfl.ai/blog/flux-3-mimic Comments URL: https://news.ycombinator.com/item?id=49033127 Points: 240 # Comments: 32 24 Hacker News — AI on Front Page community 1mo ago Flux 3 Article URL: https://bfl.ai/blog/flux-3 Comments URL: https://news.ycombinator.com/item?id=49031796 Points: 313 # Comments: 77 36 Latent.Space news-outlet 1mo ago [AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model A HUGE win for BFL! 22 arXiv — Machine Learning research 1mo ago Neural Operator Surrogates for Two-Dimensional Neutron Flux Estimation arXiv:2607.19388v1 Announce Type: new Abstract: This work extends our one-dimensional single-sweep neural-operator studies to two dimensions. We consider one-group transport with isotropic scattering. As in the one-dimensional work, we use Fourier neural operators (FNOs) to… 38 r/LocalLLaMA community 1mo ago Mage-Flow - An Efficient Native-Resolution Foundation Model for Image Generation and Editing - Microsoft Models: (Check Model cards for so much sample demo images) https://huggingface.co/microsoft/Mage-Flow https://huggingface.co/microsoft/Mage-Flow-Turbo https://huggingface.co/microsoft/Mage-Flow-Edit Mage-Flow is a compact 4B-scale generative stack for efficient text-to-image… 30 r/MachineLearning community 1mo ago Anyone heading to Jeju for KDD? Let's meet up! 🙋[D] Hey all! Is anyone else going to be at KDD in Jeju? Would love to connect with fellow attendees. I work on interpretability, fairness, and editing of text-to-image models, so I'd especially love to meet people working in these areas. But honestly, we can chat about anything:… 4 Hugging Face Daily Papers research 1mo ago Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Abstract Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers… 20 Hugging Face Daily Papers research 1mo ago Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Abstract Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention… 8 Page 1 of 4 · 176 articles Older →