New Stable Diffusion Model Brings Fast AI Art to Consumer Computers

Stability AI and its collaborators announce Stable Diffusion, a groundbreaking text-to-image model that runs efficiently on consumer GPUs. The highly efficient system democratizes AI art generation by requiring less than 10 GB of video memory.

Stability AI and a collaborative group of researchers announce the initial release of Stable Diffusion, a text-to-image model that creates stunning art in seconds. The team, led by researchers from Runway and LMU Munich, builds this breakthrough on prior latent diffusion models to empower billions of people with fast, high-quality image generation.

Unlike previous heavy models, Stable Diffusion runs on consumer-grade graphics cards by requiring under 10 GB of VRAM. This efficiency allows the system to generate 512x512 pixel images in just a few seconds without needing expensive enterprise hardware, truly democratizing the creative process for everyday users and researchers.

The model trains on a specialized subset of the LAION-5B dataset called LAION-Aesthetics, which filters images based on aesthetic quality using a new CLIP-based model. Stability AI currently grants access to researchers through Hugging Face and conducts large-scale beta testing with over 10,000 users generating millions of images ahead of a full public release.

Read More at the original source →