Stable Diffusion Brings Fast Text-to-Image Generation to Consumer Hardware

Stability AI unveils Stable Diffusion, a highly efficient text-to-image model that runs on consumer GPUs. The initial release goes out to researchers ahead of a planned public launch.

Stability AI and its collaborators announce the initial release of Stable Diffusion, a groundbreaking text-to-image model that creates stunning art in seconds. The model represents a major speed and quality breakthrough because it runs efficiently on consumer-grade graphics cards rather than requiring massive enterprise hardware. Hugging Face hosts the model weights for researchers while the development team prepares for a full public release.

The project builds upon prior latent diffusion research from Runway and LMU Munich, incorporating insights from various conditional diffusion models and the broader AI community. Stability AI trains this model on the LAION-Aesthetics dataset, a carefully filtered subset of LAION 5B that uses CLIP-based ratings to select visually appealing images. The company utilizes its 4,000 A100 Ezra-1 ultracluster to train the model over the past month.

By requiring under 10 GB of VRAM to generate 512x512 pixel images in just a few seconds, Stable Diffusion successfully democratizes AI media generation. Over 10,000 beta testers are currently putting the model to the test, generating millions of images to explore the boundaries of latent space. Stability AI hopes this cooperative approach continues to bring the gift of creativity to billions of people worldwide.

Read More at the original source →