Stability AI Releases Open-Source Stable Diffusion Image Generator
Stability AI launches Stable Diffusion, a highly accessible open-source text-to-image model that runs efficiently on consumer hardware. The model allows both commercial and non-commercial use under a new ethical license.
London-based startup Stability AI officially releases Stable Diffusion, an open-source text-to-image model that competes directly with DALL-E 2 and Google's Imagen. This advanced system results from a massive collaboration involving RunwayML, LMU Munich, EleutherAI, and LAION, and builds upon existing latent diffusion research combined with various conditional diffusion models.
The model trains on LAION-Aesthetics, a carefully filtered subset of the massive 5.85 billion-pair LAION 5B dataset, using an Ezra-1 AI ultracluster of 4,000 A100 GPUs. Unlike competing systems that require expensive cloud infrastructure, Stable Diffusion efficiently generates 512x512 pixel images in seconds while only consuming about 6.9 GB of VRAM on standard consumer GPUs.
Stability AI makes the model widely accessible by releasing it under a Creative ML OpenRAIL-M license through HuggingFace, which permits commercial use while enforcing ethical guidelines. Users experiment with the generator for free through the DreamStudio interface, which provides 200 initial credits before charging a small fee for additional image generations.