Meta's Make-A-Video AI Generates Creepy State-of-the-Art Clips From Text
Meta unveils Make-A-Video, a new text-to-video AI system that combines image diffusion with unsupervised video learning to achieve state-of-the-art results. Despite the impressive technical leap, the generated clips feature a distinctly surreal and nightmarish quality.
Meta researchers introduce Make-A-Video, a new artificial intelligence system that generates short video clips directly from text prompts. The system builds on existing image diffusion techniques, working backward from visual static to create a target image, and pairs this with unsupervised training on unlabeled video data to understand sequential frames.
The AI effectively combines its knowledge of realistic image generation with an understanding of video motion without requiring explicit human guidance on how to merge the two. This approach sets a new state-of-the-art in text-to-video generation, outperforming previous systems in spatial and temporal resolution, faithfulness to text, and overall quality.
Despite the impressive technical leap, the resulting videos possess a distinctly surreal and nightmarish quality. The motion resembles stop-animation, and visual artifacts give the objects a strange, furry texture that makes the clips feel both dreamlike and deeply unsettling.