Meta Unveils Text-to-Video AI Amid High Computing Costs

Meta introduces Make-A-Video, a new AI system that generates short video clips from text prompts. The technology builds on recent text-to-image AI breakthroughs but faces significant computational and ethical hurdles.

Meta unveils Make-A-Video, an artificial intelligence system that generates five-second video clips from simple text prompts. Users type descriptive phrases, such as a dog flying in a superhero cape, and the AI produces a video that looks like a trippy old home movie. This development represents the next logical step beyond the text-to-image AI systems that dominate the tech industry this year.

Creating these videos requires massive computational power because a single short clip demands hundreds of images. This heavy computational lift means only large tech companies currently have the resources to build such systems. Additionally, researchers face a challenge with training data, as large-scale data sets pairing high-quality videos with descriptive text simply do not exist yet.

To overcome this data limitation, Meta combines three open-source image and video data sets to train the model. Labeled still images teach the AI to recognize objects, while video databases demonstrate how those objects move in the real world. Experts note that the resulting model shows a promising understanding of 3D shapes, camera rotation, and depth, even though the technology remains strictly in the research phase and is not yet available to the public.

Read More at the original source →