OpenAI Enters Video Generation Market With Impressive Sora Model
OpenAI unveils Sora, a new AI model that generates high-definition videos up to a minute long from text prompts or still images. While the results show remarkable progress over competitors, the model still struggles with complex physics and spatial consistency.
OpenAI launches Sora, a new generative AI model that creates high-definition video from text prompts or still images. The model produces 1080p movie-like scenes up to a minute long, featuring multiple characters, various types of motion, and detailed backgrounds. Sora also extends existing video clips by filling in missing details automatically.
The generated samples appear significantly more advanced than current text-to-video technologies from competitors like Runway, Google, and Meta. Sora demonstrates a strong grasp of language and physical world concepts, allowing it to output videos in multiple styles such as photorealistic, animated, and black-and-white while maintaining better coherence and fewer physically impossible movements than previous systems.
Despite the impressive demonstrations, Sora is not without noticeable flaws. Some videos featuring humanoids look overly synthetic or lack background activity, and common AI artifacts still appear, such as objects moving in impossible directions or limbs merging with objects. OpenAI openly admits that the model struggles with accurately simulating complex scene physics and occasionally confuses spatial details like left and right.