Google DeepMind Showcases Multimodal AI Advances at NeurIPS 2023

Google DeepMind presents over 180 papers at NeurIPS 2023, highlighting multimodal AI that bridges language, video, and physical actions. The team also demos groundbreaking tools for weather forecasting, materials discovery, and AI watermarking.

Google DeepMind presents over 180 research papers at NeurIPS 2023 in New Orleans, showcasing a strong push toward more multimodal, robust, and general AI systems. The team highlights its largest and most capable AI model, Gemini, alongside live demonstrations of cutting-edge technology for global weather forecasting, materials discovery, and watermarking AI-generated content. This massive presence at the world's largest AI conference emphasizes their commitment to solving complex, real-world challenges.

A major focus of their research involves bridging the gap between different types of data, such as text, images, and video. The researchers demonstrate that diffusion models like Imagen classify images by recognizing shapes rather than textures, which closely mimics human visual processing. Additionally, they reveal how simply predicting captions from images significantly improves computer-vision learning, surpassing current methods and showing strong potential to scale up for future applications.

The team also introduces UniSim, a universal simulator of real-world interactions that generates realistic experiences in response to actions by humans, robots, and digital agents. By leveraging video generation and closed captioning, these models successfully transfer knowledge to guide actual robot actions. Researchers are also creating digital agents that navigate software interfaces using screenshots and mouse clicks exactly like a human would, paving the way for more helpful everyday assistants in both physical and digital worlds.

Read More at the original source →