Google Launches Gemini 1.0 to Rival GPT-4 Across Text, Image, and Video

Google unveils Gemini 1.0, a natively multimodal AI model that outperforms human experts on key benchmarks and powers the updated Bard. The system comes in three sizes to run on everything from smartphones to massive data centers.

Google unveils Gemini 1.0, its most capable and general AI model that natively understands text, code, audio, images, and video. Unlike previous multimodal models that stitch separate components together, Gemini pre-trains on different modalities from the start using Google's TPU v5p hardware. This native approach gives the model sophisticated reasoning capabilities, allowing it to digest massive amounts of data, such as summarizing 200,000 scientific papers in about an hour.

The new model debuts in three distinct sizes to serve various computing needs. Gemini Ultra handles highly complex tasks in data centers, Gemini Pro scales across a wide range of tasks, and Gemini Nano runs efficiently on mobile devices. Google integrates Gemini Pro into Bard today, bringing these advanced capabilities directly to consumers.

In benchmark tests, Gemini Ultra surpasses OpenAI's GPT-4 in text-based reasoning, math, and code evaluations. Notably, it becomes the first model to outperform human experts on the MMLU benchmark with a score of 90.0%. Gemini Ultra also beats GPT-4V across image, audio, and video tests without relying on external text-extraction tools.

Read More at the original source →