Google Unveils Gemini, a Multi-Modal AI Model Surpassing Human Experts
Alphabet releases Gemini, a new multi-modal AI model developed by Google DeepMind that outperforms human experts on key language understanding benchmarks. The model launches in three distinct sizes to power various Google products and developer tools.
Alphabet releases Gemini, a next-generation multi-modal AI model developed by Google DeepMind. This new system makes history as the first AI model to outperform human experts on the Massive Multitask Language Understanding (MMLU) benchmark, achieving a score above 90 percent. CEO Sundar Pichai highlights that Gemini outperforms OpenAI's ChatGPT and achieves state-of-the-art results in 30 out of 32 leading academic benchmarks.
Gemini stands out because of its native multi-modal capabilities, allowing it to seamlessly understand, generate, and reason across text, images, and code simultaneously. The model's architecture prioritizes efficiency and scalability, enabling rapid integration with existing developer tools and APIs. This highly adaptable design fosters collaboration across the AI community and drives future innovation.
Google rolls out Gemini in three distinct sizes to serve different computing needs. The largest version, Gemini Ultra, targets highly complex tasks, while the medium-sized Gemini Pro currently powers Google's Bard chatbot. Meanwhile, the highly efficient Gemini Nano runs directly on mobile devices, making its debut on the Google Pixel 8 Pro smartphone, though early user reactions remain mixed regarding occasional hallucinations.