Google Unveils Multimodal Gemini AI to Rival OpenAI

Google introduces Gemini, a new family of AI models that processes text, images, audio, and video to compete with OpenAI. The technology currently powers an updated Bard chatbot and runs directly on the latest Pixel smartphones.

Google unveils Gemini, its most advanced class of transformer-based AI models that processes text, images, audio, and video simultaneously. This new multimodal system features a 32k context window and comes in three distinct sizes to handle various computing needs. The largest version, Gemini Ultra, tackles complex reasoning tasks, while the mid-tier Gemini Pro focuses on efficient, broad performance, and the smallest Gemini Nano operates directly on mobile devices.

The tech giant immediately integrates Gemini Pro into its Bard chatbot, replacing the older PaLM 2 language model to improve text understanding and summarization. Although Gemini boasts impressive multimodal capabilities, the current Bard update only processes and generates English text. Google plans to upgrade other core services, including Search, Ads, Chrome, and productivity apps like Gmail and Google Docs, with Gemini Pro in the coming months.

For on-device tasks, the Pixel 8 Pro smartphone utilizes Gemini Nano to summarize audio recordings and generate quick text message replies. Google intends to expand these on-device AI features and empower third-party Android developers through a new service called AICore. Running on Android 14, AICore provides safe access to the Nano model via open-source APIs so developers can build their own intelligent mobile applications.

Read More at the original source →