Apple Researchers Unveil State-of-the-Art Multimodal AI Model MM1
Apple researchers publish a new paper demonstrating breakthroughs in training AI on both text and images. The MM1 model achieves top performance as the company invests heavily to catch up in the generative AI race.
Apple researchers achieve a significant breakthrough in artificial intelligence with the unveiling of MM1, a new multimodal large language model. Detailed in a recently published research paper, MM1 demonstrates state-of-the-art performance by training on a careful mix of image captions, interleaved image-text data, and text-only information. This diverse training approach allows the AI system to excel at complex tasks like visual question answering, image captioning, and natural language inference.
The study reveals that scaling the visual components of AI models plays a critical role in overall performance. Researchers find that the choice of image encoder, image resolution, and image token count heavily influence the model's capabilities, while the design of the vision-language connector is less important. Furthermore, the largest 30 billion parameter version of MM1 exhibits impressive in-context learning, successfully performing multi-step reasoning across multiple images using chain-of-thought prompting.
This research emerges as Apple aggressively ramps up its investments in artificial intelligence to compete with industry rivals like Google and Microsoft. The tech giant reportedly spends around $1 billion annually on AI development and is actively working on an internal large language model framework called Ajax and a chatbot known as Apple GPT. These efforts ultimately aim to integrate advanced generative AI features directly into future Apple products and software ecosystems.