imagen Google Researchers Unveil Imagen, a Highly Advanced Text-to-Image AI Model A new AI system called Imagen combines large language models with diffusion techniques to generate highly realistic images from text descriptions, outperforming competitors in human evaluations.
deepmind New Flamingo Visual Language Model Excels at Few-Shot Learning Tasks DeepMind introduces Flamingo, a visual language model that seamlessly processes interleaved images, videos, and text to master new tasks with only a handful of examples. The system outperforms heavily fine-tuned models across a wide range of visual benchmarks.
flamingo New Flamingo Visual Language Model Masters Few-Shot Learning Tasks Researchers introduce Flamingo, a visual language model that seamlessly processes interleaved images, videos, and text to achieve state-of-the-art few-shot learning. The model outperforms competitors by rapidly adapting to new tasks using only a handful of examples.
dall-e 2 Researchers Test DALL-E 2 Reasoning Skills With Challenging Prompts A new study by Gary Marcus, Ernest Davis, and Scott Aaronson puts DALL-E 2 to the test with difficult prompts to evaluate its common sense and reasoning abilities. The results show that while the AI sometimes succeeds, it struggles to consistently produce accurate images for complex text inputs.
google cloud Google Cloud Unveils New AI Tools to Simplify Business Workflows Google Cloud introduces Translation Hub, updated Document AI features, and a new Vertex AI computer vision tool to help everyday business users easily integrate AI into their workflows.
mit MIT Researchers Show Synthetic Data Can Rival Real Images for AI Training MIT researchers develop a generative model that creates synthetic data to train AI for image classification, potentially saving millions in dataset costs. The method sometimes outperforms models trained on real data.
john deere John Deere Autonomous 8R Tractor Relies on AI for Safe Farming John Deere utilizes deep learning and computer vision to operate its new fully autonomous 8R Tractor safely. The machine uses advanced perception technology and a robust control stack to navigate fields without human intervention.
google Google Releases TensorFlow 3D to Advance Spatial Scene Understanding Google introduces TensorFlow 3D, a new open-source library designed to help researchers process 3D sensor data for applications like autonomous driving and augmented reality.
google Google Releases Open AI Models With Emotion Detection, Alarming Experts Google launches the PaliGemma 2 family of open AI models that claim to identify emotions in images. Critics warn the technology relies on flawed science and risks spreading harmful biases.
advex ai Advex AI Launches to Tackle Manufacturing Vision Data Scarcity Advex AI officially launches with $3.6 million in funding to help manufacturers train computer vision systems using synthetic data. The startup generates thousands of realistic images from small sample sets to solve critical AI training bottlenecks.
openai OpenAI Launches GPT-4o Model With Real-Time Voice and Vision OpenAI introduces GPT-4o, an omni-modal AI model that processes text, speech, and video simultaneously. The upgrade brings real-time voice conversations and enhanced vision capabilities to ChatGPT.
openai OpenAI Expands GPT-4 Vision Access Despite Ongoing Safety Questions OpenAI opens GPT-4 with vision to all developers through the new GPT-4 Turbo API. Independent researchers praise the model's accuracy but highlight unresolved safety and privacy concerns.
google deepmind Google DeepMind Showcases Multimodal AI Advances at NeurIPS 2023 Google DeepMind presents over 180 papers at NeurIPS 2023, highlighting multimodal AI that bridges language, video, and physical actions. The team also demos groundbreaking tools for weather forecasting, materials discovery, and AI watermarking.
openai GPT-4 Powers Be My Eyes Virtual Volunteer for Visually Impaired Users OpenAI's latest model, GPT-4, makes its debut in Be My Eyes as a virtual volunteer that analyzes images to provide instant visual assistance. The new feature helps blind and low-vision users complete everyday tasks like reading labels, navigating gyms, and cooking without waiting for a human helper.
pixyle ai Pixyle AI Raises $1 Million to Fix E-Commerce Product Search Pixyle AI uses advanced computer vision to automatically tag and categorize fashion items, helping online retailers reduce cart abandonment. The startup just secures a €1 million seed round to expand its B2B visual search technology.