Meta Unveils Data2vec AI to Process Speech, Images, and Text
Meta introduces data2vec, a single AI model trained to understand speech, images, and text simultaneously. The company plans to integrate this multi-modal technology into future augmented reality and metaverse products.
Researchers at Meta develop a new AI algorithm named data2vec that processes speech, images, and text using a single self-supervised model. Unlike traditional AI systems that focus on one specific type of data, this multi-modal approach recognizes speech in audio, classifies objects in images, and analyzes grammar or emotions in text.
Meta views this technology as a crucial step toward blending physical and digital environments for the metaverse. CEO Mark Zuckerberg explains that systems like data2vec help computers understand the world through a combination of sight, sound, and words, which could eventually power AI assistants built into augmented reality glasses.
The transformer-based neural network learns by predicting missing representations across all three data formats, guessing the next group of pixels, speech utterances, or words. To train this versatile algorithm, researchers utilize a mix of Nvidia V100 and A100 GPUs to process 960 hours of speech audio, millions of text words, and ImageNet-1K images.