Google Unveils PaLI, a Multilingual AI for Images and Text

Google demonstrates PaLI, a new AI model that understands 109 languages to perform complex tasks involving both text and images. Meanwhile, other advancements showcase highly emotional text-to-speech generation and new tools protecting artists from unauthorized AI training.

Google researchers introduce PaLI, an AI system that understands 109 languages to perform tasks combining text and images. Trained on both visual and linguistic data, this model handles complex duties like image captioning, object detection, and optical character recognition across a massive variety of languages.

Advancements in speech synthesis also make headlines as Play.ht reveals a new text-to-speech model that produces highly emotional and realistic audio. While still in early stages and not quite ready for full-length audiobook narration, the impressive quality points to expanding practical applications as the technology continues to mature.

In response to growing concerns over data usage, a Berlin-based group launches Source+ to give visual artists, musicians, and writers control over whether their work trains AI systems. This opt-in and opt-out framework addresses the ongoing tension between content creators and the machine learning models that scrape their intellectual property.

Read More at the original source →