Meta Releases Open-Source AI Model Capable of Translating 200 Languages
Meta AI open-sources NLLB-200, a massive 54.5B parameter model that outperforms state-of-the-art systems in translating over 200 languages. The release includes an expanded FLORES-200 benchmark to evaluate machine translation across low-resource languages.
Meta AI releases NLLB-200, an open-source artificial intelligence model that translates between more than 200 languages. This massive 54.5 billion parameter Mixture of Experts model trains on over 18 billion sentence pairs and outperforms current state-of-the-art translation systems by up to 44% on benchmark evaluations.
The No Language Left Behind project focuses specifically on supporting low-resource languages that possess fewer than one million publicly available translated sentences. To build this robust system, Meta researchers collect multilingual training data by hiring professional human translators and mining text directly from the web.
Alongside the translation model, Meta updates and open-sources the FLORES-200 benchmark dataset to evaluate machine translation in over 40,000 directions. This release continues Meta's ongoing effort to break down language barriers, building on previous milestones like the LASER library and the M2M-100 multilingual translation model.