Google Unveils 540-Billion Parameter PaLM to Surpass Human NLP Benchmarks
Google introduces PaLM, a massive 540-billion-parameter language model trained on over 6,000 TPU chips that outperforms human baselines on key reasoning and language evaluations.
Google Research unveils the Pathways Language Model (PaLM), a massive 540-billion-parameter AI system that outperforms average human performance on the BIG-bench benchmark. The model breaks current records on 28 out of 29 natural language processing tasks and demonstrates impressive abilities in logical inference and joke explanation. By utilizing an autoregressive decoder-only Transformer architecture, PaLM pushes the boundaries of what large-scale language models can achieve.
To train this enormous system, Google leverages the largest TPU cluster known to date, which consists of 6,144 chips running on their custom Pathways infrastructure. This advanced training setup allows the model to handle the immense computational demands that typically prevent such massive architectures from fitting on a single accelerator. The integration of a new chain of thought prompting method further boosts PaLM's ability to reason through complex problems step-by-step.
Developers Sharan Narang and Aakanksha Chowdhery emphasize that PaLM represents a major step toward Google's ultimate goal of creating a single AI system capable of generalizing across millions of tasks. By successfully combining massive scaling capabilities with innovative architectural choices, this model proves that future AI systems can understand diverse data types with remarkable efficiency. This breakthrough lays essential groundwork for the next generation of highly adaptable artificial intelligence.