Massive 204-Task Benchmark BIG-bench Challenges Large Language Models
A collaborative team of 444 researchers introduces BIG-bench, a diverse 204-task benchmark designed to test the true limits of large language models. The project evaluates models ranging from millions to hundreds of billions of parameters to predict their future transformative effects.
A massive collaborative effort involving 444 authors from 132 institutions introduces BIG-bench, a highly difficult and diverse benchmark designed to evaluate large language models. Named in homage to Alan Turing's imitation game, this comprehensive suite contains 204 tasks that current AI models cannot fully solve. The benchmark aims to predict the potentially transformative effects of these rapidly advancing systems.
BIG-bench evaluates dense and sparse transformer models spanning a massive scale from millions to hundreds of billions of parameters. The suite includes a wide variety of novel tasks covering diverse topics and languages, alongside a streamlined subset called BIG-bench Lite for faster evaluation. It supports both JSON and programmatic task types to thoroughly test model capabilities.
The research team provides detailed evaluation results across models of varying sizes and establishes baseline scores using human evaluators. By offering an open API and tools for lightweight task creation, BIG-bench gives the AI community a crucial resource for understanding the current limitations and near-future potential of large language models.