Researchers introduce Flamingo, a visual language model that seamlessly processes interleaved images, videos, and text to achieve state-of-the-art few-shot learning. The model outperforms competitors by rapidly adapting to new tasks using only a handful of examples.