Tech Giants Turn to Synthetic Data to Train Next-Generation AI Models
Major tech companies like OpenAI and Meta are increasingly relying on synthetic data to train and fine-tune their AI models. While this approach reduces reliance on expensive human annotators, it introduces new risks related to AI hallucinations and potential model collapse.
Major tech companies increasingly embrace synthetic data to train and fine-tune their artificial intelligence models. OpenAI recently unveils a new ChatGPT feature called Canvas, which relies on a customized version of GPT-4o fine-tuned entirely with synthetic outputs distilled from its o1-preview model. This strategy allows OpenAI to rapidly improve the model and enable new user interactions without depending on costly human-generated data.
Meta follows a similar path by using synthetic captions generated by its Llama 3 models to develop Movie Gen, its new suite of AI video creation tools. Although human annotators step in to correct errors and add detail to these automated captions, the bulk of the data groundwork remains largely automated. Industry leaders like OpenAI CEO Sam Altman predict that AI will eventually produce high-quality synthetic data capable of training future models entirely on its own.
However, this synthetic-data-first approach carries significant risks that developers must carefully navigate. Because the models generating the synthetic data inevitably hallucinate and contain inherent biases, these flaws directly transfer into the new training data. To prevent dangerous model collapse where an AI system loses its creativity and accuracy, companies must rigorously curate and filter synthetic data just as they do with human-generated information.