Synthetic Data Fills Critical Gaps in AI Training

Artificial intelligence developers increasingly rely on synthetic data to train algorithms when real-world information is scarce or too sensitive to use. While this computer-generated data promises to reduce bias, it ultimately remains limited by the original data used to create it.

Artificial intelligence developers increasingly turn to synthetic data to train machine learning models when real-world information is unavailable or too sensitive to use. By using AI to generate artificial datasets that share the same statistical properties as genuine data, researchers overcome significant limitations, such as the recent creation of a much-needed African fashion dataset to address a glaring gap in existing computer-vision training materials.

A wave of startups and universities now makes this technology widely accessible to the public. Companies like Datagen and Synthesis AI supply digital human faces on demand, while other firms focus on synthetic data for finance and insurance. Meanwhile, MIT's Data to AI Lab offers open-source tools through the Synthetic Data Vault to help developers create a wide variety of synthetic data types.

This rapid expansion relies heavily on generative adversarial networks (GANs), which excel at producing realistic but entirely fake examples. Proponents argue that synthetic data helps avoid the widespread bias found in many traditional datasets, but critics warn that the output is only as fair as the input. For instance, a GAN trained on a limited number of Black faces produces less lifelike synthetic Black faces, meaning the underlying bias simply perpetuates in a new form.

Read More at the original source →