USC Study Reveals Widespread Bias in AI Commonsense Databases

USC researchers discover that between 3.4% and 38.6% of "facts" in popular AI commonsense databases contain human biases. These flawed foundations threaten the fairness of virtual assistants, chatbots, and automated content generators.

Researchers from the USC Information Sciences Institute examine two popular commonsense knowledge databases, ConceptNET and GenericsKB, to determine if the data is fair. These databases provide the foundational facts that artificial intelligence systems use to think and respond like humans. Unfortunately, the team finds that human biases heavily taint these supposedly neutral repositories of common knowledge.

The study utilizes COMeT, a knowledge graph completion algorithm, to analyze the information and measure the extent of the prejudice. Depending on the specific database and the metrics applied, the researchers discover that 3.4% of ConceptNET and up to 38.6% of GenericsKB data contains biases. This means a significant portion of the "facts" fed into AI systems unfairly stereotypes or marginalizes people based on race, gender, sexuality, or nationality.

Because tech companies widely use these databases to power virtual assistants like Siri and Alexa, as well as marketing copywriting tools and media bots, the biased data directly impacts everyday consumers. The USC team highlights that AI developers blindly trust these crowdsourced datasets, assuming they represent pure scientific truth rather than human prejudice. This underlying contamination ultimately causes the resulting AI technologies to generate unfair and discriminatory outputs.

Read More at the original source →