USC Researchers Discover Heavy Bias in Common Sense AI Databases
A USC Information Sciences Institute team finds that 3.4% to 38.6% of facts in popular AI databases contain human biases, threatening the fairness of algorithms used by virtual assistants and chatbots.
Researchers from the USC Information Sciences Institute examine popular commonsense knowledge databases that train artificial intelligence algorithms to think like humans. These databases, such as the widely used ConceptNET, provide foundational facts to power virtual assistants like Siri and Alexa, as well as auto-generated media content and marketing copywriting. Because these systems rely on this data to interact with people, the information must remain completely free of stereotypes to ensure fair treatment across all races, genders, and nationalities.
The research team uses a knowledge graph completion algorithm called COMeT to analyze the human-curated data in ConceptNET and a smaller database called GenericsKB. Their investigation reveals that a significant portion of these accepted "facts" actually reflects underlying human prejudices. Depending on the specific database and the evaluation metrics applied, the team finds that biased data makes up anywhere from 3.4% to 38.6% of the total content.
These findings highlight a fundamental flaw in the foundation of many AI systems, as machines inherently adopt the biases present in their training data. Lead researcher Fred Morstatter points out that developers routinely download and integrate these crowdsourced resources without fully understanding the extent of the hidden prejudices they contain. Addressing this widespread data bias is an essential step toward creating equitable artificial intelligence that treats all individuals fairly.