OpenAI Develops New Method to Reduce GPT-3 Bias and Toxicity
OpenAI publishes a study claiming a simple new approach mitigates bias, toxicity, and harmful outputs in its GPT-3 language model. The technique allows developers to guide a model's tone and ethical behavior using specific prompts and small datasets.
OpenAI releases a new study detailing a simple method to improve the ethical behavior of GPT-3. The technique gives developers the ability to dictate the tone and personality of the language model based on specific prompts. By using a specialized dataset called PALMS, the lab aims to mitigate harmful outputs tied to gender, race, and religious prejudices that exist within the model's training data.
Large language models historically struggle with amplifying systemic biases found in their training data. OpenAI notes that unmitigated models frequently place derogatory words near female pronouns and falsely associate certain religions with terrorism. In extreme cases, previous tests show that medical chatbots powered by GPT-3 respond to vulnerable patients by encouraging self-harm.
External researchers praise the new method for its surprising simplicity and high efficiency with small datasets. However, OpenAI cautions that there is no universal standard for appropriate model behavior, as desirable outputs differ heavily depending on the social context. The lab also acknowledges that forcing models to adhere to strict linguistic norms risks alienating minority speakers and discouraging their engagement with the technology.