OpenAI Launches InstructGPT to Fix Toxicity and Improve Instruction Following

OpenAI introduces InstructGPT, a fine-tuned version of GPT-3 that uses human feedback to reduce toxic outputs and follow instructions more accurately. While this marks a major step in AI alignment, the enhanced capability also introduces new risks for potential misuse.

OpenAI overhauls its GPT-3 language model and introduces InstructGPT as the new default tool to address complaints about toxic language and misinformation. By using reinforcement learning from human feedback, the company fine-tunes the AI to be safer, more helpful, and better at following specific instructions in English.

The development process involves a three-step procedure that includes supervised fine-tuning, the creation of a separate reward model, and additional reinforcement learning to align the system with human preferences. Although InstructGPT sometimes scores lower on traditional NLP benchmarks than GPT-3, it performs much better in real-world scenarios because it adapts directly to what humans actually want.

This release represents a significant breakthrough in solving the AI alignment problem, but it also brings a dangerous drawback. Because InstructGPT follows instructions so effectively, malicious users could potentially exploit this improved capability to generate harmful or deceptive content on a larger scale than the previous model allowed.

Read More at the original source →