OpenAI Builds CriticGPT to Catch Code Errors Human Trainers Miss

OpenAI introduces CriticGPT, a new AI model designed to help human trainers spot code mistakes in ChatGPT responses. This tool addresses the growing challenge of human workers struggling to evaluate increasingly advanced AI outputs.

OpenAI introduces a new AI model called CriticGPT to assist human trainers in catching code errors made by ChatGPT. As generative AI models become more advanced through Reinforcement Learning from Human Feedback, human workers increasingly struggle to identify flawed answers when the chatbot knows more than they do. CriticGPT steps in as a supplementary tool to bridge this widening knowledge gap.

The system operates as an augmentation tool rather than an autonomous feedback loop, meaning it simply helps human evaluators do their jobs more effectively. OpenAI reports that people assisted by CriticGPT outperform those without the AI helper 60 percent of the time when reviewing ChatGPT code. This approach proves especially necessary given that OpenAI relies on low-paid crowdsourced workers who often lack advanced computer science expertise.

According to OpenAI's recently published research paper, large language models catch substantially more inserted bugs than qualified humans paid for evaluation. By deploying CriticGPT based on GPT-4, the company aims to maintain the quality of its model refinement process even as its chatbots outshine their human teachers. This development highlights the emerging industry trend of using AI to evaluate and improve other AI systems.

Read More at the original source →