OpenAI's Whisper Transcription Tool Found to Fabricate Text Regularly
Researchers discover that OpenAI's Whisper audio transcription tool frequently hallucinates, adding false information like racial commentary and fake medical treatments to transcripts.
Software engineers and academic researchers raise serious concerns about OpenAI's Whisper transcription tool, which frequently fabricates text instead of accurately converting audio to words. Unlike typical AI chatbots where users expect some creative generation, a transcription tool is supposed to strictly follow the spoken audio, making these errors particularly alarming.
Experts report that Whisper invents highly problematic content, including racial commentary and imagined medical treatments. A University of Michigan researcher finds hallucinations in 80 percent of public meeting transcriptions, while other developers observe false information in more than half of the over 100 hours of audio they analyze.
These persistent inaccuracies pose a significant threat as organizations integrate Whisper into hospitals and other critical medical environments. OpenAI acknowledges the issue, stating that the company continually works to improve model accuracy and noting that its usage policies explicitly prohibit utilizing Whisper in high-stakes decision-making contexts.