GPT-4 Launch Delivers Multimodality But Leaves Key Details Secret
OpenAI officially launches GPT-4 with highly anticipated multimodal capabilities, though the initial release only supports image and text inputs. The company keeps the exact parameter count of the new model completely hidden from the public.
OpenAI releases GPT-4 to the public, sparking global excitement while leaving several critical anticipated features unaddressed. The biggest upgrade is the introduction of multimodality, which allows the model to process both image and text inputs to generate text outputs for tasks like dialogue systems and translation.
Despite previous hints that the model would handle video and audio, the developer demo only showcases image integration. OpenAI president Greg Brockman notes that this image feature is merely a preview and is not yet publicly available, though it demonstrates an impressive ability to understand visuals and explain why an image is funny or read handwritten notes.
OpenAI completely avoids discussing the parameter size of GPT-4, leaving the rumor of a massive 100-trillion-parameter model unresolved. Although CEO Sam Altman previously dismissed that specific number, the lack of official transparency about the model's actual capacity keeps tech enthusiasts guessing about its true scale.