At a press conference on Monday, May 13, 2024, OpenAI unveiled its latest language model, GPT-4o, which outperforms GPT-4 and is now available to all users. At the same time, OpenAI launched a new ChatGPT app for macOS.
AISYSNEXT, which closely follows the latest developments in artificial intelligence, walks you through the main GPT-4o announcements.
Here are the details.
GPT-4o: OpenAI’s new gem
Omnimodel: a major technological step forward
GPT-4o delivers greater speed and improved performance in certain areas, particularly voice and images. In addition, GPT-4o will soon be able to handle video, including real-time video.
Future outlook
In the future, improvements will enable more natural, real-time voice conversations, as well as the ability to interact with ChatGPT through real-time video. For example, you could show ChatGPT a live sports game and ask it to explain the rules.
Performance examples
On its page, OpenAI presents a few examples of what GPT-4o can do, notably in creating and iterating on visuals, with results that are often impressive. However, when we reproduced these examples, we did not always get results that were as satisfying.
Advanced speech recognition and image analysis capabilities
Technical comparison
In a technical comparison, OpenAI states that GPT-4o reaches levels similar to GPT-4 Turbo for text, reasoning and coding, but that it “sets new standards for multilingual, audio and vision capabilities.” On speech recognition, the results show that the error rate of GPT-4o is markedly lower than that of Whisper, OpenAI’s previous speech recognition model.
Omnimodel: a unique approach
OpenAI explains that the new omnimodel is a single model trained end to end across text, vision and audio, “which means that all inputs and outputs are processed by the same neural network.” By contrast, voice mode with GPT-3.5 and GPT-4 requires three different models, which causes delays and a loss of information.
Benefits of the new process
This process means that the main source of intelligence, GPT-4, loses a lot of information: it cannot directly observe tone, multiple speakers or background noise, and it cannot produce laughter or singing or express emotion, according to OpenAI.
GPT-4o available to all users
Availability
GPT-4o is currently available to users on the paid ChatGPT Plus and Team plans. Enterprise plan subscribers will have to wait a few more weeks. The new model is also built into the free version of the chatbot, but with a message limit up to five times lower than the one for ChatGPT Plus users.
Limited use for free users
The number of messages free users can send with GPT-4o will be limited based on usage and demand. When a user reaches the limit, ChatGPT automatically switches to GPT-3.5 so the conversation can continue, according to OpenAI.
Available features
Starting now, users of the free version of ChatGPT can try features previously reserved for paying subscribers, such as web access, data analysis, image analysis and custom chatbots. To try it, simply click GPT-3.5 or GPT-4 in the upper-left corner of the interface and choose GPT-4o.
Conclusion
GPT-4o marks a major step forward for language models. With its improved capabilities and its availability to a larger number of users, it ushers in a new era for AI applications.
Feel free to contact us at AISYSNEXT to discuss your web projects and how we can help you deliver them using the latest JavaScript technologies.
Put our AI-focused expertise to work by contacting us:




