- ChatGPT-maker OpenAI has announced that it had enabled its GPT-3.5 and GPT-4 models to study images and analyse them in words, while its mobile apps will have speech synthesis so that people can have full-fledged conversations with the chatbot.
- The Microsoft-backed company had promised multimodality during the release of GPT-4.
- Google’s new yet-to-be-released multimodal large language model called Gemini, is already being tested in a bunch of companies.
- OpenAI is also reportedly working on a new project called Gobi which is expected to be a multimodal AI system from scratch, unlike the GPT models.
|