Are ‘visual’ AI models actually blind?
New multi-modal language models, such as GPT-4o and Gemini 1.5 Pro, may not truly understand images and audio as expected.
MAIN POINTS
- Latest language models are described as multi-modal.
- They are expected to understand images, audio, and text.
- A study suggests these models might not genuinely comprehend visual and auditory data.
TAKEAWAYS
- Multi-modal capabilities of new models are under scrutiny.
- Understanding of non-text data by these models is questionable.
- Further research is needed to validate their multi-modal claims.