Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

521–521 of 521

Sep 25, 2023Monday

OpenAI News

ChatGPT can now see, hear, and speak

OpenAI says ChatGPT now supports seeing, hearing, and speaking. The post body is empty, so it does not disclose model versions, rollout timing, regional limits, pricing, or API scope. The real watchpoints are voice latency, vision limits, and access paths.

Why it matters: OpenAI added voice conversation and image understanding to ChatGPT, with speech synthesis from a few seconds of real audio and image input built on multimodal GPT-3.5 and GPT-4.