OpenAI added vision, listening, and speaking to ChatGPT on September 25, 2023. My read is straightforward: treat this first as a product-distribution move, then as a model story. The headline is huge, but the body here is empty. We do not get model names, rollout timing, latency, pricing, regional limits, or API scope. Without those conditions, “see, hear, and speak” is a very incomplete claim.
I’ve always thought multimodal launches live or die on latency more than on demos. Stitching ASR, TTS, and vision into one surface is not the hard part anymore; making it feel conversational is. If voice round-trip latency is bad, users stop treating it like an assistant and start treating it like a voice memo box. This post does not disclose whether audio is streamed, how fast first token arrives, or whether voice is full duplex. Same issue on vision. “Can see” tells you almost nothing unless you know whether it handles OCR, charts, UI screenshots, dense documents, or only simple image Q&A.
In 2023 context, this looked like OpenAI closing the interface gap around ChatGPT. Around that period, the market was moving from “text chatbot” toward “always-available assistant” on mobile. I remember GPT-4V had already been introduced around then, so this did not read like a brand-new research unlock. It read like OpenAI packaging existing model components into the consumer product shell that mattered most. That distinction matters. Distribution through camera and voice changes retention and usage frequency fast, even if the underlying stack is still a composed system rather than one elegant native multimodal model.
My pushback is on the framing. Putting see, hear, and speak in one line invites people to infer a unified multimodal model experience. The article, at least in the material provided here, does not say that. It does not tell us whether this is one model, multiple models chained together, or some hybrid orchestration. That is not a technical footnote. It determines cost, latency, failure modes, and how durable the advantage is. If the stack was basically Whisper plus TTS plus GPT-4V-style vision routing, that is still a strong product step, but it is not the same thing as proving native multimodal maturity.
So for practitioners, I would ignore the headline glow and ask four operational questions instead: who gets access, on which surfaces, at what voice latency, and with what hard vision limits. Until those are disclosed, this announcement says more about OpenAI trying to own the user interface than about multimodality being solved.