Grok added a video input path, but the product boundary is still missing. The post claims full-video uploads plus real-time analysis, summarization, translation, scene explanation, and context extraction. It gives no duration cap, supported formats, latency target, batching, API access, or enterprise rollout. For practitioners, those details matter far more than “native multimodal.”
I don’t buy the “understands full video” framing yet. Gemini and GPT-4o-class systems already pushed video or frame-sequence understanding into demos and product surfaces. The hard part is long-video sampling, audio-video alignment, timestamped citations, and cost. If Grok only works well on X-native clips, it is a consumer feature. If it handles long meetings, surveillance footage, lectures, and grounded time references, then it starts to matter in workflows.