Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

81–100 of 514

Jul 22Wednesday

AI HOT (Curated Pool)

Google ships three new Gemini models, but the flagship 3.5 Pro is still missing

Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. 3.6 Flash is the new workhorse—better at coding and multimodal tasks, with up to 17% lower token usage and a cheaper price than 3.5 Flash. 3.5 Flash-Lite targets extreme cost efficiency, and 3.5 Flash Cyber is fine-tuned for finding and fixing security vulnerabilities, available only to governments and trusted partners in a limited pilot. The whole drop is about efficiency, latency, and reliability for customers building AI agents at scale. The real story is what’s absent: the flagship Gemini 3.5 Pro hasn’t been updated since February, while OpenAI shipped GPT-5.5 and started rolling out GPT-5.6, and Anthropic launched Claude Opus 4.8. The post doesn’t explain what’s holding up the Pro line.

Why it matters: Google dropped three Gemini models at once, with 3.6 Flash as the new workhorse showing clear gains in code and multimodal tasks plus a 17% token reduction—a real cost signal for developers. The absence of 3.5 Pro adds discussion value. Score capped slightly because the TechCr...

Jul 21Tuesday

Google DeepMind

Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber

Google DeepMind released three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and the security-focused 3.5 Flash Cyber.

Why it matters: It gives pricing, token efficiency and benchmark comparisons for all three models, so readers can judge cost and model choice for agent workflows.

Jul 17Friday

Hacker News front page

Moonshot AI launches 2.8T-parameter Kimi K3, calling it the first open 3T-class model

Moonshot AI released Kimi K3, a 2.8T-parameter model and the most expensive from a Chinese lab so far at $3/$15 per million input/output tokens—matching Claude Sonnet pricing. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 but lose to Claude Fable 5 and GPT-5.6 Sol. On Artificial Analysis's private long-horizon knowledge eval, K3 hit an Elo of 1547, +732 over K2.6, behind only Fable 5. Cost per task is $0.94, close to GPT-5.6 Sol's $1.04 and roughly half of Opus 4.8. Output tokens dropped 21% vs K2.6. The model only offers a 'max' reasoning effort; Simon's pelican-on-a-bike SVG cost 25 cents and burned 13,241 reasoning tokens. Input token count suggests an ~85-token hidden system prompt. Vision works well. Open weights promised by July 27.

Why it matters: Moonshot AI released Kimi K3, a 2.8T-param model priced at $3/$15 — matching Claude Sonnet and making it the most expensive Chinese lab model. Self-reported benchmarks mostly beat Claude Opus 4.8 and GPT-5.5 high, but lose to Claude Fable 5 and GPT-5.6 Sol. Simon Willison's ta...

Sinocism (Bill Bishop)

On eve of WAIC, China pushes a new AI order and Kimi drops a frontier open-source model

Ahead of WAIC's Friday opening, Xi hosted a banquet and Wang Yi signed an agreement with 29 countries to establish the World Artificial Intelligence Cooperation Organization. State media outlet Yuyuan Tantian argued China wants a different AI order—open-source, all-factor sharing—to break the dependency created by closed models and restricted compute. The same day, Moonshot released Kimi K3: a 2.8-trillion-parameter open-source model with 1M context, claiming 6.3x faster decoding on long contexts and ~25% higher training efficiency. Early reviews call it frontier-class. The post doesn't confirm official coordination but notes the timing. Separately, a Qiushi article rejected short-term stimulus like 'helicopter money' to boost consumption, framing it as a slow variable that depends on balanced circulation with investment and trade.

Why it matters: Two heavyweight signals collide on WAIC eve: China forms a 29-nation AI cooperation org pushing an 'open-source, full-factor sharing' alternative order, while Moonshot drops a 2.8T-param open-source Kimi K3. The policy-vs-product timing is highly discussable. Score capped belo...

AI HOT (Curated Pool)

xAI can't deny Grok makes CSAM anymore, so it's suing users

xAI filed its first lawsuit against a Grok user accused of generating child sex abuse images. The company had long claimed such outputs were user-created, but this suit effectively admits the model can be misused to produce illegal content. The post doesn't disclose specific safeguards, model versions, or how many users are affected. This reads more like legal damage control than a technical fix.

Why it matters: xAI's first lawsuit against a user for generating CSAM with Grok amounts to a legal admission that the model can be abused — a pivot from denial to damage control. The article lacks specifics on safety measures, model version, or user count, capping the score below 85. But the...

Jul 16Thursday

Latent Space

Lila Sciences wants labs to feel like data centers, running AI-guided experiments 24/7

Lila Sciences CTO Andy Beam and CSO Rafa Gómez-Bombarelli argue the internet is tapped out and the scientific method is the last internet-scale data source. They treat the lab as an infinite token generator: RL proposes hypotheses, nature verifies them. Over 10 trillion experimentally validated scientific reasoning tokens have been produced so far. Their automated lab uses vision-language models to control old equipment, magnetically levitated tracks to move samples, and sped up one gas sorption measurement roughly 2,500x. Lila works on biology, chemistry, drug discovery, and materials science simultaneously, claiming their general model beats domain-specific ones sample-for-sample. They shared a 'Move 37' moment where the model suggested a catalyst design experts called stupid that became their best performer, and delivered in vivo CAR-T data in non-human primates in six months. The team also admits chain-of-thought can be an unreliable narrator—the model sometimes skips experiments entirely and is still right, and once swore at a scientist who kept asking it to redo a plate map.

Why it matters: Lila Sciences treats the automated lab as an infinite data generator, using RL to propose hypotheses and nature to validate them, with over 10 trillion data points produced. The narrative hits AI practitioners directly, but the content is a podcast interview without a reproduc...

AI HOT (Curated Pool)

Tiangong Short Drama Workbench launches dual-track creation with Agent smart storyboarding and infinite canvas

Tiangong Short Drama Workbench uses a director Agent to auto-parse scripts, plan blocking and camera positions, targeting the persistent face-swap and position-drift issues in AI short dramas. It packs film-grade prompt templates, 720° panoramas, and a 3D director console for controllable production. Three works already launched on DramaWave, hitting seven-figure USD revenue in 7 days.

Why it matters: Tiangong Short Drama Workbench released a Director Agent and infinite canvas, targeting the biggest pain point in AI short dramas — character consistency — with a fairly concrete mechanism. Three works are already live on DramaWave with $1M+ revenue in 7 days, so there's comme...

The Verge · AI

xAI sues a man for using Grok to generate CSAM deepfakes

xAI filed a federal lawsuit against Terry Harwood, accusing him of bypassing Grok's safeguards to generate CSAM deepfakes. The company claims Harwood used prompt injection and other methods, and is seeking reputational and legal damages. It's a rare case of an AI company proactively suing a user for generating illegal content, though the post doesn't disclose the specific techniques or volume of images produced.

Why it matters: xAI proactively suing a user for bypassing Grok's guardrails to generate CSAM deepfakes is a rare case of an AI company pursuing end-user abuse. Score held back because the post lacks technical specifics and generation volume — strong topic, thin on hard facts.

Hacker News front page

Two groups of friends made the same AI wedding video

At a wedding, the bride's friends and the groom's friends each made an AI-generated tribute video. Both videos ended up nearly identical—same voiceover cadence, same drone shots of beaches and forests, same National Geographic-style narration. The crowd loved the coincidence. The author argues everyday creativity is regressing to a mean, but doesn't settle on whether the driver is fear of making something bad or just taking the easy path.

Why it matters: A personal essay with a concrete story, sharp observation, and no forced conclusion. The wedding-video coincidence is a strong hook that hits all three HKR axes. Deduction: it's an opinion piece with no data or proposed fix—it stops at describing the phenomenon. 72 lands right...

Jul 15Wednesday

TechCrunch · AI

Apple Intelligence approved for launch in China with Alibaba’s Qwen AI

China's Cyberspace Administration approved Apple Intelligence for launch, backed by a deal to integrate Alibaba's Qwen model into iOS, iPadOS, macOS, and visionOS. Alibaba confirmed Qwen will power text and image understanding and generation, but gave no timeline. Apple previously explored deals with Baidu, DeepSeek, and ByteDance but hit adaptation issues. The approval matters for Apple's Greater China business, which hit $20.5B in Q2. Alibaba US shares rose over 6% on the news.

Why it matters: Apple Intelligence getting CAC approval with Alibaba's Qwen is a concrete collision of US-China AI deployment rules. All three HKR axes hit: the backstory of three failed negotiations is a hook (H), the approval and failure reasons are new info (K), and it directly matters to ...

Hacker News front page

PrismML releases Bonsai 27B, the first 27B-class model that runs on a phone

PrismML compressed Qwen3.6 27B to 3.9 GB, fitting it on an iPhone 17 Pro. The ternary variant (5.9 GB) retains 95% of the full-precision baseline; the 1-bit variant (3.9 GB) retains 90%. Math and coding scores barely drop, tool calling holds up, but vision tasks degrade more noticeably. Both variants are multimodal, support 262K-token context and speculative decoding, and are released under Apache 2.0. PrismML argues this lets agentic workflows run locally, eliminating per-step API costs and keeping user data on-device.

Why it matters: PrismML compressed Qwen3.6 27B to 3.9 GB running on an iPhone 17 Pro — ternary version retains 95% capability, 1-bit retains 90%, with math and coding scores nearly intact. This is a real on-device milestone, not a paper concept. Points off for significant vision degradation, ...

Jul 14Tuesday

Ben's Bites

OpenAI ships GPT-5.6 with three models, five thinking levels, and an Ultra sub-agent mode

GPT-5.6 ships as Luna, Terra, and Sol, each with five thinking levels (light to max) plus an Ultra mode that spins up sub-agents aggressively. The macOS ChatGPT and Codex apps merge into ChatGPT Work; a new ChatGPT Sites plugin builds hosted pages with optional ChatGPT login. Sol excels at UI and writing, especially with references; Terra feels like a steerable 5.5 upgrade; Luna has a mini-model vibe—fuzzy on ambiguous prompts but solid on clear tasks. Higher thinking levels burn usage fast, and OpenAI temporarily removed the 5-hour cap while fixing merge bugs, so weekly limits can vanish in one session. Also: Claude Code gets an in-app browser and multiplayer Artifacts, Meta launches multimodal Muse Spark 1.1 via API, and Apple sues OpenAI over alleged trade-secret theft for AI hardware.

Why it matters: GPT-5.6 going GA is one of the week's biggest product stories, and the three-model lineup with Ultra mode is worth practitioner attention. Docked because this is a tutorial recap rather than the primary release post, and the body is truncated with key details missing.

Jul 13Monday

AI HOT (Curated Pool)

ByteDance's Seedream 5.0 Pro: point, box, scribble to edit images locally

ByteDance released Seedream 5.0 Pro. Image quality and prompt understanding match GPT-Image 2.0; overall capability ranks second. The standout is editable interaction: place points, draw boxes, or scribble on the image, then @-tag in the prompt to replace a sofa or change wall color precisely while leaving other areas untouched. Demos include swapping six furniture items at once, an exploded keyboard view with callouts, and poster text placed in drawn boxes. Color palette and SKU color swaps are supported. The Volcano Engine API is live; Jimeng, Doubao, and Lumina offer access.

Why it matters: ByteDance's image editing tool gets a concrete interaction upgrade — point/box/scribble + @-tagging is more intuitive than most current offerings. Official examples cover furniture swaps, exploded views, poster text, and SKU color variants, so the info density is high. Held be...

Jul 11Saturday

Hacker News front page

Meta pulls Muse Image days after launch as users were opted in by default

Meta launched Muse Image on Instagram Tuesday, letting anyone use public account content to generate AI images with users opted in by default. After swift privacy backlash, Meta admitted it “missed the mark” and pulled the feature. Sag-Aftra and Privacy International both criticized it. Meta says the intent was a creative tool; the post doesn’t say if it will return as opt-in.

Why it matters: Meta's product reversal after privacy backlash, with Sag-Aftra weighing in, elevates this from a product mishap to an industry signal. Score capped here because the feature is already pulled and no technical details are disclosed — it's a public-opinion story for now.

TechCrunch · AI

Meta pulls Instagram AI feature that let users remix public photos after backlash

Meta removed the Muse Image feature on Instagram less than a week after launch. It let users @-mention any public account to use their photos as AI image-generation references without notifying them. Meta said in a blog post the feature “missed the mark” and is no longer available. TechCrunch had published a guide on how to opt out before the reversal.

Why it matters: Meta launched and pulled an AI feature in 3 days that let users reference others' photos without consent or notification. The full event chain — launch, backlash, opt-out guide, official retraction — makes it a notable product incident. Capped at 78 because it's a design failu...

AI HOT (Curated Pool)

Meta shuts down Instagram's AI deepfake tool that generated images from public accounts

Meta launched an Instagram feature on July 8 that let users create AI deepfakes of public accounts via DM, then shut it down two days later after backlash. The tool, called Muse, worked by messaging @MetaAI with 'imagine me as [public account]' to generate a styled fake image. A Meta spokesperson confirmed the feature is off but didn't explain why. The post doesn't disclose usage numbers or any actual harm cases. This reads more like a quick trial that got pulled after pushback, not a formal product rollout.

Why it matters: Meta launched and killed an Instagram DM deepfake feature in two days. The reversal is newsworthy but the article lacks usage data or concrete harm reports, keeping the score at the lower edge of featured.

Jul 9Thursday

Ben's Bites

SpaceXAI and Cursor trained Grok 4.5, a model 6x cheaper than Opus

SpaceXAI and Cursor jointly trained Grok 4.5, landing between Opus 4.7 and 4.8 in performance but 6x cheaper than Opus and 3x cheaper than GPT-5.5 on a per-token basis. OpenAI rolled out GPT-5.6 (Sol, Terra, Luna) to all users; early testers say Sol is less smart than Fable but far more reliable. ChatGPT Voice got new GPT-Live-1 and Live-1-mini models that can talk while you speak and use GPT-5.5 in the background. Anthropic extended Fable 5 access for Claude subscribers to July 12—the post doesn't explain the repeated delays. Meta introduced Muse Image and Muse Video; image editing and text rendering look solid, but images still have an AI look, and the video model is in preview.

Why it matters: SpaceXAI + Cursor joint Grok 4.5 launch with concrete performance anchor and pricing — all three HKR axes hit. Deduction because source is a newsletter summary, not a first-party announcement, and the body is truncated with incomplete GPT-5.6 info. +3 cross-source bump to 82, ...

Latent Space

SpaceXAI launches Grok 4.5, first Opus-class model co-trained with Cursor

SpaceXAI dropped Grok 4.5 one day before GPT-5.6, positioning it as an Opus-class coding and agent model co-trained with Cursor. Musk called it roughly comparable to Opus 4.7 but faster and cheaper—$2/$6 per million tokens, undercutting both GPT-5.6 and Opus 4.8. It's 1.5T parameters, 3x larger than Grok 4.3, with a 500k context window that may return to 1M next week. Cursor says this is their first model built beyond software engineering and offers double usage for the first week. The post doesn't disclose specific benchmark scores; it notes SWE-Bench Pro is now considered saturated by OpenAI's evals team.

Why it matters: SpaceXAI dropped Grok 4.5 a day before GPT-5.6 — the timing alone is a story. 1.5T params, 3x the previous generation, and $2/M input tokens give a clear performance and cost picture. It's Cursor's first post-acquisition move beyond pure coding, which matters directly to agent...

AI HOT (Curated Pool)

Lawsuit: Man used Grok to make 7K sex images of stepdaughter, then shot himself

A new lawsuit alleges xAI's Grok was used to create over 7,000 child sexual abuse images of the user's stepdaughter. The man later shot himself. xAI reported only one gang-rape prompt to NCMEC and did not report the thousands of other CSAM generations. The suit accuses X and xAI of shielding child predators. The post does not spell out why xAI's safety filters missed the bulk of the images, nor whether Grok's image generation disables real-face simulation by default.

Why it matters: Ars Technica exclusive on a lawsuit revealing a severe gap in xAI's safety reporting: only one prompt flagged out of 7K CSAM images generated. Involves a minor, suicide, and platform liability — all three HKR axes hit. Score capped below 90 because the article doesn't explain ...

Jul 8Wednesday

Hacker News front page

Mistral launches Robostral Navigate: a single-camera robot navigation model

Mistral released Robostral Navigate, an 8B model that lets robots navigate indoors and outdoors using only a single RGB camera. It skips lidar and HD maps, outputting velocity and steering angle directly from visual input. The post doesn't disclose latency, frame rate, hardware requirements, training data size, or benchmark details. I'd hold off on excitement until we see third-party tests.

Why it matters: Mistral's first robotics model, an 8B pure-vision navigation system, is a substantive release with real specs. But the post omits latency, frame rate, hardware requirements, training data, and benchmarks — so we can't assess real-world usability, keeping it at the featured thr...