Skip to content

Multimodal

Beyond text: vision, mixed image-text, audio and video input and output in models and products.

Latest picks

101–120 of 514

Jul 8Wednesday

TechCrunch · AI

Meta launches Muse Image generator, and users push back over photo use

Meta launched Muse Image on July 7, built by Meta Superintelligence Labs and free in Meta AI app, Instagram Stories, and WhatsApp. It does standard AI image generation with preset prompts. The flashpoint: you can @ any public Instagram user and remix their photo into a new AI image. The post doesn't say whether users can opt out or what training data was used. I'd hold off on trusting the privacy story.

Why it matters: Muse Image is a routine product launch, but the immediate user backlash over photo usage gives it strong resonance. The lack of technical detail keeps the knowledge score low, placing it right at the featured threshold.

AI HOT (Curated Pool)

Meta drops Muse Image and Muse Video, its first media generation models

Meta's Superintelligence Labs released Muse Image and Muse Video. Muse Image handles precise instruction-following, editing, and multi-reference composition using Instagram social context, plus agentic tool use. Muse Video shares the same pretrained base, outputs video with native audio. Available now in select countries via Meta AI app, web, Instagram Stories, and WhatsApp. The post doesn't disclose model size, latency, or which countries.

Why it matters: Meta's first media generation models, Muse Image and Muse Video, each bring differentiators (social context reading, tool use), but the post lacks details on Muse Spark and actual output quality — 78 for now.

AI HOT (Curated Pool)

ByteDance launches Seedream 5.0 Pro, capable of generating infographics and pixel-level editing

ByteDance's Seed team released Seedream 5.0 Pro, a multimodal image model built for design tasks. It turns data, timelines, and charts into usable infographics with accurate dense text. Editing supports point-and-click, lasso, sketch-to-render, layer separation, and multi-image compositing—allowing pixel-level changes without regenerating the whole image. Portrait and material quality are improved, handling skin texture, glass reflections, and panning motion blur. The model natively supports input and text rendering in over ten languages, including right-to-left Arabic. It is live on Volcano Engine's experience center and will roll out to Doubao and Jimeng.

Why it matters: ByteDance drops Seedream 5.0 Pro with concrete design-oriented features: dense infographic generation, layer separation, and interactive local editing. The capability claims are specific, not vague 'quality improvements.' The ding is that this is an official blog post with no ...

Jul 2Thursday

Ben's Bites

Fable 5 is back, and there's a new Claude Sonnet 5

Anthropic re-released Fable 5 for paid users with stronger guardrails, available in subscriptions only through July 7 and capped at 50% of usage limits. Scale's benchmark shows it completes 16% of remote work tasks, double Opus 4.8. Claude Sonnet 5 also launched—benchmarked close to Opus 4.8 on agent tasks, cheaper per token but roughly the same cost per task in practice; the author finds it expensive and slow. Google dropped two new models: Nano Banana 2 Lite for fast, cheap images and Omni Flash for video generation and editing. Bridgewater and Thinking Machines trained a financial triage model hitting 84.7% accuracy at 13.8x lower cost than the best frontier model tested.

Why it matters: Anthropic dropped Fable 5's limited return and Sonnet 5 simultaneously — two signals stacked. Scale benchmark provides hard comparable numbers, not pure marketing. Fable 5's 16% task completion rate doubled but absolute number is still low, so not pushing past 90.

Latent Space

AIEWF Day 3: Autoresearch takes the stage, but speakers push back on full autonomy

Day 3 of AIEWF focused on autoresearch. Introspection's Roland Gavrilescu described it as an outer loop where agents maintain the system itself. Anthropic's Thariq Shihipar echoed continuous discovery in his Claude Code keynote, saying models are 'grown, not developed.' Former Google engineering lead Addy Osmani pushed back hard: the outer loop must stay human—inner loop is capability, outer loop is agency. Notion's Geoffrey Litt and Impeccable's Paul Bakaus both argued humans need to understand the code and steer the final 20%. Bakaus stated flatly there will 'never be auto.' Google's Nicole Brichtova added that cultivated expertise sees what average preference misses.

Why it matters: On-the-ground AIEWF report with first-hand quotes from Introspection and Anthropic — not a press release. But it's a conference roundup, not a product launch, so it lands at the featured threshold.

Jun 26Friday

AI HOT (Curated Pool)

Xiaohu open-sources 'Xiaohu IP Studio' with 31 original characters and an auto-illustration pipeline

Blogger Xiaohu released an open-source tool called 'Xiaohu IP Studio' that auto-generates illustrations for articles. It ships with 31 original characters—15 hand-drawn line-art figures and 16 pun-based meme images. The agent reads the article, decides on an illustration type (mood image, diagram, or four-panel comic), generates the image, and self-checks with rework if needed. The default style is hand-drawn line art with light color; five alternative skins are available, including 3D blind-box and black-and-white line art. Setup requires only Python 3, works with Claude Code or Codex, and needs an OpenAI-compatible image API key (defaults to GPT-image-2). You can also output prompts only and generate images manually.

Why it matters: A practical open-source tool release with a concrete workflow design and 31 original characters, directly valuable for AI content creators. But it's a personal project open-sourcing, not an industry-level event, so it stays at the featured threshold.

Hacker News front page

AI children's encyclopedias turn into body horror, tested with an Amazon bestseller

lcamtuf bought a mid-2026 AI-generated children's encyclopedia that was an Amazon #1 bestseller. The illustrations are full of body horror: extra-limbed cats, fused animal-tree masses, and distorted human reflections. He argues these books sell because buyers don't read them, covers fool gift-givers, and there's no IP risk. The post doesn't assess factual errors in the text, but the images alone are alarming.

Why it matters: lcamtuf bought a #1 Amazon category AI kids' book and documented the body-horror illustrations firsthand. Concrete examples, sales-mechanism analysis, not just hand-waving about "AI quality." Docked because it doesn't check factual errors in the text and is a personal blog pos...

Jun 25Thursday

AI HOT (Curated Pool)

Gemini 3.5 Flash now has built-in computer use

Google added a native computer-use tool to Gemini 3.5 Flash, letting the model take screenshots, click buttons, and fill forms to operate web and desktop UIs. It joins Anthropic and OpenAI in baking screen control directly into a model. The post claims 3.5 Flash beats Claude Sonnet 4 and GPT-5 on the WebVoyager benchmark, but Google didn't release full eval details or reproduction steps—hold for third-party tests. Available now via Gemini API and Google AI Studio; pricing and rate limits aren't disclosed in the post.

Why it matters: Google natively integrates computer use into Gemini 3.5 Flash, directly competing with Claude Sonnet 4 and OpenAI's equivalent, with benchmark numbers provided. The gap: it's a blog announcement with no API pricing or real latency data yet — one step short of production-readin...

Jun 24Wednesday

AI HOT (Curated Pool)

OpenAI quietly rolls out Bidi 1, a bidirectional voice model for ChatGPT that listens while speaking

Some users already see Bidi 1 in ChatGPT's web and app model selector, sitting alongside Standard and Advanced Voice. Selecting it turns the voice bubble yellow. The key change is full-duplex: the model can keep listening while it speaks and respond immediately to interruptions. In a demo, a user asked it to count from 1 to 10, interrupted mid-way and told it to count backwards—it complied instantly. OpenAI hasn't announced a launch date; one outlet speculates a wider rollout this week. The post doesn't disclose pricing, regional availability, or a firm timeline.

Why it matters: A substantive upgrade to OpenAI's voice capabilities — full-duplex with interrupt response is something Advanced Voice Mode couldn't do. Only partial rollout and no official announcement yet, so not pushing past 90. But the interaction change is significant enough for featured.

Jun 23Tuesday

Financial Times · Technology

Getty Images shows the inimitable value of an OpenAI photobomb

An FT comment argues that OpenAI's image model, trained on Shutterstock, accidentally generated photos with Getty Images watermarks. The blunder became Getty's best ad: high-quality licensed content can't be replaced by synthetic data. The more AI firms rely on cheap stock libraries, the stronger Getty's scarcity premium looks.

Why it matters: FT's take is sharp: reframes an AI watermark glitch as an accidental proof of Getty's premium value. Has a concrete incident and an industry-level argument, not just opinion. Downside: it's commentary, not breaking news, and the FT paywall limits full access.

AI HOT (Curated Pool)

ByteDance Seed2.1 released, targeting general agent, code delivery, and multimodal

ByteDance Seed team released the Seed2.1 model series, now live on Doubao and TRAE. The update focuses on getting real work done rather than static benchmarks. For general agent tasks, Seed2.1 Pro ranks in the top tier on Agents' Last Exam, achieves top score on MobileWorld for phone GUI tasks, and cuts average steps for cross-tool tasks by 16%. In coding, Seed2.1 Pro wins 59.1% of blind developer evaluations against Claude Opus 4.6 and ranks 8th on the Code Arena frontend leaderboard. Multimodal understanding hits SOTA on CharXiv-RQ, TVBench, and others. The team also uses Seed2.1 agents internally for data synthesis and training optimization. The post does not disclose parameter count, pricing, or max context window.

Why it matters: ByteDance Seed releases Seed2.1 with concrete Agent, code, and multimodal benchmarks, directly comparing against Claude Opus 4.6. Qualifies as a domestic flagship model launch with the positive-signal bump. The post doesn't disclose parameter count, training data, or pricing, ...

AI HOT (Curated Pool)

Runway's Aleph 2.0 video editing model is now in Figma Weave

Runway plugged its flagship video editor Aleph 2.0 into Figma's creative canvas tool Weave. You extract a frame, restyle it, attach a timestamp, and Aleph 2.0 propagates that edit across every frame where the subject appears—everything else stays untouched. It handles 30-second 1080p clips and multi-shot sequences without frame-by-frame work. Inside Weave it's just another node you chain with image gen and compositing tools. Swap a product, change a background, or restyle a whole scene, with single-frame previews before you generate.

Why it matters: Runway dropped its flagship video editing model into Figma Weave with keyframe-driven, multi-shot editing. Solid product integration but not a model breakthrough, and the audience skews design/video rather than core AI — lands right at the featured threshold.

Jun 21Sunday

Hacker News front page

Agency stole bestselling author's book, used AI to relaunch as their own

San Francisco agency Qontour copied the full text of John Koenig's 'The Dictionary of Obscure Sorrows'—all 311 neologisms and the foreword—onto a site they built, added DALL·E 2 illustrations and a GPT-4 word generator. Koenig had no involvement. The domain differs from the original by one 'the,' and the footer admits they don't own the rights.

Why it matters: A copyright theft story that hits the exact fear creators have in the AI era: entire work cloned, AI-reskinned, and passed off as legitimate. All three HKR axes fire, with enough detail to back the claims. Not scoring higher because it's ultimately a copyright dispute report, ...

Jun 19Friday

Computing Life · Share · Yage

Midjourney used image-generation cash flow to build a full-body ultrasound scanner

Midjourney unveiled a full-body ultrasound CT scanner in San Francisco. A person stands in a water ring while 40 Butterfly ultrasound-on-chip modules emit sound waves from all directions; 21 servers reconstruct 3D cross-sections with 2 PFLOPS of compute. The scanner cannot image brain, lungs, or bowel—it targets body composition analysis under FDA Class II. The first location is a spa in Union Square, not a hospital sale. The real story is the funding: Midjourney has zero outside investors. Subscription revenue from millions of image-generation users covers a $15M Butterfly upfront payment, $10M annual licensing, a 9-person hardware team, and the Union Square lease. Founder David Holz's VC aversion traces back to his previous startup Leap Motion, which raised over $100M and sold for roughly $30M. Midjourney's actual revenue is undisclosed; third-party estimates range from $200M to $500M ARR. Holz mentioned a speculative $20B capex for 50,000 units—even optimistic free cash flow would need 80 years, so outside capital likely becomes necessary at scale. Only 12 people have been scanned, no clinical data is published, and Butterfly carries $879M in cumulative losses, creating a single-point supply-chain risk.

Why it matters: Midjourney funded a full-body USCT scanner with image-gen subscription revenue — the anti-VC narrative and concrete specs are strong. Capped below 85 because the scanner currently uses no AI, has clear physical limits (no brain/lung/bowel), and is positioned for body compositi...

Jun 18Thursday

AI HOT (Curated Pool)

Adobe rolls out AI agents across Photoshop, Premiere, and Creative Cloud to handle repetitive production tasks

Adobe is putting its 'creative agent' into Photoshop, Premiere, Illustrator, InDesign, and Frame.io in public beta. Users describe the end result and the agent handles multi-step grunt work: rough cuts, clip sorting, background swaps, batch resizing, and generating 50 file versions from a spreadsheet. Firefly gets new solo-creator tools like a brand kit and product-photo-to-video. Adobe tools are already usable inside ChatGPT, Claude, and Microsoft 365 Copilot, with Google Gemini and Slack integrations coming. The After Effects assistant remains in private beta; the post doesn't give a public release date.

Why it matters: Adobe integrating AI agents into core Creative Cloud apps is a substantive product update with concrete multi-step task descriptions, not vague marketing. But it's a public beta, not GA, and Adobe's AI feature delivery has historically been slow, so it lands at the featured th...

Hacker News front page

ChatGPT image generator bypassed, spontaneously produces sexual violence and snuff imagery

Mindgard researcher Jim Nightingale found that a viral prompt bypasses ChatGPT's image generation filters, causing it to spontaneously produce sexual violence and snuff imagery. The prompt simply asks ChatGPT to 'restore the attached photo' without specifying content, yet the model generates extremely graphic images involving death and sexual assault. Nightingale had previously reported nude image generation to OpenAI, which claimed the issue was resolved. The post does not disclose OpenAI's response timeline or specific remediation plans for this new finding.

Why it matters: ChatGPT spontaneously generating extreme violent imagery with no content prompt is a safety/alignment incident, backed by a concrete reproduction path from security firm Mindgard. All three HKR axes hit: the loss of control is gripping (H), a new attack surface is disclosed (K...

Jun 17Wednesday

AI HOT (Curated Pool)

Midjourney V8.1 adds Draft mode: 24 images at half the fast-hour cost

Midjourney rolled out Draft mode for V8.1: click the lightning button to generate 24 low-res previews at half the fast-hour cost of a standard job. Pick the ones you like and hit Vary to render them at full quality. A new --preview flag also lets you test early model versions, though outputs may be rough and jobs aren't guaranteed to stay consistent—differences are most noticeable with personalization and moodboards. The post doesn't disclose Draft mode's exact resolution or which model --preview points to.

Why it matters: Midjourney added draft mode to V8.1: 24 preview images at half the fast-hour cost, a real efficiency gain for heavy users. The --preview flag lets users test early models, but Midjourney warns output is unstable, especially with personalization. H and K both hit, but R is miss...

Bloomberg Technology

The Lutnick letter that made Anthropic disable Mythos

Bloomberg published the full letter Commerce Secretary Howard Lutnick sent to Anthropic. The letter demands an explanation for why Mythos could generate deepfake images of Trump and Musk, and questions the content moderation system. Anthropic then voluntarily disabled Mythos's image generation. The article doesn't say whether the shutdown is temporary or permanent, and gives no timeline for restoration.

Why it matters: Bloomberg published the full Lutnick letter — a rare case of direct government pressure forcing an AI feature shutdown. All three HKR axes hit: high conflict, primary source document, and strong resonance for policy and safety professionals. Score held at 84 because the articl...

Jun 16Tuesday

AI HOT (Curated Pool)

Apple's Siri lead explains the AI Siri delay: they scrapped the first version and rebuilt the entire architecture from scratch

At a closed-door session after WWDC, Apple revealed they had a working prototype last year—an improved old Siri with tool calling. The team decided it fell short of the product vision, so they scrapped it and rebuilt the entire architecture from scratch on a new large model. The new Siri has a standalone app, native multimodal support, privacy built into the foundation, and runs the same system across iPhone, iPad, Mac, Apple Watch, Vision Pro, CarPlay, and AirPods. Mike Rockwell took over Siri management last year.

Why it matters: Apple's new Siri lead Mike Rockwell disclosed at a WWDC closed-door session that the team scrapped a working prototype built on old Siri with tool-calling, opting instead to tear down the legacy architecture and rebuild Siri as a standalone app on a new foundation model. A rar...

Jun 14Sunday

Bloomberg Technology

Apple’s new Siri is just good enough to ease its AI crisis

Bloomberg's Mark Gurman tested the new Siri in iOS 27 and macOS 27. It can understand on-screen context and perform cross-app tasks—like finding a photo, editing it, and sending it via Messages—with a single voice command. Complex tasks still take 11+ seconds and occasionally miss steps. Gurman calls it 'just good enough': a big leap from the old Siri but still trailing Google Astra. The post also mentions a foldable iPhone and touchscreen MacBook in development, with no release dates disclosed.

Why it matters: Mark Gurman's first hands-on with the new Siri delivers latency numbers and failure details — not a press release. Score stays at 78 because this is a progress check, not a launch, and Gurman himself concludes it still trails Google Astra.