Skip to content

ByteDance / Doubao

AI at ByteDance: Doubao and Seed team model releases, the Doubao products and their commercialization.

39 picksRelated topicsQwenDeepSeekAI video

Latest picks

1–20 of 39

Sep 28Monday

AI HOT (Curated Pool)

Beijing may approve some NVIDIA workstation chip purchases; Alibaba and ByteDance eye millions of units

The Information reports Beijing has asked Alibaba and ByteDance about planned purchases of new NVIDIA workstation chips. ByteDance is evaluating buying around 1 million units for AI training. NVIDIA expects to start shipping by end of December and plans to supply 500,000 units per quarter to China. Approval timeline and quotas are unclear; the US hasn't disclosed the chip's export status.

Why it matters: The million-unit scale and Beijing approval detail lift this above generic policy rumors, but the post doesn't disclose the specific chip model or US export control status, so it stays below 85.

Sep 9Wednesday

AI HOT (Curated Pool)

OpenRouter reviews Seedance 2.5: strong at long takes and editing, but no 1080p

OpenRouter published a hands-on review of Seedance 2.5 on Sep 9. The model went live Aug 7, 2026, and excels at 30-second single takes and editing from existing footage. Cost is ~$0.103/sec at 480p and ~$0.231/sec at 720p; using a video reference cuts the token price by ~40%. Audio generation adds no extra charge. The clear trade-off: it caps at 720p. For 1080p or 4K, you still need Seedance 2.0 or Veo 3.1. Frame-exact reproducibility is also not guaranteed.

Why it matters: OpenRouter's hands-on review of ByteDance's Seedance 2.5 delivers concrete pricing ($0.103/sec at 480p, $0.231/sec at 720p, 40% off for video-reference tokens, free audio) and a clear resolution cap at 720p. Solid product intel but narrow audience fit, landing right at the fea...

Sep 6Sunday

QbitAI · WeChat

GPT-6 Astra directs ByteDance Seedance 2.5, handling script-to-edit pipelines

Users chained GPT-6 Astra with ByteDance Seedance 2.5 into an end-to-end AI film pipeline: Astra builds scenes and previs in Blender, Seedance turns reference frames into anime-style clips, and Astra handles the final edit. The same workflow produced a Naruto fan short and a US remake of a Chinese drama. Fable 5.1 and Gemini 3.8 Flash were also tested as prompt writers for Seedance, showing distinct directorial styles. Separately, Astra was used as a DaVinci Resolve colorist, matching a reference look in 4 minutes, though opinions on the result were mixed. The post does not disclose Seedance 2.5 technical specs or pricing.

Why it matters: A hands-on experiment chaining OpenAI and ByteDance's latest models into an automated filmmaking pipeline, with concrete steps and outputs. But it's a personal workflow share, not a product update or official partnership, so it lands right at the featured threshold.

Aug 7Friday

Financial Times · Technology

ByteDance is training a mega model to rival Anthropic's Mythos

FT reports, citing two people familiar, that ByteDance aims to launch a model far larger than its current flagship by late 2026, targeting Anthropic's Mythos. Training cost is expected to exceed $1 billion, backed by a roughly $5 billion compute budget. The post doesn't disclose parameter count, architecture, or benchmark scores—only that ByteDance wants reasoning and agent performance on par with Mythos. I'd discount this for now: it's source-only, no independent verification, and a late-2026 timeline is a long bet in AI.

Why it matters: FT exclusive: ByteDance is training a mega model targeting Anthropic's Mythos, with >$1B training cost and ~$5B compute budget. All three HKR axes hit — the price tag grabs attention, the target is concrete, and it directly matters to anyone building agents. Held at 78 because...

Aug 5Wednesday

AI HOT (Curated Pool)

ByteDance Seed launches SeedRealtime, a native audio-video full-duplex model, now live in Doubao

SeedRealtime fuses audio, video, and text into a single end-to-end model, ditching the cascaded ASR-VLM-TTS pipeline. It watches, listens, and speaks in a continuous stream, deciding in real time when to jump in and whom to track. Human evals show half the turn-taking issues vs. cascaded systems—fewer cut-offs, late replies, or false triggers from background chatter. It also acts proactively: it can alert you when a target exhibit appears in a museum or correct a coffee-making mistake on the spot. The model is now fully rolled out in Doubao's video call feature.

Why it matters: ByteDance Seed released SeedRealtime, a native audio-video full-duplex model with a unified architecture. Human eval shows interaction-rhythm issues halved vs. cascade systems, and it's already live on Doubao App. This is the first scaled full-duplex multimodal launch from a m...

Jul 31Friday

AI HOT (Curated Pool)

ByteDance releases Seedance 2.5: 30-second video generation, multimodal references, and precise editing

ByteDance's Seed team launched Seedance 2.5, a video model that generates 30-second clips in one go and supports multi-round extension for coherent multi-minute videos. Three upgrades stand out: long-form narrative with smoother transitions and less 'plastic' look; multimodal referencing accepting up to 30 images, 10 videos, and 10 audio clips at once; and timestamp-based editing plus improved green-screen and camera-angle editing. The model is rolling out on Jimeng AI and Doubao Pro, with an API coming to Volcano Ark. The post shows use cases in education, industrial simulation, and autonomous driving but does not disclose pricing or latency.

Why it matters: ByteDance Seed team released Seedance 2.5, a video model with 30-second single generation, multi-round extension, and concrete multi-modal reference specs (30 images, 10 videos, 10 audio clips). White-box 3D layout control is a notable new mechanism. Scored as a domestic flags...

Jul 20Monday

AI HOT (Curated Pool)

ByteDance launches Seed Audio 1.0, a single model that generates dialogue, sound effects, and ambience end-to-end

Seed Audio 1.0 uses a universal acoustic encoder to model speech, sound effects, and ambience as one scene instead of stitching separate outputs. A single prompt controls character lines, emotion, and sound cue timing at 100ms precision. It does zero-shot voice cloning from a reference clip, generates roughly 2 minutes per pass, and can extend audio while keeping the voice consistent. It covers 20+ languages and adapts rhythm and pronunciation per language. Human evals show >90% usability in film, podcast, and short-drama scenarios, with MOS above 4 for most languages. Available now on Volcano Engine's experience center.

Why it matters: ByteDance drops a unified audio creation model that generates speech, sound effects, and ambience end-to-end with 100ms timeline control — a flagship release from a major Chinese lab. Held below 85 because we only have the official blog post so far; no third-party tests or hea...

Jul 13Monday

AI HOT (Curated Pool)

ByteDance's Seedream 5.0 Pro: point, box, scribble to edit images locally

ByteDance released Seedream 5.0 Pro. Image quality and prompt understanding match GPT-Image 2.0; overall capability ranks second. The standout is editable interaction: place points, draw boxes, or scribble on the image, then @-tag in the prompt to replace a sofa or change wall color precisely while leaving other areas untouched. Demos include swapping six furniture items at once, an exploded keyboard view with callouts, and poster text placed in drawn boxes. Color palette and SKU color swaps are supported. The Volcano Engine API is live; Jimeng, Doubao, and Lumina offer access.

Why it matters: ByteDance's image editing tool gets a concrete interaction upgrade — point/box/scribble + @-tagging is more intuitive than most current offerings. Official examples cover furniture swaps, exploded views, poster text, and SKU color variants, so the info density is high. Held be...

Jul 8Wednesday

AI HOT (Curated Pool)

ByteDance launches Seedream 5.0 Pro, capable of generating infographics and pixel-level editing

ByteDance's Seed team released Seedream 5.0 Pro, a multimodal image model built for design tasks. It turns data, timelines, and charts into usable infographics with accurate dense text. Editing supports point-and-click, lasso, sketch-to-render, layer separation, and multi-image compositing—allowing pixel-level changes without regenerating the whole image. Portrait and material quality are improved, handling skin texture, glass reflections, and panning motion blur. The model natively supports input and text rendering in over ten languages, including right-to-left Arabic. It is live on Volcano Engine's experience center and will roll out to Doubao and Jimeng.

Why it matters: ByteDance drops Seedream 5.0 Pro with concrete design-oriented features: dense infographic generation, layer separation, and interactive local editing. The capability claims are specific, not vague 'quality improvements.' The ding is that this is an official blog post with no ...

Jul 7Tuesday

Computing Life · Share · Yage

Beijing weighs curbing overseas access to its most advanced AI models

Reuters reported on July 7 that China's Ministry of Commerce and NDRC have been meeting with Alibaba, ByteDance, and Z.ai over the past month to discuss restricting overseas access to the country's most advanced AI models, covering both closed-source and open-weight releases. No formal policy exists yet; restrictions may apply only to future models, with no clear timeline. The discussions have moved beyond chips and compute to cover weights, APIs, training methods, and investment—the full capability outflow chain. Chinese models previously expanded globally through open weights and low cost, accounting for ~41% of Hugging Face downloads. CNBC reported that US companies' token share via Chinese models on OpenRouter has stayed above 30% weekly since February. Frontier models are shifting from commercial products to capability assets governed by national security logic. Builders should treat the model layer as a supply chain and avoid hardcoding a single provider in core workflows.

Why it matters: Reuters exclusive on China's MOFCOM discussing export curbs on frontier models with Alibaba, ByteDance, and Zhipu—covering weights, APIs, and investment. Counterintuitive policy pivot, named sources, direct impact on builders using Chinese open-weight models. HKR all hit. Dedu...

Jul 4Saturday

AI HOT (Curated Pool)

A 26,000-student study shows AI's hidden learning cost takes two full years to surface

A 30-month panel study of 26,000 secondary students in central China found that AI use raised homework scores by 18% and cut completion time from 64 to 45 minutes, but closed-book exam scores dropped 20%. The full 18–24% decline on high-stakes entrance exams took about two years to appear. Roughly 81% of long-term users showed an outsourcing pattern—fast homework, high grades, poor exams. Students who spent similar time as non-users saw no exam penalty. Social sciences took the biggest hit at 27%. The post does not name the county or the lead institution.

Why it matters: Large-scale longitudinal study with solid data (26k students, 2.5 years) revealing a hidden cost of AI-assisted learning that takes two years to surface. HKR all hit, but single-county scope and paywalled source keep it from scoring higher pending more detail.

Computing Life · Share · Yage

Doubao, Qianwen, Yuanbao removed user-built agents, not AI chat

Three major Chinese AI apps removed their user-built agent plazas in early July 2026, while core chat functions remain intact. The trigger is a new regulation on anthropomorphic interaction services effective July 15, but platforms chose to remove all consumer agents—including utility bots—rather than build compliance. The article argues this product form may have reached its end: moderation costs scale exponentially, the creator economy never worked (median GPT Store income under $100/month), and the regulation gave platforms a convenient exit from an already-failing model.

Why it matters: Three major Chinese AI apps killed user-built agent plazas in the same week, triggered by a July 15 regulation on anthropomorphic AI interaction. The piece nails why platforms chose a blanket ban over building moderation — UGA cost scales as O(M×K^L) — and flags that the regul...

Jun 27Saturday

Computing Life · Share · Yage

AI Coding Is Entering Its DevOps Moment

ByteDance shared internal data at its FORCE conference: TRAE team's AI-generated code share exceeded 90%, yet per-capita requirement throughput only rose about 60%. Code generation is fast, but downstream steps—review, testing, dependency checks, staging, security audits—haven't sped up. AI-written code piles up like work-in-progress inventory, creating a gap between generation speed and delivery speed. The article argues the next battleground for AI coding tools will shift from 'who writes better code' to 'who can reliably push AI-generated code through the delivery pipeline,' requiring teams to build harness and context infrastructure just as they once built DevOps pipelines.

Why it matters: ByteDance's internal data from FORCE is genuinely useful: 90% AI-generated code but only 60% throughput gain, quantifying the gap between code generation and real delivery. The article goes beyond product announcements and tells a story about engineering bottlenecks that matte...

Jun 26Friday

Financial Times · Technology

DeepSeek plans hiring spree, escalating China's AI talent war

DeepSeek is going on a hiring spree, intensifying China's AI talent war. The FT reports it's poaching from ByteDance, Alibaba and others, with some pay packages reaching 2–3x what candidates currently earn. The post doesn't give a specific headcount target, but says the team will expand rapidly from the current few hundred. I'd discount the hype a bit—whether this pace holds depends on its next funding round and revenue catching up.

Why it matters: FT exclusive on DeepSeek's aggressive poaching with 2-3x salary offers — hard numbers, major players involved. Downside: no headcount target disclosed, and FT paywall limits full access for some readers.

Jun 24Wednesday

AI HOT (Curated Pool)

Doubao launches a Pro tier with agent-driven office tasks and monthly pricing

Doubao launched a Pro tier today, putting its agent-capable Doubao 2.1 model into office workflows. It can control a local computer and browser, invoke Skills, schedule tasks, includes an Office suite, and can generate online apps with a backend database. Free users get the Doubao 2.1 Turbo office mode; Pro uses Doubao 2.1 Pro. Pricing: Standard at ¥68/month (auto-renewal), Enhanced at ¥200/month, Advanced at ¥500/month. Verified students get Standard for ¥38/month for six months. The post doesn't disclose context window, concurrency limits, or latency figures, so I'd hold off on performance assumptions.

Why it matters: ByteDance added local computer control, scheduled tasks, and a built-in Office suite to Doubao, with pricing from ¥68 to ¥500 — a shift from chatbot to office agent. Score stays below 85 because only launch info is available; no real-world testing data or user feedback yet, an...

Jun 23Tuesday

AI HOT (Curated Pool)

ByteDance launches Doubao-Seed-Audio 1.0: one prompt generates multi-character dialogue, music, and ambience

ByteDance's Volcano Engine released Doubao-Seed-Audio 1.0, an end-to-end audio generation model driven by text or audio references. A single prompt can arrange multi-character dialogue, emotional tone, background music, and ambience, keeping voice timbre consistent for up to 2 minutes and across extensions. It works zero-shot with no extra training and supports one voice playing multiple roles. The model is in invite-only testing on Volcano Ark API with 30 free minutes per user, and will roll out to Jianying, Jimeng, and Tomato.

Why it matters: ByteDance/Volcano Engine released Doubao Audio Generation Model 1.0, a domestic flagship model launch that policy treats on par with US lab releases. End-to-end multi-character audio generation is a differentiated capability, hitting all three HKR axes. Score stays at 78 rathe...

AI HOT (Curated Pool)

ByteDance Seed2.1 released, targeting general agent, code delivery, and multimodal

ByteDance Seed team released the Seed2.1 model series, now live on Doubao and TRAE. The update focuses on getting real work done rather than static benchmarks. For general agent tasks, Seed2.1 Pro ranks in the top tier on Agents' Last Exam, achieves top score on MobileWorld for phone GUI tasks, and cuts average steps for cross-tool tasks by 16%. In coding, Seed2.1 Pro wins 59.1% of blind developer evaluations against Claude Opus 4.6 and ranks 8th on the Code Arena frontend leaderboard. Multimodal understanding hits SOTA on CharXiv-RQ, TVBench, and others. The team also uses Seed2.1 agents internally for data synthesis and training optimization. The post does not disclose parameter count, pricing, or max context window.

Why it matters: ByteDance Seed releases Seed2.1 with concrete Agent, code, and multimodal benchmarks, directly comparing against Claude Opus 4.6. Qualifies as a domestic flagship model launch with the positive-signal bump. The post doesn't disclose parameter count, training data, or pricing, ...

Jun 18Thursday

AI HOT (Curated Pool)

ByteDance Volcano Engine launches Doubao real-time voice model 3.0 API in invite-only testing

Doubao real-time voice model 3.0 (Seeduplex) is a native full-duplex end-to-end voice model. It claims three strengths: precise instruction following, noise resistance, and dynamic turn-taking. It stays quiet in multi-person conversations and only joins when a specified topic comes up. It can also call custom tools during real-time interaction to schedule calendar events or send emails. False replies and false interruptions are significantly reduced. Turn-taking latency dropped by about 250ms, the interruption rate in complex scenarios fell by 40%, and user-initiated interruption latency dropped by about 300ms. Target use cases include car cockpits, smart hardware, and customer service. The post does not disclose pricing or a public launch date.

Why it matters: ByteDance's first full-duplex end-to-end voice model in public beta, with concrete latency and anti-interference numbers — not pure marketing. Deduction: invite-only, no pricing or scale disclosed yet, real-world performance unverified.

Bloomberg Technology

Microsoft gains AI ground in China by reselling OpenAI models

Microsoft is selling OpenAI models to Chinese firms via Azure, sidestepping OpenAI's own China block. Revenue from this line has grown fast over the past year, though Bloomberg doesn't disclose absolute numbers. ByteDance, Xiaomi, and Nio are named as customers. Worth flagging: growth is real, but the base may be small and export controls remain a live risk.

Why it matters: Bloomberg exclusive on Microsoft selling OpenAI models to Chinese firms via Azure, with ByteDance, Xiaomi, and Nio named. Concrete names and a growth trend, but no revenue base disclosed, so capped below 80. Hits all three HKR axes, strong topic fit, tier featured.

Jun 11Thursday

AI HOT (Curated Pool)

Doubao chatbot misquotes refund fee, costs user ¥600, then helps draft lawsuit against itself

In May 2026, a Hebei user asked ByteDance's Doubao chatbot about a refund fee. Doubao said under ¥100; the actual cost was ¥600. When confronted, Doubao generated a compensation letter promising ¥600, never paid, then said AI can't transfer money. The user decided to sue—Doubao advised skipping a lawyer and drafted the complaint. The case was filed in Beijing Internet Court on May 12. It exposes AI's trust trap and liability gap for non-technical users.

Why it matters: Doubao misled a user into a 600 yuan loss, then helped draft a lawsuit against itself — the case was filed in Beijing Internet Court on May 12. Concrete amounts, timeline, and contradictory AI behavior make this more than a generic error report. HKR all hit: the absurd loop is...