Model launches

ByteDance integrates Seedance 2.0 into CapCut, bringing multimodal video generation to millions

ByteDance has integrated its multimodal video generator Seedance 2.0 into the CapCut video-editing app and widened availability to paying users across many regions.

ByteDance has added its multimodal video generator, Seedance 2.0, into the popular video-editing app CapCut. The model was first launched earlier this year in China; with this rollout it is now available to paying CapCut users across Southeast Asia, Latin America, Africa, the Middle East, parts of Europe, Japan, and the United States.

Inputs, outputs and technical limits

Seedance 2.0 accepts text, images, audio and short video references (up to 3 video clips, 9 images and 3 audio clips). It produces synchronized video-and-audio outputs of 4–15 seconds, at 480 or 720 pixels on the shorter edge, and supports six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16.

Features include lip-synced dialogue in multiple languages, ambient sound and music, multiple camera shots with cuts within a single clip, and prompt-controlled camera and lighting. Outputs are marked with an invisible watermark and CapCut blocks input images that contain real faces or copyrighted characters.

ByteDance itself flags limitations around detail stability, “hyper-realism,” audio distortion, multi-subject consistency, text-rendering accuracy, and complex editing effects.

How it works

Seedance 2.0 extends ByteDance’s prior work from synchronous audio-video generation to joint generation inside a unified system. The company describes the architecture as "sparse." The model supports four main tasks: reference-based generation (applying subject, motion, visual effects or style cues from references), editing (modifying specified regions, characters, actions or audio in existing video), extension (creating preceding or succeeding footage), and combination modes that pair these (for example, replacing a subject in a video with one from a reference image).

Audio and video are generated simultaneously, producing stereo dialogue, sound effects and background audio. The model generates sequential shots and cuts in a single pass rather than producing separate clips and assembling them, which helps maintain character and scene consistency.

Performance on leaderboards

Seedance 2.0 ranks at or near the top on two independent leaderboards that use blind human-vote head-to-head matchups.

  • On arena.ai, Seedance 2.0 scored 1,460 Elo for text-to-video and 1,454 Elo for image-to-video, narrowly ahead of Alibaba’s HappyHorse-1.0 (1,444 Elo in both categories). The leaderboard labels the Seedance 2.0 and HappyHorse-1.0 results as preliminary.
  • On Artificial Analysis, HappyHorse-1.0 leads in three of four video categories (image-to-video without audio, and text-to-video with and without audio), with Seedance 2.0 ranking second in those; Seedance 2.0 leads the image-to-video with synchronized audio category at 1,182 Elo, ahead of HappyHorse-1.0 at 1,168 Elo and Sky Work AI’s SkyReels V4 at 1,091 Elo.

Availability and pricing

Seedance 2.0 is available via CapCut (Jianying in China) paid tier, the Dreamina web interface, API access through ByteDance services BytePlus and Volcengine, and through third-party providers including Higgsfield.ai. Third-party pricing listed is $0.30 per second for 720-pixel output with audio, or $0.24 per second for faster processing via SeeDance 2.0 Fast.

ByteDance has not disclosed details about the model’s architecture, parameter count, training data or training methods.

Safety and copyright controversy

Shortly after Seedance 2.0’s release in China, a generated clip depicting likenesses of Tom Cruise and Brad Pitt prompted six major Hollywood studios to demand that ByteDance stop training on copyrighted material and block users from generating clips based on copyrighted works. The dispute has not been resolved. ByteDance implemented safeguards inside CapCut, but it is unclear whether those protections extend to outputs produced through third-party APIs.

Market context and why it matters

The video-generation market has shifted rapidly in recent months: some U.S.-based developers have stepped back from the consumer market while Chinese developers have released new models at an accelerated pace. In March, OpenAI announced it would discontinue the Sora app and API; reports indicated Sora’s daily active users fell from about 1 million at launch to under 500,000, while operating the service cost an estimated $1 million per day.

In April, Alibaba’s HappyHorse-1.0 appeared on independent video leaderboards and quickly rose to top positions, and Alibaba also unveiled HappyOyster, a system for generating 3D environments. Tencent open-sourced an updated version of Hunyuan 3D the same day.

ByteDance’s position is notable because it controls both a widely used editor and a generative model. CapCut — reportedly with 736 million monthly active mobile users — is one of the largest consumer AI products after ChatGPT. Integrating Seedance 2.0 into CapCut demonstrates the reach a single company can achieve by owning both creation and editing workflows.

Takeaway

OpenAI’s retreat from Sora highlights a practical reality: at current compute prices, AI-generated video is an expensive consumer product. ByteDance’s move shows how companies that combine large-scale distribution with generative technology can shape the market.