Topics / AI video input flexibility

Attributed viewpoint

Multi-modal input composition

Seedance 2.0 accepts up to 9 images, 3 video clips, 3 audio files, and a text prompt simultaneously—assigning distinct creative roles to each: composition from images, camera movement from video, rhythm from audio, and descriptive intent from text.

Behind the viewpoint

Translations are for reading; original excerpts remain the evidence.

How to make remarkable videos with Seedance 2.0 – Replicate blog

Original excerpt

Most video models take a text prompt and give you a clip. Seedance 2.0 works differently. You can feed it up to 9 images, 3 video clips, 3 audio files, and a text prompt. The model understands how to use each piece. You can pull the composition from a photo, the camera movement from a video clip, the rhythm from an audio track, and describe how it all works together in words.