话题 / AI视频输入灵活性

观点转述

多模态输入组合

Seedance 2.0可同时接收最多9张图像、3段视频片段、3个音频文件及一条文本提示,并为每类输入赋予明确的创意角色:图像提供构图,视频提供镜头运动,音频提供节奏,文本提供描述性意图。

观点背后的信息

译文仅辅助阅读;核查观点请以原始摘录为准。

如何使用Seedance 2.0制作惊艳视频——Replicate博客

大多数视频模型仅接受文本提示并输出一段视频。Seedance 2.0的工作方式则不同:用户可向其输入最多9张图像、3段视频片段、3个音频文件及一条文本提示。该模型能理解如何分别利用每一类输入——可从一张照片提取构图,从一段视频片段提取镜头运动,从一段音频轨道提取节奏,并用文字描述各要素如何协同运作。

原始摘录
Most video models take a text prompt and give you a clip. Seedance 2.0 works differently. You can feed it up to 9 images, 3 video clips, 3 audio files, and a text prompt. The model understands how to use each piece. You can pull the composition from a photo, the camera movement from a video clip, the rhythm from an audio track, and describe how it all works together in words.