ByteDance officially launched the Seedance 2.5 video generation model, doubling the single-video generation duration from 15 seconds to 30 seconds. The model will be gradually launched on Ji Meng AI and Dou Bao Professional Edition, and API services will also be integrated into Volcano Argo soon, opening up to multiple scenarios such as film, advertising, education, industrial manufacturing, and even autonomous driving.

image.png

The official said that Seedance 2.5 continues the unified multimodal audio-visual joint generation architecture, with core breakthroughs in the comprehensive improvement of long narrative capabilities, multimodal reference, and editing capabilities. In simple terms, AI videos are no longer just generating scattered clips, but can complete a complete creation with a beginning, middle, and end.

Organize multi-shot storytelling within 30 seconds, maintaining consistency through multiple extensions

In the official demonstration example, a clip of a singer performing in one shot fully presented the narrative logic: the camera moved from the gap of the red curtain to the warm backstage dressing room, the young female singer adjusted her earphones, staff reminded her to go on stage, then she walked through the passage and interacted with her partner, took the microphone, and finally stepped onto the stage—the camera pulled back to show the entire stadium, with the audience, lights, fluorescent sticks, and cheers all captured. The model can organize multiple logically connected shots such as setup, development, turning point, and conclusion within 30 seconds.

The multi-round extension capability is equally impressive—the model can continue to generate another 30 seconds based on an existing video, while maintaining the consistency of the main characters, scenes, visual style, as well as sounds and sound effects. In the official example, a boy ran out of the subway car holding a soccer ball, and the male lead chased and finally caught the boy, with the continuous action seamless and without any sense of detachment.

Feed 30 images and 10 videos at once, handling group storytelling effortlessly

Seedance 2.5 supports users to input up to 30 images, 10 video clips, and 10 audio clips as reference materials in one go. The model can comprehensively understand elements such as composition, scene, style, characters, and props in different materials, and accurately apply them to video generation according to instructions. In scenes with multiple people appearing together, it can simultaneously restore multiple characters' appearances and voices while keeping the main subjects stable—officially demonstrated 30-second music concert clip used a 16:9 horizontal screen with a cinematic realistic style, and the reference materials covered a pianist, a cellist, a violinist, a vocalist, an orchestra, a choir, and the audience. When AI video evolves from "special effect toys" to being able to handle complete narratives and group scheduling, the toolkit for short video creation may need to be rewritten again.