Model Brief
Seedance 2.5: A complete video workflow
Explore ByteDance Seed’s new video model for 30-second audiovisual stories, multimodal references, and targeted editing.
Seedance 2.5 builds on Seedance 2.0’s unified audio-video generation architecture. Its focus is no longer just producing a short clip, but carrying an idea through longer storytelling, reference-led direction, and precise revision. The result is a model designed around the full journey from concept and shot planning to generation, extension, and selective editing.
What Seedance 2.5 can do
Tell a 30-second story in one pass
Generate a synchronized audio-video sequence of up to 30 seconds, with stronger continuity across shots, scene changes, and narrative beats. Multi-round extension can continue the story while preserving its subjects, setting, pacing, and audiovisual language.
Extend one sequence into a longer work
Continue an existing result through multiple rounds instead of splitting every idea into isolated clips. The model is designed to carry characters, environments, narrative pace, visual style, sound effects, and transitions forward, making multi-minute storytelling more practical.
Direct with up to 50 references
Combine as many as 30 images, 10 videos, and 10 audio clips in one generation. Seedance 2.5 can draw on their characters, voices, environments, props, composition, motion, and style while handling multi-subject and multi-scene ideas.
Turn spatial plans into finished shots
Clay renders can act as a production scaffold for scene structure, character poses, blocking, trajectories, and camera angles. Seedance 2.5 can then apply finished materials, atmosphere, and lighting while respecting the spatial relationships in that scaffold.
Control and revise specific moments
Timestamp prompts can direct narrative beats, camera perspectives, movement, and rhythm. Targeted edits can then change a character, action, plot point, background, or camera move without rebuilding the whole sequence.
Built for more controlled production
- Clay-render references can define spatial structure, blocking, motion paths, camera angles, and physically grounded lighting.
- Expanded editing covers green-screen replacement, camera-perspective changes, and reference-based revisions.
- The model improves image, sound, motion, transitions, texture, skin, lighting, and color while keeping audio and visuals synchronized.
- Potential uses extend beyond film and advertising to education, industrial demonstrations, simulation, and synthetic training data.
- ByteDance notes that complex physical motion and interactions among multiple subjects still have room to improve.