Today, we are officially launching Seedance 2.5, the new-generation video creation model. Since the merchandise of Seedance 2.0, we person noticed a displacement successful what users expect from video creation models: from simply generating a clip to completing a imaginative work. Building connected the unified multimodal audio-video joint-generation architecture of Seedance 2.0, Seedance 2.5 centers connected foundational procreation and reference-based generation, delivering awesome breakthroughs successful long-form storytelling, multimodal reference, and editing. Grounded successful real-world usage cases, it opens up greater imaginative imagination and control, and further unlocks productivity.
Key highlights include:
Up to 30 seconds per generation, pinch multi-round extensions: Seedance 2.5 tin make high-quality, 30-second audio-video clips successful a azygous walk and supports aggregate rounds of extension. It besides improves changeable transitions and segment changes for stronger continuity successful longer videos, and delivers notable gains successful image, audio, and mobility quality, resulting successful a much natural, polished ocular value than commonly seen successful AI-generated video. As a result, users tin nutrient high-quality multi-minute contented pinch a accordant audiovisual language, bringing a complete communicative to life successful 1 take.
Fully upgraded multimodal referencing: Users tin now input up to 30 images, 10 video clips, and 10 audio clips arsenic reference materials successful a azygous pass. The exemplary besides strengthens a scope of reference capabilities, including clay render, motion, and imaginative references, enabling it to amended grasp the creator's intent and recognize analyzable ideas that span aggregate subjects, scenes, and changeable changes.
More precise and unchangeable editing capabilities: Seedance 2.5 offers timestamp-level power for targeted editing of audio and video content, notably improving ratio and controllability. The exemplary besides enhances precocious editing features, specified arsenic greenish screen, camera perspective, and reference-based editing, to meet the rigorous demands of professional, analyzable fields for illustration movie and advertising.
With advancements successful long-form storytelling, multimodal reference, and editing, Seedance 2.5 goes beyond longer single-pass video generation. The exemplary amended understands imaginative intent and delivers the travel from thought to vanished video pinch greater control. Now, we'd for illustration to induce you to watch a short imaginative film, produced end-to-end by Seedance 2.5.
您的浏览器不支持视频播放。Today, Seedance 2.5 is rolling retired connected Jimeng AI, Doubao Pro, and different platforms, pinch API entree coming soon via BytePlus ModelArk. We induce you to springiness it a effort and stock your feedback.
Project homepage:
https://seed.bytedance.com/seedance2_5
Access:
Jimeng Web -> Video Generation -> Select Seedance 2.5
Doubao Pro -> Video Generation -> Select Seedance 2.5
Seedance 2.5 extends single-pass video procreation from 15 to 30 seconds and further strengthens its storytelling successful longer videos. Within 30 seconds, the exemplary tin shape aggregate logically connected shots truthful that a communicative unfolds done setup, development, turning points, and resolution, alternatively than simply extending a azygous moment. For example, successful a one-take clip of a singer's shape performance, the exemplary portrays the afloat communicative of the vocalist interacting pinch unit successful the dressing room, past stepping done the backstage corridor, gathering the dancers, and stepping onto the shape pinch them for the performance, alternatively of only the infinitesimal of stepping connected stage.
您的浏览器不支持视频播放。T2V prompt: One-take handheld gimbal search shot. The camera slow pushes successful done a spread successful a dense reddish curtain and enters a warm-toned backstage dressing room. A young female singer, pinch her backmost to the camera, is adjusting her earpiece arsenic a unit personnel reminds her it's clip to spell on. She turns toward the camera and starts singing citypop. The camera pulls backmost and tracks her arsenic she passes done the curtain into a dim backstage corridor, interacting people pinch her dancers on the way; 1 unit personnel hands her a microphone. She and the dancers past measurement onto the stage, and the camera arcs astir to the back, gradually revealing the red-and-black shape design, LED screens, spotlights, haze, and reflective floor. The camera yet pulls retired to a wide changeable of the arena, showing the packed audience, ray boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
Thanks to the model's multi-round hold capability, users tin smoothly append consequent shots to existing video outputs. Throughout the hold process, it maintains the consistency of main characters, environments, and communicative pacing. This allows users to output videos lasting respective minutes astatine once, reducing the effort required to divided clips, many times splice footage, and hole transitions.
您的浏览器不支持视频播放。R2V prompt: Extend the video. Continue from the visuals and subjects successful @Video 1 and make different 30-second clip, keeping the characteristic subjects, scene, ocular style, and sound effects consistent. The small boy runs on the train carriage holding a shot ball. When the subway stops, the broadside doorway opens and he instantly dashes out, pinch the antheral lead chasing aft him. The 2 tally crossed the level and retired onto the street, startling passersby and vehicles on the way. The antheral lead yet catches up and grabs him. The boy looks up, aggrieved. The antheral lead's anger slow fades; he pats the boy's caput and shows a helpless smile.
In position of ocular presentation, the exemplary achieves smoother transitions betwixt camera movements. The main taxable remains unchangeable crossed aggregate cuts, and the audio and visuals stay successful sync, resulting successful highly coherent long-form videos. For example, successful a Peking Opera scene, the camera executes a graceful information cookware pursuing the lead actor's flowing sleeves, while the main taxable and inheritance stay wholly consistent. The swinging of the sleeves forms earthy arcs successful the air, intimately adhering to real-world physics.
您的浏览器不支持视频播放。R2V prompt: 16:9 widescreen, cinematic texture, azygous continuous take, soft camera movement, nary cuts. Scene reference: @Image 4. 0–5s: Open pinch a close-up of the Overlord from @Image 2. The camera slow circles his precocious assemblage and transitions into a mean shot. The Overlord spins and turns, his assemblage and backmost flags sweeping quickly past the lens to shape a earthy occlusion, and the camera follows done to Consort Yu's broadside successful @Image 1. 6–10s: The camera steadily circles Consort Yu successful a mean changeable from @Image 1, pursuing her h2o sleeves done the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She past draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The antheral warrior from @Image 3 enters pinch an aerial flip. The Overlord takes halfway shape while the warrior advances and retreats connected the other broadside successful a combat exchange. Consort Yu stands somewhat down and to the broadside of the Overlord, weaving successful water-sleeve movements to group softness against strength. The camera slow pulls backmost from a medium-close changeable of the warrior to a afloat shape view. At the end, each 3 look the assemblage and onslaught a synchronized Peking opera finale pose.
Additionally, to reside the overly artificial look often seen successful AI-generated videos, Seedance 2.5 systematically optimizes specifications specified arsenic entity textures, tegument and oculus features, lighting, and colour saturation. The exemplary besides minimizes uncontrolled occurrences successful subtitles and inheritance music, delivering last products that intimately lucifer the cinematic value of live-action footage.
Seedance 2.5 further strengthens its multimodal reference procreation capabilities. It allows users to input up to 30 images, 10 video clips, and 10 audio clips arsenic reference materials successful a azygous pass. A larger measurement and wider assortment of references tin amended seizure the user's intent, producing analyzable videos pinch much subjects, richer scenes, and much elastic camera work.
The exemplary comprehensively understands elements specified arsenic ocular composition, scenes, styles, characters, and props crossed each materials, applying them to the video procreation process arsenic instructed. Even successful analyzable scenarios for illustration multi-character shots aliases group storytelling, it tin sphere the appearances and voices of aggregate characters while keeping each subject's characteristics stable.
您的浏览器不支持视频播放。R2V prompt: A 30-second performance series successful 16:9 landscape, pinch cinematic realism, authentic performance hallway lighting and shadows, lukewarm aureate shape lighting, and the ambiance of a general classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly travel @Image 5. Reference @Images 6 to 10 for the remainder of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the assemblage seating. The lead vocalist walks from halfway shape toward the beforehand edge. The pianist is positioned by the piano. The orchestra is arranged connected some sides and toward the rear. The choir stands astatine the backmost of the stage. Open pinch a high-angle wide changeable of the afloat performance hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and originates singing. The camera people moves crossed the violin, cello, and orchestra arsenic they execute together, pinch the violin emotion agleam and the cello warm. In the second part, the choir joins in. The lead vocalist concisely makes oculus interaction pinch front-row assemblage members, who respond pinch a grin and a flimsy nod. In the closing shot, the camera pulls back. The singing ends, and the assemblage joins successful the applause.
Seedance 2.5 besides enhances circumstantial reference capabilities including clay render, motion, and imaginative referencing, giving users finer power complete subjects, actions, and camera activity successful the frame. For instance, pinch clay render referencing, users tin build a scene's spatial structure, characteristic poses, mobility paths, and camera angles utilizing textureless 3D models. The exemplary past uses this building to make the video, ensuring that the creation and blocking of analyzable shots intimately lucifer the creator's expectations. Additionally, Seedance 2.5 improves lighting control. By leveraging the spatial accusation from the clay render, it generates realistic lighting effects that travel beingness laws, specified arsenic ray root direction, colour temperature, intensity, and protector projection. This results successful much earthy ray and protector successful the last output.
您的浏览器不支持视频播放。R2V prompt: Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, taxable trajectory, and blocking. Refer to @Image 2 for characteristic design, scene, materials, lighting, color, and fairy-tale atmosphere, and render the achromatic exemplary arsenic a dreamy, warm, 3D animated short pinch a childlike imagination feel. The communicative unfolds arsenic follows: formation done a imagination entity → mythical beasts flying alongside done a oversea of clouds → a dive into the water → weaving done the heavy pinch manta rays → passing done a mirrored rift successful spacetime → picking stars from the cosmos → transforming backmost into the chamber → begetter tucking successful the broad → the image book closes and holds connected the last frame.
In video creation, users typically request to power pacing during the procreation process and besides refine specifications afterward. The nonstop 2nd an action occurs, the precise timing of a camera cut, and whether a character's activity successful a circumstantial clip requires accommodation each profoundly effect the last result. More precise and reliable editing capabilities let creators to accurately bring their ideas to life, improving ratio and reducing the costs of repetitive generation.
Seedance 2.5 supports precise contented editing via timestamps. During the procreation phase, users tin usage prompts to power the narrative, camera perspective, movement, and wide hit for a circumstantial clip frame, aligning the output much intimately pinch their imaginative intent. After generation, users tin make targeted modifications to characters, actions, aliases crippled elements wrong circumstantial clips, each while maintaining continuity and realism earlier and aft the edits.
Seedance 2.5 besides elevates aggregate editing features, specified arsenic greenish surface editing, camera position editing, and reference-based editing, to meet the rigorous demands of master fields for illustration movie and advertising. In greenish surface editing, for example, the exemplary tin switch backgrounds and show wholly different stories while keeping the main taxable intact. Furthermore, it excels astatine rendering really the taxable responds to the beingness rules of the caller environment. This includes the fluttering guidance of clothes, the authorities of hair, gait rhythm, and lighting interaction, ensuring the taxable blends harmoniously pinch the scene.
您的浏览器不支持视频播放。R2V prompt: Using @Video 1, render the green-screen background, obstacles, wardrobe, and supporting characters. 0–4s: outdoor training, switch the obstacles pinch rocks, bricks, tires, and woody crates. 4–10s: locker room, friends offering encouragement. 10–15s: world match, switch the training poles pinch original defenders and a goalkeeper, and the protagonist scores. Overall photorealistic, cinematic quality.
您的浏览器不支持视频播放。R2V prompt: Edit @Video 1. Keep the characters, actions, and ocular style unchanged. Adjust only the camera movement. A 15-second segmented camera plan: 0–4s, a micro-FPV move skims tightly past the pan, past follows the popping toast and whip-pans to the coffee; 4–7s, push successful and way laterally on the rim of the pan, pursuing the fried ovum arsenic it flips up and lands backmost successful place; 7–11s, quickly emergence to a top-down view, past descend astatine a dependable pace, sweeping crossed the sheet and keys; 11–15s, usage a handheld close-up to travel the hands pinch a accelerated lateral whip, past push successful connected the meal and propulsion backmost to a mean two-shot. Keep the full series smooth, continuous, and stable.
As the model's capabilities evolve, Seedance 2.5 is reaching deeper into broader manufacture scenarios specified arsenic acquisition and manufacturing. In education, the exemplary has begun to participate existent learning settings. For example, Seedance 2.5 tin move the humanities context, characters, and storylines down a instruction into much vivid and immersive visuals. It besides helps teachers nutrient instructional videos much efficiently, turning absurd contented — technological principles, humanities events, experimental procedures — into move demonstrations. This not only lowers the obstruction to producing acquisition materials but besides allows for highly elastic contented customization.
您的浏览器不支持视频播放。An illustration of Doubao Learning app's "Doubao Classroom" scenario
R2V prompt: Expressive Eastern painterly style. A thoroughfare segment successful Lin'an during the Southern Song dynasty. Several children tally and outcry done the bustling street, chanting, "I move around, and location he is, wherever the lantern lights turn dim." The camera follows the children arsenic they run, sweeping past the lively street. The camera past tilts up to uncover Xin Qiji from @Image 1. Xin Qiji turns his head, and successful the region stands a man among the fading lantern lights. The changeable stays continuous throughout.
In sectors for illustration business manufacturing, embodied intelligence, and autonomous driving, Seedance 2.5 is becoming integrated into highly circumstantial accumulation workflows. The exemplary tin make high-quality synthetic video information that helps train robots' cognition and manipulation skills. It is besides being utilized for business simulations, process training, and instrumentality demonstrations. For autonomous driving, the exemplary tin simulate long-tail scenarios, specified arsenic utmost upwind and analyzable postulation conditions, providing much divers samples for strategy testing and training.
您的浏览器不支持视频播放。R2V prompt: Reference the camera work, composition, changeable scale, spatial relationships, portion positions, exemplary structure, assembly order, and mobility paths from @Clay Render 1. Reference the materials, lighting, color, reflections, and ambiance from @Image 1, and move the clay render into a high-end, photorealistic car assembly sequence.
Seedance 2.5 marks a important measurement guardant successful knowing and rendering the existent world, elevating video procreation from clip-level outputs to broad imaginative workflows. At the aforesaid time, we admit location is still room for improvement, peculiarly regarding the beingness plausibility of analyzable motions and the stableness of scenes involving interactions among aggregate subjects.
Looking ahead, the Seed squad will proceed to research much coherent storytelling, present a much intuitive procreation and editing experience, and further deepen the model's grasp of real-world physics. We dream the Seedance models will go much vivid, much controllable, and amended astatine knowing users' intent, helping much users definitive their imaginative ideas while continuing to research and service broader manufacture needs.
English (US) ·
Indonesian (ID) ·