The Doubao Seedance 2.0 series (hereinafter referred to as the Seedance 2.0 series) models natively support joint audio and video generation, with exceptional semantic understanding and multimodal interaction capabilities. This article introduces the prompt usage methods and related techniques for the Seedance 2.0 series models, helping you efficiently generate high-quality videos that meet your needs.
Basic Formula
The Seedance 2.0 series models support simultaneous reference to multimodal materials such as videos, images, and audio, accurately locking in features like character appearance, action effects, visual style, and voice timbre, significantly lowering the barrier to prompt writing. Leveraging this advantage, we can guide the model using a simple basic formula, utilizing the characteristics of multimodal materials to quickly generate videos that meet specific requirements.
Reference video generation can be subdivided into three types of tasks: multimodal reference, editing videos, and extending videos. You can choose the basic prompt formula based on the task type.
Multimodal Reference
Extract some elements (such as subject, style, scene, sound effects) from the material to generate a brand-new video.
<Image N>'s <Subject N>, generate...<Video N>'s 's <Action/Camera Movement/Style/Sound Effect>, generate...<Audio N>'s timbre, generate...Editing Videos
Make partial or global modifications to the original video. Unmentioned parts remain unchanged by default.
<Element Features> + <Timing of Appearance> + <Location of Appearance><Video N>, changing its <Original Feature> to <New Feature>Extending Video
Extend the original video along the time dimension, requiring consistency in audio-visual style, subject, and narrative.
<Video N>, generate...<Video 1> + <Transition Description> + then <Video 2> + <Transition Description> + then <Video 3><Video N>" to refer to the video. Do not use "reference<Video N>" to avoid being misidentified as a reference task.Combined Tasks
The above three task types can also be used in combination.
<Image/Video N>'s [Reference Dimension], strictly edit <Video X>, [Specific Editing Content]Advanced Formula
Seedance 2.0 is essentially a multimodal AI director: it simultaneously reads your text prompts, images, videos, and audio, and internally decomposes them into a "spatial layer" (what is in the frame) and a "temporal layer" (how things change over time) to understand and generate visuals.
Therefore, a good prompt is not a mere "descriptive text" but an "engineering instruction": who, in what scene, doing what action, how the camera moves, and in what temporal order, each directed to the spatial and temporal layers. The specific formula is as follows:
Advanced Prompt Formula: Precise Subject + Action Details + Scene Environment + Lighting and Color + Camera Movement + Visual Style + Image Quality + Constraints
Simply put, first lock in "who" is doing "what," then specify "where" and "what atmosphere," then tell the model "how to shoot," and finally tighten the result with style, quality, and constraints. Below is a detailed breakdown of each element.
1 Defining the Subject
In actual reference materials, an image often contains multiple subjects. To precisely reference a specific object in the material, a clear subject definition is necessary. The subject can be a person, prop, scene, etc.
<Image/Video N>'s [Core Subject Features]
Define
<Subject N>When objects from multiple materials point to the same subject, they should be uniformly bound:
[...], Image 2's [...] defined as **<Subject N>**When there are multiple subjects in the video, each must be defined separately and distinguished using labels. Labels should be unique and stable, and must be consistently used in subsequent descriptions to avoid reference confusion.
<Subject N>@<Image N> to emphasize the binding relationship between the subject and the material. For example: Zhang San@Image 1.<Image/Video N> to refer to the subject. Since the model cannot directly associate the Asset ID with the reference content, do not directly replace <Image/Video N>.
2 Using Shot Timing
The model's internal modeling decouples space and time. Therefore, the ideal form of a complex video prompt is a timeline-based storyboard: break the video into several shots, dynamically describing each shot in the order of events: who + where + doing what + how the camera moves.
Practical Advice
Use shot order, writing a simple "Shot 1 / Shot 2 / Shot 3" storyboard for each segment, then combine them into a complete prompt.
Specific Rules
Shot 1, Shot 2, Shot 3 to organize content in the order of events (primary first, secondary later). There is no mandatory limit on the duration of each segment; let the model naturally generate the rhythm based on the plot.3 Action Description Requirements
Abstract emotion | Externalized into actions and details |
Sadness | Head lowered, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the hem of the clothes, tears welling up but not falling. |
Joy | Unable to suppress a smile, brows and eyes relaxed, steps becoming light, unconsciously humming a tune, unable to resist spinning around. |
Nervousness / Anxiety | Frequently checking watch, fingers tapping the table incessantly, rapid breathing, evasive eyes, unconsciously biting nails. |
Anger | Fists clenched tightly, jawline tense, chest heaving violently, eyes sharp as knives, words squeezed through gritted teeth. |
Relief | Letting out a long sigh of relief, completely relaxing tense shoulders, revealing a long-lost, faint smile, looking up into the distance. |
4 Camera Movement Writing
The model has a strong understanding of camera movement terms, so standard cinematography terminology can be used directly, such as "medium shot, close-up, full shot, slow push-in, smooth pan, fixed camera."
5 Quality, Style, and Constraint Words
Quality, style, and constraint words are key to controlling video generation results. They define the creative boundaries for the model, unify quality and artistic tone, and avoid visual defects and random deviations, ensuring stable and compliant final output.
1 Quality
Define image clarity, detail texture, and lighting quality to enhance the basic quality of the output.
Example: High definition, rich in detail, cinematic texture, natural colors, soft lighting
2 Style
Set the overall artistic style and visual tone to unify the artistic atmosphere of the frame.
Example: Cyberpunk cool blue-purple tones, retro film, Japanese fresh style
3 Constraint Words
Constraint words are very important; they effectively avoid visual defects, deformities, and unreasonable elements, constraining the generation boundaries and stability.
Common Constraint Word Templates:
6 Practical Examples
Demonstrate how to use the advanced formula and various elements introduced above to write prompts.
Material Preparation :
Prompt :
The girl from @Image 1 is the main character. @Image 2 serves as the dormitory scene style reference. Reference the camera movement from @Video 1.
Shot 1: At dusk, the girl @Image 1 walks briskly to the dormitory door @Image 2. The camera follows steadily in a mid-shot. Warm yellow sunlight streams into the corridor from the window. She pauses at the door, takes a deep breath, with a slightly nervous expression.
Shot 2: the girl @Image 1: Pushes open the door and enters the dormitory. The camera cuts to an indoor mid-shot. Roommates look up at her while organizing books. One of them smiles and asks {How did the exam go? Did you pass?}. The camera slowly switches between half-body close-ups of the characters.
Shot 3: the girl @Image 1: First looks down with a dejected expression. The camera gives a close-up of her face. Then she looks up, unable to hold back a smile, and laughs loudly saying {Just kidding!}. The roommates chase and playfully hit her. The camera slowly pulls back, ending on a full shot of the dormitory filled with laughter.
The entire video is in a high-definition cinematic documentary style, with warm tones and soft lighting. Characters' faces are stable and undistorted, actions are natural and smooth, with no stuttering or flickering. Ambient sound effects blend naturally with @Audio 1.
Material Preparation :
Prompt :
The woman in red from @Image 1 is the female lead. The woman in black from @Image 2 is the opponent. The scene references the cliffside bamboo forest environment from @Image 3. The overall camera movement and action rhythm reference @Video 1. The background sound effects sync with @Audio 1.
Shot 1: At dusk, the camera starts from the woman in red @Image 1A medium side shot slowly pushes in. She stands at the cliff's edge, lifting a wine flask to drink. Her robes gently sway in the mountain wind. The camera orbits half a circle around her, moving from a frontal view to her back. In the distance, a black-clad figure is faintly visible among the bamboo grove.
Shot 2The camera zooms and fades to a long shot. A drone's perspective overlooks the entire cliff and bamboo grove. The two figures stand at opposite ends of the cliff. The mountain wind lifts their robes and dust, and the rhythm quickens slightly with the drumbeat.
Shot 3The camera cuts back to a ground-level close-up. The two slowly draw their swords and face off. The woman in red @image 1 her expression shifts from nonchalant to cold and sharp. The woman in black @image 2 her gaze is resolute, the tip of her sword trembling slightly. The camera smoothly follows them as they circle each other, finally freezing on a close-up of the moment just before their swords clash.
The overall scene has a cinematic feel of misty rain and martial arts rivers and lakes. Cold tones with low saturation, a film-grain texture, and rich layers of light and shadow. The characters' faces and body proportions remain stable and undistorted. Movements are fluid and natural, without stiffness, glitches, or lag.
Other Tips
Text Generation
The Seedance 2.0 series models support generating common text. The model can automatically match the appropriate style and color based on the context, and also supports specifying the text's color, style, appearance method, timing, and position in the prompt. When writing, prioritize using common characters, and avoidUncommon CharactersandSpecial Symbolsto ensure optimal presentation. Currently supported in scenarios such as ad copy, subtitles, and speech bubbles. For specific writing methods and examples, seeText Generation.
Video Extension vs. Segmented Stitching
In practice, both methods are often combined—for example, using extension to generate a coherent dialogue, then stitching in cutaway shots or transition clips to balance immersion and pacing.
Asset Configuration Strategy
Assets are typically divided into four 'functional roles':
Recommended Configuration (Total 4-5 assets): 1-2 character images (close-up/full body) + 1 scene image + 1 camera reference video + 1 audio clip.
Important Notes
Language Standards
Dialogue language must be consistent; avoid mixing Chinese and English (except for proper nouns).
Special Character Standards
Using symbols appropriately in prompts helps the model accurately understand different types of information:
Information Type | Symbol | Example |
Music | () | (Fast-paced rock music plays in the background) |
Sound Effect | <> | <Distant dog barking> |
Dialogue | {} | {Hello, world}. If the dialogue is in a minority language (not Chinese or English), the language must be indicated, e.g., say in Japanese {こんにちは}. |
Subtitles | 【】 | 【Chapter 1: Departure】 |
Frequently Asked Questions
Character ID Drift
Typical Phenomenon
The generated character's appearance does not match the reference image, or a 'face swap' (ID drift) occurs mid-video, causing the character to resemble a celebrity and get blocked by review.
Root Cause Analysis
Insufficient effectiveness of facial reference images
Solution
Strengthen the independence and weight of facial references:
Before optimization | After optimization |
The character in the video "changes face" midway, resembling a celebrity. | The face remains consistent with the reference image throughout. |
Video Contains Subtitles
Typical Phenomenon
The prompt did not request subtitles, but the generated video contains them.
Solution
Video Contains Logo/Watermark
Typical Phenomenon
The prompt did not mention watermarks, but the generated video contains logos/watermarks from other video platforms.
Solution
Add explicit constraint instructions in the prompt: 'Do not generate watermarks,' 'Do not generate logos.'
Style Drift
Typical Phenomenon
Expecting a 2D or 3D anime style, but the input reference image is realistic, and the prompt does not emphasize the video style, causing the generated video to drift into a realistic style.
Solution
Add explicit style constraint words in the prompt, such as '2D Japanese anime style' or '3D Chinese comic style.' For more precise style control, convert the reference image to the target style before generating the video.
Before optimization | After optimization |
Xianxia style drifted into a realistic style. | Maintain a 3D Chinese anime CG xianxia style. |
Jump Cut at Video Extension Seam
Typical Phenomenon
After using the video extension function to generate a new video and stitching it with the original, there may be jumps or backtracking at the seam.
Solution
Currently, it is recommended to fix this through post-editing by aligning keyframes. Future model iterations will address this fundamentally.
Before optimization | After optimization |
At the transitions (5th second, 20th second), there is a Sudden frame jump or Content regression | Smoother transitions |
Twin Problem
Typical Phenomenon
When there are many characters in the frame and multi-view character images are used as reference assets, the generated video may easily contain two identical characters in the same frame.
Root Cause Analysis
Solution
Clearly define each character in the prompt, specifying the correspondence between characters and reference images. It is recommended to annotate the corresponding reference image after the character name and maintain consistent formatting.
Example: Zhang San (corresponding to Image 1) throws the green bankbook at standing Li Si (corresponding to Image 2).
2. Add Global Constraint Instructions
Add a fixed constraint at the end of the prompt:Throughout the video, characters with identical appearance, clothing, and accessories are strictly prohibited. Do not generate duplicate clones or twin effects. Only keep the corresponding single character in the same frame; no character duplication or replication..
3. Optimize Reference Assets
Prioritize using single, independent photos for character reference images; avoid using multi-view assets.
4. Simplify and Optimize Prompts
Do not use the full script directly as a prompt. Overly redundant content can confuse the model. Streamline irrelevant descriptions and ensure instructions are clear and focused.
Before optimization | After optimization |
At the 8th second of the video, "twins" appear. |
Video Extension Quality Degradation
Typical Phenomenon
When using a model-generated video as input for extension, quality degradation occurs. Multiple extensions compound the degradation, especially causing mottled color blocks in facial areas.
Solution
Currently, the following methods can mitigate quality degradation. Future model iterations will address this fundamentally:
Before optimization | After optimization |
Directly continue writing the model output. | Convert the model output into a white-model video before continuing. |
Special Effects Not Meeting Expectations
Typical Phenomenon
When describing specific effects via text in prompts, the generated effect may not match expectations. For example, specifying 'the number '2999' appears with a countdown animation' may result in a number scrolling effect with random jumps, not following standard countdown logic.
Solution
It is recommended to usereference videosto define effects: Input the target effect video as a reference asset to the model, allowing it to accurately understand the effect's form and motion logic, resulting in a more aligned output. For example: The appearance of the number '2999' references Video 1.
Before optimization | After optimization |
The number scrolling effect has a noticeable chaotic jumping feel. | After adding the correct number scrolling effect video, the effect met expectations. Effect Animation |
Too Many Reference Characters
Typical Phenomenon
When the number of reference characters exceeds 4, model output stability decreases. The generated video may have an incorrect number of characters (e.g., fewer or more) or duplicate characters.
Solution
Currently, the following methods can mitigate this. Future model iterations will address this fundamentally:
Before optimization | After optimization |
Input 8 reference characters, output video has 9 people. | Generate 2 images from the 8 characters, then use image-to-video. |
Noise at Video End
Typical Phenomenon
When the video contains narration, there may be abrupt clicking sounds or truncated noise at the end.
Solution
Regenerate the video, or use editing tools like CapCut to applyvolume envelopeaudio fade-out to the end of the audio track to eliminate truncation noise.
Specific Steps (Using CapCut):

Before optimization | After optimization |
Noise at the end | No noise at the end |
Inaccurate Chinese pronunciation
Typical Phenomenon
The model tends to mispronounce polyphonic characters, rare characters, and visually similar characters.
Solution
Replace easily mispronounced text with commonly used homophones with the same pronunciation to avoid pronunciation errors and restore the expected audio effect. Note that this is only an optimization method and cannot completely avoid all pronunciation issues.
Example: Rewrite '螭龙山' in the prompt as the homophone '吃龙山'.
Before optimization | After optimization |
The pronunciation of "螭" in "螭龙山" is incorrect. | After changing it to "吃龙山", the pronunciation is correct. |
Inaccurate timbre reference
Typical Phenomenon
When using a reference audio to specify a timbre, the final generated video's audio timbre may deviate significantly from the reference.
Solution
Before optimization | After optimization |
Timbre does not match the input audio. Input audio | After adding timbre feature descriptions to the prompt, the timbre matching improved significantly. "Use @audio1 to say in a low, warm, slightly gritty middle-aged male voice" |
Appendix: Prompt Examples
Showcasing prompt examples for using the Seedance 2.0 series model in different scenarios, helping you achieve more precise multimodal reference control and text generation functions.
Text Generation
Slogan
Reference prompt template:
「Text Content」+「Timing of Appearance」+「Position of Appearance」+「Manner of Appearance」,「Text Features (Color, Style)」
Multi-image Reference > Logo Reference.
Reference example:
[Final Result]
[Prompt]
Hand-drawn comic style. Three people sit together eating the fried chicken from Image 1. The atmosphere is friendly and pleasant. The image gradually blurs, and the text “Happiness in Seedance” appears in the center.
[Reference Material]

▲ Image 1
Subtitles
Reference prompt template:
Subtitles appear at the bottom of the screen. The subtitle content is "...". The subtitles must be fully synchronized with the audio rhythm.
Reference example:
[Final Result]
[Reference Material]

[Prompt]
Generate a video with voiceover. A deep, calm male voice says: "In the vast universe, our world is but a fleeting moment. Yet within it, life thrives against all odds." The scene should transition slowly from night to dawn, with stars gradually fading and the sun rising from behind the mountains. Subtitles matching the lines appear at the bottom of the screen.
[Final Result]
[Reference Material]

[Prompt]
The two people in the image are chatting in the office. The woman speaks first, saying: "You always arrive right on the dot. Do you enjoy that just-in-time feeling?" The man smiles and responds: "I have my own rhythm." The conversation is casual and natural. Corresponding subtitles appear at the bottom of the screen.
Speech bubbles
Reference prompt template:
「Character」says: “…”, a speech bubble appears around the character as they speak, with the dialogue written inside.
Reference example:
[Final Result]
[Reference Material]

[Prompt]
Image 1The two from are wearing sportswear and running on the school track. The girl looks at the boy and says confidently with a smile: "We can definitely do it!" The camera cuts to a close-up of the boy, who hesitantly replies: "Are you sure?" The camera cuts back to a medium close-up of the girl, who says cheerfully: "Yes!" The mood is bright and determined. Speech bubbles appear around the speaking characters, containing the corresponding lines.
[Final Result]
[Reference Material]

[Prompt]
Reference Image 1, Image 2 the image of the girl from. The girl is in a strawberry field. She picks one, takes a bite, smiles, and says: “This is the real deal!” A speech bubble appears around the girl with the dialogue written inside.
Image Reference
The Seedance 2.0 series models support both multi-view subject references and multi-image references such as scene images and storyboard images.
If there are requirements for image order during use, upload in order. In the prompt, use Image 1, Image 2…Image n for precise reference.
Multi-view Subject Reference
Simply specify the reference object clearly. The model can respond to instructions including but not limited to the following examples.
Product:
3C Digital Products
[Final Result]
[Reference Material]

[Prompt]
Extract the camera from Image 1, Image 2, Image 3, change the background to white. The camera is on a white table. The camera focuses on the camera in close-up, then slowly rotates around the camera as the main subject, clearly showing its front, side, and back.
Household Items
[Final Result]
[Reference Material]

[Prompt]
The background is a warm-toned home scene. A mid-shot presents the thermos cup from the reference image. The camera smoothly pushes in to a close-up of the thermos cup. A hand naturally enters the frame from outside, gently grips the cup, and lifts it. The camera follows the hand's slight rotation to showcase the cup.
Character:
[Final Result]
[Prompt]
Reference Image 1, Image 2, Image 3 the image of the woman from, generate a scene of her eating cake in a coffee shop.
[Reference Material]

Multi-image Reference
Logo Reference
[Final Result]
[Reference Material]

[Prompt]
The background is a neon-lit futuristic city sky corridor, interwoven with flying vehicles and holographic advertisements. Reference the girl in Image 2. First, use a mid-shot to show the girl releasing a silver floating lamp with a holographic projection. Then, the camera pulls back to reveal countless floating lamps. The image gradually blurs, and the Logo from Image 1 appears. The overall style is a 3D cyberpunk sci-fi animation style.
Multi-subject Reference
[Final Result]
[Reference Material]

[Prompt]
Reference the cat and dog in the image. In a cozy apartment, the dog is lying down eating dog food. The cat walks over, reaches out a paw, and touches the dog. The dog stops eating upon seeing the cat. The cat snuggles up next to the dog. The scene uses warm tones.
Multi-element Reference
[Final Result]
[Reference Material]

[Prompt]
The scene is set in
Image 4
the restaurant from, bustling with customers.
Image 1The girl from is wearing
Image 2the outfit from, organizing items on the counter.
Image 3The boy from is a customer who steps forward, intending to ask the girl for her contact information.
Image 5
The logo from remains displayed in the bottom-right corner of the frame.
Multi-panel Storyboard Reference
[Final Result]
[Reference Material]

[Prompt]
Reference the storyboard in the image to generate an intense fight scene. The compositions of each panel in the image should appear in order, followed by the two characters fighting fiercely.
Reference storyboard
[Final Result]
[Reference Material]

[Prompt]
Reference Image 3The storyboard composition from the reference: a girl is waiting for her dad to finish cooking. She says: "아빠, 배고파요! 밥 다 됐어요?" The girl's appearance references
Image 1. Then the camera pans right, switching to
Image 4
's scene and composition. The dad's appearance references
Image 2. The dad replies: "거의 다 됐어, 조금만 기다려!" Then the camera cuts back to a close-up of the daughter's slightly disappointed facial expression. She says: "아직 멀었어요? 맛있는 냄새 나는데..." Then it switches to a close-up of the dad's face. He says: "이제 진짜 금방이야. '빨리빨리' 하지 말고 손부터 씻고 와!"
Video Reference
The Seedance 2.0 series models support video reference. When using, simply specify the generated content and reference object clearly.
If there are requirements for video order during use, upload in order. In the prompt, use Video 1, Video 2…Video n for precise reference.
Motion Reference
Film and Television
[Final Result]
[Reference Material]
▲ Video 1

[Prompt]
Reference Video 1Using the character actions and camera language from, generate
Image 2and
Image 1a fight scene between, with
Image 2as the left character and
Image 1as the right character. Intense background music plays.
Marketing
[Final Result]
[Reference Material]
▲ Video 1
[Prompt]
Reference Video 1 the galloping form of the horse from, generate a golden horse galloping on the grassland. Then freeze its galloping majestic posture, turning it into a horse-shaped gold pendant.
Camera Movement Reference
[Final Result]
[Prompt]
Reference Video 1Using the camera movement from, create a concept video of a tech park. Take
Image 1the high-rise building in as the visual center, using a first-person dive perspective to convey
Image 1the tech-futuristic feel of the park in.
[Reference Material]
▲ Video 1

▲ Image 1
Effect Reference
Film and Television
[Final Result]
[Reference Material]
▲ Video 1

▲ Image 1
[Prompt]
Reference Video 1 the golden particle effect from, let the character in Image 2 play the flute while being surrounded by the same particle effect.
Gameplay effects
[Final Result]
[Reference Material]
▲ Video 1

▲ Image 1
[Prompt]
Reference Video 1Using the effects from, make
Image 1the girl in grow the same wings, with the wing generation trajectory matching.
Video Editing
The Seedance 2.0 series models support video editing, including adding, deleting, or modifying elements, extending the video forward or backward, and track completion.
If there are requirements for video order during use, upload in order. In the prompt, use Video 1, Video 2…Video n for precise reference.
Element Addition, Deletion, and Modification
Adding Elements
[Final Result]
[Reference Material]
▲ Video 1
[Prompt]
Add snacks like fried chicken and pizza to the countertop in Video 1.
Deleting Elements
[Final Result]
[Reference Material]
▲ Video 1
[Prompt]
Clear
Video 1Remove other parts and tools from the desktop, keeping it clean and tidy, with only the items in their hands on the table.
Modifying Elements
[Final Result]
[Reference Material]
▲ Video 1

▲ Image 1
[Prompt]
Replace the perfume in Video 1 with the face cream in Image 1, keeping the action and camera movement unchanged.
Video Extension
Extend Backward
[Final Result]
[Reference Material]
▲ Video 1
[Prompt]
Generate the content after Video 1. Two late men run over to them. The five finally meet and chat amicably.
Extend Forward
[Final Result]
[Reference Material]
▲ Video 1
[Prompt]
Extend Forward
Video 1, an over-the-shoulder shot of the man in white. The man in white says: “It’s not that bad. You're just stressed. Everyone goes through this, you just need to keep going.”
Track Completion
Reference example:
[Final Result]
[Prompt]
Video 1, the moment a leaf falls to the ground, it triggers a golden particle effect. A gust of wind blows, then Video 2.
[Reference Material]
▲ Video 1
▲ Video 2