All posts

Seedance 2.5 Prompt Cookbook: Recipes for AI Video Creation

Published on Aug 27, 2026 47min read

Seedance 2.5 prompt cookbook

Seedance 2.5 official model page. Source: ByteDance Seed.

Seedance 2.5 is easier to use when you stop treating the prompt as a sentence and start treating it as a production brief.

The official model positioning centers on 30-second storytelling, multimodal reference control, audio-video joint generation, and editing. The practical challenge is not finding more adjectives. It is telling the model which asset controls which character, which action belongs in which phase, what must remain unchanged, and what the final frame should look like.

This cookbook turns those principles into reusable prompt patterns for text-to-video, image-to-video, multi-reference creation, video editing, extension, keyframes, storyboards, transitions, and one-take content.

Research note: This cookbook is based on the supplied prompt guide and ByteDance Seed’s public Seedance 2.5 materials. Examples are prompt patterns, not guarantees. Generation quality depends on the source assets, task complexity, available controls, and review iterations.

Seedance 2.5 quick facts

Capability Practical use
30-second generation Fit a complete short scene or a tightly staged sequence into one generation
Multi-round extension Build longer narratives while checking boundary continuity
Image, video, and audio references Bind identity, motion, environment, voice, music, or sound effects to explicit roles
Video editing Change a defined object, region, time range, or sound category while preserving the master video
First and last frames Control the beginning and ending states of a continuous action
Multi-keyframe planning Organize a sequence of distinct visual states
Audio-video generation Coordinate dialogue, music, effects, and ambient sound
Flexible reference control Separate what a reference contributes from what it must not contribute

The universal prompt formula

For a simple text-only generation, start with:

text
[Subject] in [scene and environment] performs [main action or event].
The visual style is [style, light, color, material, or mood].
The camera uses [shot size, position, movement, focus, or transition].
Sound includes [dialogue, ambience, sound effects, or music].

A compact example:

text
A ceramic artist in a quiet studio at sunrise removes a pale blue cup from a pottery wheel and places it at the center of a wooden shelf.
Soft morning light enters through the window, revealing the fine gloss of the wet clay and the clean workbench.
Start with a medium shot of the hands shaping the cup, slowly push toward the surface texture, then cut to a frontal view of the shelf.
Keep the low rotation of the wheel, wet-clay friction, and subtle studio ambience. No subtitles.

When to omit details

Do not force every field into every prompt. Omit camera, sound, or style when it is irrelevant. Generation parameters belong in the interface or API configuration, not in the prose prompt, unless the product explicitly requires otherwise.

Recipe 1: text-to-video

Use this when there are no reference assets.

text
[FORMAT]
Create a [video type] lasting approximately [duration].

[SUBJECT]
The main subject is [identity, appearance, clothing, or object structure].

[EVENT]
The subject performs [one primary action or event].

[ENVIRONMENT]
The scene takes place in [location, time, weather, spatial relationships].

[VISUAL LANGUAGE]
Use [lighting, palette, material, lens feel, texture, and overall style].

[CAMERA]
Start with [shot]. Move [direction and speed]. Focus on [subject]. End with [visible final state].

[SOUND]
Include [dialogue, ambience, effects, or music]. Use [language and accent] for dialogue. [No subtitles / subtitles: ...].

Rule: Put one major state change in each stage. Too many actions in one sentence often causes missed events or excessive cuts.

Recipe 2: reference mapping

When using multiple assets, never ask the model to “use all these references” without assigning roles.

text
@image1 defines [Character A]'s face, hairstyle, and clothing. Do not use its background.
@image2 defines [Character B]'s face and clothing. Do not transfer Character A's wardrobe.
@image3 defines [Prop A]'s structure, material, and color. It belongs only to Character A.
@image4 defines [Location A]'s layout, architecture, and light. Do not use the people shown in the image.
@video1 provides [specific movement, camera path, or rhythm]. Do not use its identity, wardrobe, or setting.
@audio1 provides [voice, dialogue, ambience, sound effect, or music].

[Character A] and [Character B] keep their identities, clothing, positions, and dialogue separate.
Use only the assets needed in each scene.

Multi-reference example

text
@image1 defines the restorer's face, hairstyle, and dark green apron.
@image2 defines the archivist's face, short hair, and beige jacket.
@image3 defines the sample box's shape and worn brass clasp.
@image4 defines the restoration room's workbench, windows, and cool morning light.
@image5 defines the gallery's layout and warm spotlighting; do not use the people shown there.
@video1 provides the action rhythm of opening a box and examining its contents; do not use its person or setting.

Scene 1: In the restoration room, the restorer opens the sample box and inspects one artifact under a desk lamp. The box remains on the restorer's right side, which appears on the left side of the image.
Scene 2: In the gallery, the archivist checks the inventory number beside a display case while holding the record board with both hands.
Keep identities, clothing, prop ownership, screen direction, and room layouts consistent. Do not place every reference subject in every scene.

Recipe 3: dialogue, music, sound effects, and subtitles

Use explicit markers when sound categories must remain separate:

text
Music: (quiet piano under the scene)
Sound effect: <a distant clock bell>
Dialogue: {Good morning. The doors are open.}
Subtitle: 【Chapter One: Arrival】

For non-English dialogue:

text
Dialogue language: natural American English.
The young woman speaks in a calm, conversational voice: {I thought you were not coming.}

To remove or preserve sound:

text
Remove the original background music. Preserve dialogue, lip movement, room ambience, and action sounds. Do not add subtitles.

Recipe 4: a 30-second story with stages

Break a dense sequence into states with visible end conditions.

text
[Goal]
Create a 30-second product story about a florist preparing and delivering one bouquet.

[Stage 1: 0-8 seconds]
The florist stands behind the workbench. Loose flowers, scissors, and wrapping paper are visible. The florist trims the stems. End state: the bouquet is held in the florist's left hand and the scissors rest on the right side of the workbench.

[Stage 2: 8-20 seconds]
Continue from Stage 1. The assistant unfolds the paper while the florist places the bouquet inside and ties a green ribbon. End state: the finished bouquet lies at the center of the workbench, ribbon facing the camera.

[Stage 3: 20-30 seconds]
The assistant lifts the bouquet and places it on the pickup shelf. End state: the bouquet is centered on the shelf and both people stand behind the workbench.

[Continuity]
Keep both identities, clothing, flower colors, prop ownership, screen direction, workbench layout, and ambient sound consistent.

Time ranges allocate rhythm; they are not frame-accurate edit points. Avoid demanding several precise actions in one second.

Recipe 5: video editing

Define one master video, one target change, and an explicit preservation boundary.

text
[Edit goal]
Edit @video1. Only from 4 to 7 seconds, change the cool blue light on the right wall to warm amber light.

[Master]
@video1 is the only master. It controls the people, room layout, actions, composition, camera movement, sound, and event order.

[Edit scope]
Change only the right wall and the area illuminated by it. Allow skin tones to respond naturally to the new light.

[Preserve]
Keep identities, clothing, facial expressions, positions, gestures, room structure, camera movement, dialogue, ambience, and event order from @video1.

Object replacement

text
Edit @video1. Replace the yellow folding lamp with the white folding lamp defined by @image1.
@video1 remains the master for the desk, book, hands, camera, occlusion, timing, and motion.
@image1 defines only the white lamp's shape, structure, and material; do not use its background.
There is exactly one lamp throughout. The replacement inherits every appearance, rotation, hand occlusion, path, speed change, and exit time of the original lamp.
Do not change the book, desk, hands, background, camera, or event order.

Recipe 6: video extension

For a backward extension, end on the original first frame. For a forward extension, begin by matching the original last frame.

text
@video1 is the source video to extend forward.
The first frame of the extension directly continues @video1's final frame: preserve the subject's pose and direction, prop position, background, camera, lighting, sound state, and motion trend.
Then let the orange paper airplane continue toward the right side of the frame and leave the image while the white curtain moves slightly.
Keep the same subject, prop structure, classroom layout, camera axis, and ambient sound. The airplane remains one continuous object and does not duplicate.

Check both sides of the boundary. The goal is natural visual, motion, and sound continuity—not pixel-identical splicing.

Recipe 7: first frame, last frame, and keyframes

text
@image1 is the first frame. It defines the perfume workshop composition, subject position, pose, props, and camera direction.
@image2 is the last frame. It defines the final composition, subject position, pose, props, and camera direction.
@image3 defines the perfumer's face, hair, and green apron without changing the first or last frame composition.
@image4 defines the glass bottle's shape, material, and label position.

The perfumer picks up the dropper and bottle, adds amber fragrance, gently shakes the bottle, closes it, and places it at the center of the table. Begin at @image1 and end at @image2.
Keep the perfumer, bottle count, table layout, warm side light, and camera axis consistent.

For multiple keyframes:

text
Use @image1 through @image4 as ordered keyframes, read in sequence.
@image1: the orange paper airplane rests on the left side of a wooden desk.
@image2: the same airplane is lifted by one hand.
@image3: it passes the window while the curtain moves right.
@image4: it lands on the middle shelf.
Use continuous motion between states. Keep the airplane's size, orange material, folds, flight direction, classroom layout, and camera axis stable.

Recipe 8: storyboard and white-model references

A storyboard is for shot order and approximate composition, not pixel-level reproduction.

text
@image1 provides a four-panel storyboard. Read left to right, top to bottom. Do not use its sketch style, labels, or placeholder characters.
@image2 defines the ceramic artist's face, short hair, and gray apron.
@image3 defines the blue-glazed cup's proportions, glaze color, and curved handle.

Shot 1: wide view of the quiet studio.
Shot 2: medium side view of the artist shaping wet clay.
Shot 3: close-up of the rim and handle connection.
Shot 4: medium close view of the finished cup placed on the shelf.
Use a realistic documentary style. Preserve wheel rotation, wet-clay friction, and studio ambience.

For a rough white model, decide what it contributes:

  • Coarse white model: motion paths, positions, camera movement, cuts, light changes, and sound rhythm.
  • Detailed white model: complete structure and motion, ready for material, color, character, or environment re-rendering.
text
@video1 is a coarse white-model reference. It provides the walking path, vehicle movement, fixed camera, one push-in, two cuts, and light changes. Do not use its gray geometry or empty setting.
@video1's tall cylinder becomes the presenter; its rectangular block becomes the moving display cart.
@image1 defines the presenter. @image2 defines the cart. @image3 defines the exhibition hall.
Keep the white model's timing, positions, camera, and cuts while rendering the final scene as a bright documentary-style technology exhibition.

Recipe 9: one-take creation from images

text
[Asset roles]
@image1: market entrance and opening environment.
@image2: the traveler walking through the street.
@image3: lantern stalls and handmade details.
@image4: three friends eating together.
@image5: night river and reflections.
@image6: the group portrait at the bridge.
@video1: editing rhythm, hand-drawn stickers, and transition style only; do not use its people or location.

[Sequence]
Show @image1 through @image6 in order: arrival, exploration, meal, walk, and group portrait.

[Motion]
Use slow push-ins and subtle parallax on environments. Add only natural blinking, head turns, raised cups, and cloth movement to people. Keep faces, clothing, stalls, table position, and bridge rails stable.

[Packaging]
Use a bright travel-short rhythm, natural occlusion transitions, and stickers only near the frame edges.

[Sound]
Keep market voices, light tableware sounds, river wind, and restrained upbeat instrumental music.

Recipe 10: seamless transition between two videos

text
@video1 is the transition-before clip. Preserve its rainy street, red umbrella, forward push-in, and rain sound.
@video2 is the transition-after clip. Preserve its circular skylight, upward camera movement, and quiet indoor reverb.

At the end of @video1, the red umbrella approaches the camera and covers the frame. The circular edge of the umbrella transforms into the metal ring of @video2's skylight. The red surface fades into white daylight.
End on @video2's opening composition and continue the camera movement smoothly upward. Rain sound fades into footsteps and indoor room tone.

Recipe 11: emotion and performance

Replace abstract emotion words with visible actions.

Weak:

text
The actor becomes emotional and hopeful.

Stronger:

text
Applause begins behind the curtain. The young actor's fingers stop on the program, and their eyes slowly turn toward the stage. Their shoulders remain tense.
After the announcement ends, they release a quiet breath, their shoulders lower, and a restrained smile appears. Their eyes become wet, but they remain still and do not leave the backstage area.

Recipe 12: camera language

A technical camera term is more reliable when translated into visible behavior.

text
Rack focus from the foreground leaves to the person in the background. The leaves gradually become soft while the person's face becomes sharp.

Tracking shot: the camera moves horizontally at the skateboarder's speed. Keep the skateboarder sharp while the wall becomes a horizontal blur.

Whip-pan transition at 5 seconds: move rapidly left until the foreground bookshelf fully covers the frame, then reveal the next scene with the same leftward motion and similar speed.

A complete production prompt

Use this as a reusable starting point for a short branded video:

text
Create a 30-second product story for a modular desk lamp.

@image1 defines the lamp's front shape, matte white housing, brass hinge, and warm light. Do not use its background.
@image2 defines the home office layout, desk position, window, and evening light. Do not use the person shown in the image.
@video1 provides the hand movement of folding and rotating the lamp; do not use its identity, room, or product.
@audio1 provides a calm neutral voice. Use natural American English dialogue: {A better desk starts with better light.}

Stage 1, 0-8 seconds: Start with the lamp folded flat on the desk. The office is quiet. A hand enters from the right and lifts the lamp arm. End with the arm upright and the light off.

Stage 2, 8-20 seconds: The hand rotates the lamp toward the notebook. The lamp turns on and creates a warm pool of light. The camera makes a slow push-in toward the hinge and illuminated page.

Stage 3, 20-30 seconds: The camera pulls back to show the complete desk. The hand leaves the frame. End with the lamp centered, the notebook readable, and the office window softly out of focus.

Keep one lamp, one hand, the desk layout, hinge structure, lighting direction, and screen direction consistent. Preserve the dialogue, subtle room tone, and a restrained warm visual style. Do not add subtitles or logos that are not present in the references.

Preflight checklist

Before submitting a Seedance 2.5 prompt, check:

  • Is the subject and main event explicit?
  • Does every reference asset have a role?
  • Did you say what each asset must not contribute?
  • Are multiple people, products, and props named separately?
  • Are assets selected by scene instead of all appearing at once?
  • Does each story stage contain one major state change?
  • Is the end state visible and testable?
  • Are identity, clothing, prop ownership, screen direction, and space consistent?
  • For editing, is there one master video and a narrow edit scope?
  • For extension, does the boundary frame, motion trend, and sound state continue naturally?
  • Are first and last frames individually defined and similarly framed?
  • Are abstract emotions expressed through visible behavior?
  • Are camera terms translated into observable motion?
  • Are subtitles, dialogue, music, and effects explicitly separated?
  • Did you avoid promising frame-perfect editing?

Capability boundaries

Seedance 2.5 prompts can improve control, but they cannot guarantee:

  • frame-perfect timing from natural-language timestamps;
  • pixel-identical continuation across clips;
  • perfectly accurate subtitles, labels, formulas, or product specifications;
  • stable identity when references are ambiguous or overloaded;
  • production-ready legal, copyright, or factual review.

For exact text, logos, subtitles, timing, and product details, use a post-production pass after generation.

Sources