Prompt like an editor,
not a gambler.Prompt như editor,
đừng như đánh bạc.
A field guide for generating AI dance & performance videos on Dreamina Seedance 2.0 — from the core prompt formula to fixing the three failures that kill most clips: face swap, body deformation, and broken continuity. Tài liệu thực chiến để tạo video nhảy & trình diễn bằng AI trên Dreamina Seedance 2.0 — từ công thức prompt nền tảng đến cách xử lý 3 lỗi giết chết nhiều clip nhất: đổi mặt nhân vật, biến dạng cơ thể, và sai raccord bối cảnh.
The core formulaCông thức chung
The foundation for every video. Master this before touching the fixes — most errors start with a weak base prompt.Nền tảng cho mọi video. Nắm chắc phần này trước khi học vá lỗi — đa số lỗi bắt nguồn từ một prompt nền yếu.
1.1ByteDance's official formulaCông thức gốc của ByteDance
Subject + Motion + Environment + Camera + Aesthetic + Audio
Only Subject + Motion are required — but great clips include all six. Write in natural, fluent sentences. Seedance 2.0 excels at semantic understanding; do not write fragmented tags like girl, dancing, 4k, masterpiece.Chỉ Subject + Motion là bắt buộc — nhưng clip đẹp thì nên đủ 6. Viết thành câu văn tự nhiên, mạch lạc. Seedance 2.0 hiểu ngữ nghĩa rất tốt; đừng viết kiểu tag rời rạc như girl, dancing, 4k, masterpiece.
| ELEMENTTHÀNH PHẦN | ANSWERSTRẢ LỜI CÂU HỎI | EXAMPLEVÍ DỤ |
|---|---|---|
| Subject | Who? What do they look like?Ai? Trông thế nào? | A young woman in her early twenties, high ponytail, cropped bomber jacket |
| Motion | Doing what?Làm hành động gì? | performs a sharp hip-hop combination — body rolls, arm waves, a heel pivot |
| Environment | Where?Ở đâu? | on a concert stage, LED wall behind her, haze in the air |
| Camera | What is the camera doing?Máy quay làm gì? | slow push-in on a gimbal, shallow depth of field, low angle |
| Aesthetic | Tone / look?Tông / chất hình? | cinematic teal-and-amber grade, 35mm anamorphic, film grain, 24fps |
| Audio | What do we hear?Nghe thấy gì? | bass-heavy hip-hop track, sneaker squeaks, crowd cheer, stage reverb |
1.2Scaling up: the multi-shot structureNâng cấp: cấu trúc multi-shot
The 6-element formula covers a single shot. Most real videos need multiple cuts — so the formula splits into two tiers: a global tier declared once (Style, Subject, Audio) and a shot tier repeated per timecode (Motion, Camera, Environment). This is the dominant pattern across the community dataset:Công thức 6 thành phần đủ cho video 1 shot đơn. Đa số video thực tế cần nhiều cut — nên công thức được tách làm 2 tầng: tầng global khai báo một lần (Style, Subject, Audio) và tầng shot lặp theo timecode (Motion, Camera, Environment). Đây là pattern áp đảo trong dataset cộng đồng:
[GLOBAL STYLE] {Genre/style}, {duration}s, {aspect ratio}, shot on {camera/lens}, {color grade}, 24fps, natural motion blur. [SUBJECT & REFERENCE] {Main character description}. Refer to the character in Image 1. [SHOT LIST] [00-03s] Shot 1 — {shot size + angle} Camera: {movement} Action: {subject} {motion} Environment/Lighting: {setting + light} [03-07s] Shot 2 — ... [07-12s] Shot 3 — ... [GLOBAL AUDIO] {ambient + foley + music}, synchronized to the on-screen action. [NEGATIVE] {everything that must not appear}
1.3References — feeding assets to lock the baselineReference — "nạp" asset để khoá baseline
Beyond text, Seedance 2.0 accepts images / audio / video as references (guide §1.2, §2.2–2.4):Ngoài text, Seedance 2.0 nhận ảnh / audio / video làm chuẩn tham chiếu (guide mục 1.2, 2.2–2.4):
| TYPELOẠI | OFFICIAL SYNTAXSYNTAX CHUẨN | USE WHENDÙNG KHI |
|---|---|---|
| Character (multi-view)Nhân vật (multi-view) 2.2.1 | Refer to the [Subject] from Image N to generate [Scene], maintaining consistent [Subject] features. | Keep the character consistentGiữ nhân vật |
| Multi-image (multi-element)Nhiều ảnh (đa yếu tố) 2.2.2 | Refer to / Extract / Combine the [elements] from Image N... while maintaining the consistency of [elements]. | Character + outfit + scene from separate imagesNhân vật + outfit + bối cảnh từ nhiều ảnh |
| Logo 2.2.2 | The logo from Image N remains in the bottom right corner throughout. | Branded videos — upload the logo as its own reference image, don't describe it in wordsVideo có branding — logo phải là ảnh ref riêng, đừng tả bằng chữ |
| Storyboard 2.2.2 | Refer to the storyboard in Image N... All storyboard frame compositions shall be presented in strict predefined order. | Force per-cut composition from a drawn storyboardÉp composition từng cut theo bảng phân cảnh vẽ sẵn |
| VoiceGiọng nói 2.3.1 | [Character] says: "[Dialogue]," referencing the voice from Audio N. | Lip-sync, dialogue with a reference voiceLip-sync, thoại có giọng chuẩn |
| Audio content / BGMNội dung audio / BGM 2.3.2 | [Intended Timing/Trigger Moment] + [Audio N] | Background music, syncing dance to the beatNhạc nền, sync nhịp nhảy |
| Motion from videoMotion từ video 2.4.1 | Refer to the [Motion Description] from Video N... keeping the motion details consistent. | Transfer choreography from a sample clipBê nguyên vũ đạo từ clip mẫu |
| Camera move from videoCamera move từ video 2.4.2 | Refer to the [Camera Movement Description] from Video N... keeping the cinematography consistent. | Copy a camera path (FPV dive, orbit...)Copy đường máy (FPV dive, orbit...) |
| VFX from videoVFX từ video 2.4.3 | Refer to the [VFX Effects Description] from Video N... keeping the special effects consistent. | Copy effects (particles, glowing wings...) on the exact same motion pathCopy hiệu ứng (particle, cánh sáng...) đúng quỹ đạo |
- Upload order = numbering order. Upload out of order → the entire
Image 1 / Image 2mapping in your prompt breaks.Thứ tự upload = thứ tự đánh số. Upload sai thứ tự → toàn bộ mappingImage 1 / Image 2trong prompt sai theo. - Audio cannot be uploaded alone (guide §2.3) — it must accompany an image or video.Audio không upload một mình được (guide 2.3) — phải đi kèm ảnh/video.
- The community also uses shorthand like
@image1,@1, or named refs like@character1(~400 prompts). The platform understands it, but the officialImage Nsyntax is safest.Cộng đồng còn dùng syntax tắt@image1,@1, hoặc đặt tên@character1(~400 prompt). Platform hiểu được, nhưng syntax chính chủImage Nvẫn an toàn nhất. - Motion Reference (2.4.1) is the cheat code for dance videos: if you have a choreography clip,
Refer to the dance moves from Video 1— accurate choreography and far fewer deformation errors, because the model no longer has to imagine the moves from words.Motion Reference (2.4.1) là "cheat code" cho video nhảy: có clip vũ đạo mẫu thìRefer to the dance moves from Video 1— vừa đúng vũ đạo, vừa giảm hẳn lỗi biến dạng vì model không phải "tưởng tượng" động tác từ chữ.
1.4On-screen text & dialogueText trên màn hình & thoại
Slogan / free text (guide §2.1.1):Slogan / text tự do (guide 2.1.1):
[Text Content] + [Timing] + [Positioning] + [Entrance Style], [Color, Font Style]
- Use common English vocabulary; avoid rare words and special symbols.Dùng từ tiếng Anh thông dụng; tránh từ hiếm và ký tự đặc biệt.
- Vietnamese names: enter without diacritics to avoid rendering errors (project experience).Tên tiếng Việt: nhập không dấu để tránh lỗi render (kinh nghiệm project).
Subtitles (guide §2.1.2 — ×190 community prompts; great for dance videos with lyrics/hooks):Subtitle (guide 2.1.2 — ×190 prompt cộng đồng; rất hợp video nhảy có lời hát/hook):
Display subtitles at the bottom-center with the text. The subtitles must be
perfectly synchronized with the audio rhythm and pacing.
// multi-speaker: The subtitles should appear sequentially as each character speaks.Speech bubble (guide §2.1.3, for comic/anime styles): [Character] says, "[Dialogue]." Speech bubbles appear around the character containing the spoken text.Speech bubble (guide 2.1.3, cho style comic/anime): [Character] says, "[Dialogue]." Speech bubbles appear around the character containing the spoken text.
Dialogue + lip-sync (×347 prompts with dialogue, ×48 lip-sync): write the line inside the Action with a voice tone, and close with:Thoại + lip-sync (×347 prompt có dialogue, ×48 lip-sync): viết thoại trực tiếp trong Action kèm tông giọng, và chốt bằng:
The characters exhibit expressive hand gestures, realistic facial expressions, and perfectly synchronized lip movements to the dialogue.
1.5Field-tested parameters (from the dataset)Thông số thực chiến (từ dataset)
- Duration: not just 10–15s — ×203 prompts run 20–30s. Longer clips drift more → use tight timecodes + denser identity locks, or chain clips with Extension.Duration: không chỉ 10-15s — ×203 prompt chạy 20-30s. Clip càng dài càng dễ drift → dùng timecode chặt + identity lock dày hơn, hoặc nối bằng Extension.
- Prompt language: the model handles many languages (dataset: 827 Chinese, 410 Japanese, 50 Korean prompts run fine). But on-screen text renders best in English (official best practice).Ngôn ngữ prompt: model ăn được đa ngôn ngữ (dataset có 827 prompt tiếng Trung, 410 tiếng Nhật, 50 tiếng Hàn chạy tốt). Nhưng text render trên màn hình thì tiếng Anh chuẩn nhất (best practice trong guide).
- Slow motion / speed ramps (×769 — the single most popular technique in the dataset): declare per segment, e.g.
speed ramp: slow-motion during the jump, back to real-time on landing. Warning: slow-mo exposes every hand/hair flaw → give a slow-mo segment one move only.Slow motion / speed ramp (×769 — kỹ thuật phổ biến nhất dataset): khai báo theo đoạn, ví dụspeed ramp: slow-motion during the jump, back to real-time on landing. Lưu ý: slow-mo phô mọi lỗi tay/tóc → đoạn slow-mo chỉ cho một động tác duy nhất. - Priority markers (×11, seen in the most technical prompts): flag your critical block with
[HIGHEST PRIORITY — ...]to concentrate the model's attention.Priority marker (×11, dùng ở các prompt kỹ thuật cao): đánh dấu block quan trọng nhất bằng[HIGHEST PRIORITY — ...]để dồn attention của model.
Fix №1 — Face swap / identity driftLỗi №1 — Face swap / đổi nhân vật
The character's face turns into a different person across cuts. The #1 failure by community keyword frequency:Mặt nhân vật biến thành người khác qua các cut. Lỗi số 1 theo tần suất từ khoá cộng đồng:
2.1Why it happensVì sao xảy ra
| CAUSENGUYÊN NHÂN | MECHANISMCƠ CHẾ |
|---|---|
| Only 1 reference image, 1 angleChỉ 1 ảnh ref, 1 góc | The model must guess the 3/4, profile and back views → it resamples a new faceModel phải tự đoán góc 3/4, profile, sau lưng → tự sample lại khuôn mặt mới |
| Vague description (a beautiful girl)Mô tả mơ hồ (a beautiful girl) | No semantic anchor → the model resamples from scratch on every cutKhông có neo ngữ nghĩa → mỗi cut model re-sample từ đầu |
| Multiple refs, no assignmentNhiều ảnh ref nhưng không gán rõ | The model blends features across refs → hybrid facesModel trộn feature giữa các ref → ra mặt lai |
| Character mentioned only at the topChỉ nhắc nhân vật ở đầu prompt | Attention decays → shots 3-4 drift the hardestAttention loãng dần → shot 3-4 là chỗ drift nặng nhất |
| Words like transform / turn into / morphDùng từ transform / turn into / morph | Directly triggers the model's morphing behaviorKích hoạt trực tiếp cơ chế morphing của model |
| Undeclared outfit change mid-clipĐổi outfit giữa cut không khai báo | The model reads it as a different personModel hiểu là người khác |
2.2The fix stack (ranked by impact)Bộ giải pháp (xếp theo độ hiệu quả)
1Multi-view reference — foundation #1Multi-view reference — nền tảng số 1 guide 2.2.1
Upload 3 images of the same person: front / three-quarter / back — same outfit, same lighting. This is the only thing that stops drift at angles a single image cannot cover.Upload 3 ảnh cùng 1 người: front / three-quarter / back — cùng outfit, cùng ánh sáng. Đây là thứ duy nhất chặn được drift ở các góc mà 1 ảnh không cover.
2Hard reference assignment (Reference Map)Gán cứng reference (Reference Map)
Without this line, the model treats 3 images as 3 people and blends their faces:Nếu không viết dòng này, model coi 3 ảnh là 3 người rồi lai mặt:
Image 1, Image 2 and Image 3 are three views (front / three-quarter / back) of ONE single person. They are the SAME person, not three different people. Do not mix or blend them into new faces.
3Written Identity Lock — mandatory even with imagesIdentity Lock bằng chữ — bắt buộc dù đã có ảnh
Images lock pixels; words lock semantics. Any angle the images don't cover is where the words save you:Ảnh khoá pixel, chữ khoá semantic. Góc nào ảnh không cover là chữ đỡ:
[CHARACTER IDENTITY LOCK — identical in EVERY shot]
{Name}, {age}, {gender}. Face: {shape}, {eyes}, {nose}, {lips},
{eyebrows}, {skin tone}, {distinguishing mark: mole / freckles / dimple}.
Hair: {length}, {color}, {style} — identical in every shot.
Maintain the exact same facial features, bone structure, hairstyle and
skin tone in every frame and every cut. No identity drift, no face
morphing, no age change, no beauty-filter smoothing.Tip: add 1-2 ultra-specific hard anchors (a mole under the left eye, silver hoop earrings, a watch on the LEFT wrist). The more specific, the harder to drift — this pattern repeats across the highest-quality community prompts.Mẹo: thêm 1-2 hard anchor cực cụ thể (nốt ruồi dưới mắt trái, khuyên tai bạc, đồng hồ đeo tay TRÁI). Càng cụ thể càng khó drift — pattern này lặp lại ở các prompt cộng đồng chất lượng cao nhất.
4Wardrobe LockWardrobe Lock
Outfit unchanged from first to last frame: {each item + shoes +
accessories}. No wardrobe change, no accessory swapping between wrists,
no color shift.5Re-anchor in EVERY shot blockRe-anchor ở MỖI shot block
Don't describe the character only at the top. Open every shot's Action with:Đừng chỉ mô tả nhân vật ở đầu. Mỗi shot mở đầu Action bằng:
Action: The same dancer from Image 1, same outfit, {motion}...6Fewer cuts / one-takeGiảm cut / dùng one-take
Fewer cuts = fewer chances to drift. For maximum identity safety → single continuous shot, one take, no cuts. And put the close-up in shot 1-2 (the model anchors the face best early in the clip), not at the end.Càng ít cut càng ít cơ hội drift. Ưu tiên giữ mặt tuyệt đối → single continuous shot, one take, no cuts. Và đặt close-up ở shot 1-2 (model neo mặt tốt nhất ở đầu clip), đừng để cuối.
7Two characters — preventing cross-blendingCase 2 nhân vật — chống lai chéo
Image 1–2 are two views of PERSON A (the woman) — the SAME woman. Image 3–4 are two views of PERSON B (the man) — the SAME man. Person A and Person B are two DIFFERENT people. Never blend the face of Person A with Person B. Never swap their faces, hairstyles or outfits.
8Identity negativesNegative cho identity
No face morphing, no identity drift, no face swap, no hairstyle change, no wardrobe change, no gender change, no age change.
9Priority marker — when everything above still isn't enoughPriority marker — khi mọi cách trên vẫn chưa đủ
Elevate the identity block to top priority to concentrate attention:Nâng block identity lên ưu tiên tối cao để dồn attention:
[HIGHEST PRIORITY — STRICT FACE & IDENTITY LOCK]
Seen in the most technical dataset prompts, often paired with absolutes like must remain fixed throughout the video, with no exceptions.Xuất hiện ở các prompt kỹ thuật cao nhất trong dataset, thường kèm yêu cầu tuyệt đối kiểu must remain fixed throughout the video, with no exceptions.
2.3Banned words vs. safe wordsTừ cấm vs từ nên dùng
| ✕ AVOIDTRÁNH | ✓ USE INSTEADTHAY BẰNG |
|---|---|
| transforms into / morphs into | match cut to / hard cut to |
| she becomes a different look | same character, new camera angle |
| beautiful girl / handsome man | 6-8 specific facial featuresmô tả 6-8 đặc điểm cụ thể |
| dreamy transition / dissolve | motivated cut on action |
Fix №2 — Face & body deformationLỗi №2 — Mặt & cơ thể biến dạng
Different from face swap (a different person across cuts) — this is the face/hands breaking apart within a shot: melting, warping mid-turn, extra fingers, tangled arms, flicker.Khác với face swap (đổi thành người khác giữa các cut) — lỗi này là mặt/tay vỡ hình ngay trong shot: tan chảy, méo khi quay đầu, thừa ngón, tay xoắn, chớp giật.
3.1Why it happensVì sao xảy ra
| CAUSENGUYÊN NHÂN | SYMPTOMBIỂU HIỆN |
|---|---|
| Too many fast moves in one instructionĐộng tác quá nhanh + quá phức tạp trong 1 câu lệnh | Twisted arms, extra legs during danceTay xoắn, chân thừa khi nhảy |
| Overlapping people/limbs in frameNhiều người/chi thể chồng lấn trong khung hình | Limbs merging between people, extra armsLẫn tay người này vào người kia, thừa cánh tay |
| Sudden 180° head/body turnsYêu cầu quay đầu / xoay người 180° đột ngột | Warped face on the middle framesMặt méo ở frame giữa |
| Close-ups of hands doing detailed actionsCận cảnh tay làm hành động chi tiết | Extra/missing fingersThừa/thiếu ngón |
| No physics / frame-rate declarationKhông khai báo physics / frame rate | Floating objects, rubbery motionVật thể trôi nổi, chuyển động cao su |
| Prompting for an "ultra smooth, ultra pretty" stylePrompt bắt style "quá mượt quá đẹp" | AI gloss, plastic skin, over-smoothed facesAI gloss, da nhựa, mặt bị "làm mịn" đến biến dạng |
3.2The fix stackBộ giải pháp
1Describe motion by the BEAT — don't cramMô tả động tác theo NHỊP — đừng nhồi
A 3-4 second shot should contain 2-3 clear moves max.Một shot 3-4 giây chỉ nên chứa tối đa 2-3 động tác rõ ràng.
// ✕ she spins, flips, drops to the floor, jumps up and waves in 3 seconds // ✓ she executes a body roll, then a clean heel pivot, and holds a freeze on the final beat
2Declare cinematic physics & motionKhai báo physics + motion chuẩn điện ảnh
The most repeated pattern across high-quality prompts:Pattern lặp lại nhiều nhất ở các prompt chất lượng cao:
24fps, 180-degree shutter, natural motion blur, coherent physics, realistic body mechanics, smooth fluid motion, no stutter.
3Force real texture against "plastic skin"Ép texture thật để chống "da nhựa" ×52 promptsprompt
Natural skin texture, visible pores, realistic facial expressions, no beauty-filter smoothing, no AI gloss.
4Manage the hands — the worst failure zoneQuản lý tay — vùng lỗi nặng nhất
- Avoid close-ups of hands doing complex actions; if unavoidable, give each hand one role — never two hard tasks for two hands at once.Tránh close-up tay đang làm động tác phức tạp; nếu bắt buộc, cho mỗi tay một vai trò — không cho 2 tay làm 2 việc khó cùng lúc.
- Multi-person scenes: state who stands where and touches what (
hands never losing contactfor a duet;no more than X people visiblefor a group).Cảnh nhiều người: khai báo rõ ai đứng đâu, chạm vào đâu (hands never losing contactcho duet;no more than X people visiblecho nhóm).
5Turns: prompt the process, not the resultXoay người / quay đầu: cho quá trình, đừng cho kết quả
// ✕ she suddenly faces the opposite direction // ✓ she turns 180 degrees over one full beat, her ponytail following the rotation naturally
6The anti-deformation negative block (community standard)Negative block chống biến dạng (chuẩn cộng đồng)
No extra fingers, no missing fingers, no warped hands, no extra limbs, no duplicated limbs, no bad anatomy, no melting or smeared faces, no distorted faces when turning, no flickering, no glitches, no artifacts, no rubber-like motion, no floating objects.
Fix №3 — Broken continuity / raccordLỗi №3 — Sai raccord bối cảnh
Symptoms: the setting mutates across cuts (day→night, a landmark jumps sides), a character walking right suddenly walks left, a glass refills itself, wet hair dries, the light changes direction.Biểu hiện: qua mỗi cut bối cảnh tự đổi (sáng→tối, mốc trái nhảy sang phải), nhân vật đang đi phải bỗng đi trái, ly nước đầy lại, tóc ướt khô lại, ánh sáng đổi hướng.
4.1Why it happensVì sao xảy ra
The model generates each shot almost independently. If the prompt doesn't state what must persist across cuts, the model freely re-invents the scene per shot. Raccord doesn't happen by itself — it must be written as law, in a [CONTINUITY RULES] block.Model generate từng shot gần như độc lập. Nếu prompt không khai báo cái gì phải giữ nguyên giữa các cut, model tự do "sáng tác lại" bối cảnh mỗi shot. Raccord không tự có — phải viết ra thành luật, trong block [CONTINUITY RULES].
4.2The fix stack — the [CONTINUITY RULES] blockBộ giải pháp — block [CONTINUITY RULES]
1Lock space & timeKhoá không gian & thời gian
All shots happen in the SAME location within the SAME continuous minute. Time of day never changes.
2The 180-degree axis + screen directionTrục 180° + screen direction
Rule #1 of film editing:Luật số 1 của dựng phim:
Keep one 180-degree axis: the character always enters from screen-left and moves toward screen-right. Never cross the axis. Never reverse the character's position across a cut.
3Spatial geography anchorsNeo địa lý (spatial geography)
Say exactly which landmark sits where — the model won't remember on its own:Nói rõ mốc nào ở đâu — model không tự nhớ:
The LED wall is always behind her, the drum riser is always screen-left, the audience is always off-frame screen-right.
4Lock lighting & colorKhoá ánh sáng + màu
Key light comes from upper screen-left in every shot; shadows fall to screen-right. Lighting direction, color temperature and one single color grade stay identical across all cuts.
5One-way physical progression — the detail that makes clips feel realTrạng thái vật lý tiến triển MỘT CHIỀU — chi tiết làm clip "thật" nhất
Physical state is progressive and never resets: sweat sheen, loose strands of hair and the drink level only move in one direction, never returning to the shot-1 state.
6Every cut needs a motiveCut phải có động cơ
Use motivated hard cuts and match cuts on action only — no dissolves, no morph transitions, no transformation effects, no jump cuts.
7Eyelines (with ≥2 characters)Eyeline (khi có ≥2 nhân vật)
Eyelines are always toward each other and consistent across cuts. Person A always occupies screen-left, Person B screen-right.
8One-take: the camera never teleportsVới one-take: camera không được "dịch chuyển tức thời"
The camera travels one continuous physical path — it never teleports, no jump in time.
9First Frame / Last Frame anchoringFirst Frame / Last Frame anchoring ×65 prompts — pro moveprompt — kỹ thuật pro
Explicitly declare the opening and closing frame states. With a clear destination, the whole shot converges to the right composition instead of drifting:Khai báo tường minh trạng thái frame đầu và frame cuối. Có đích rõ ràng, toàn shot hội tụ về đúng composition thay vì "trôi":
Last Frame: At 15 seconds, the camera completes the 360 orbit and settles back near the original centered low-angle composition. Subject remains frozen in the final pose, hair suspended mid-swing. No on-screen text.
Strongest application — chaining clips with perfect raccord: describe the Last Frame of clip 1 and reuse that exact description as the First Frame of clip 2 (or export the last frame as an image ref for the next clip).Ứng dụng mạnh nhất — nối clip giữ raccord: mô tả Last Frame của clip 1 rồi dùng chính mô tả đó làm First Frame của clip 2 (hoặc export frame cuối làm ảnh ref cho clip sau).
10Storyboard Reference — enforce raccord with an imageStoryboard Reference — ép raccord bằng hình guide 2.2.2 · ×205
Instead of describing each cut's composition in words, draw/assemble a storyboard as one image, then:Thay vì tả composition từng cut bằng chữ, vẽ/ghép bảng phân cảnh thành 1 ảnh rồi:
Refer to the storyboard in Image 1. All storyboard frame compositions shall be presented in strict predefined order.
The model follows each panel's layout → screen direction and geography align automatically; near-immune to raccord errors.Model bám đúng bố cục từng panel → screen direction và geography tự khớp, gần như miễn nhiễm lỗi raccord.
Master template & samplesTemplate tổng hợp & prompt mẫu
The Part 1 formula with all three protection layers installed. Block order matters — locks go BEFORE the shot list. Prompts below are production-ready; copy and adapt.Công thức Phần 1 sau khi cài đủ 3 lớp bảo vệ. Thứ tự block quan trọng — các block khoá nằm TRƯỚC shot list. Các prompt bên dưới dùng được ngay; copy và chỉnh theo nhu cầu.
TEMPLATEFull master templateTemplate đầy đủ›
Annotations show which error each block prevents.Chú thích cho biết block nào chống lỗi nào.
[GLOBAL STYLE] {style}, {duration}s, {aspect}, shot on {camera/lens}, {grade}, 24fps, 180-degree shutter, natural motion blur, coherent physics. One single color grade across all shots. [REFERENCE MAP — STRICT] ← Fix №1 Image 1, 2, 3 are three views of ONE single person — the SAME person, not three people. Do not blend them into new faces. Image 4 is the location reference ONLY; no person is taken from it. [CHARACTER IDENTITY LOCK — identical in EVERY shot] ← Fix №1 {face: 6-8 features + hard anchor + hair + build} Maintain the exact same facial features, bone structure, hairstyle and skin tone in every frame. No identity drift, no face morphing. Natural skin texture, visible pores, no beauty-filter smoothing. ← Fix №2 [WARDROBE & PROP LOCK] ← Fix №1 Outfit unchanged from first to last frame: {list}. No wardrobe change. [CONTINUITY RULES — RACCORD] ← Fix №3 Same location, same continuous minute. One 180-degree axis: {direction}. Geography: {left/right/back landmarks}. Key light from {direction} in every shot, one color grade. Physical state progressive, never resets. Motivated hard cuts and match cuts only. [SHOT LIST] ← Part 1 sub-formula [00-03s] Shot 1 — {shot size + angle} Camera: {movement} Action: The same {character} from Image 1 {2-3 moves on the beat}. ← Fix №2 Lighting: {key + direction, matching Continuity Rules} [03-07s] Shot 2 — ... (open with "The same {character} from Image 1, same outfit") ... [GLOBAL AUDIO] {ambient + foley + music}, all movement accents synchronized to the beat. [NEGATIVE] ← all 3 fixes No face morphing, no identity drift, no face swap, no hairstyle or wardrobe change; no extra fingers, no warped hands, no extra limbs, no melting faces, no flickering, no artifacts; no crossing the 180-degree line, no dissolve or morph transitions, no jump cuts; no text, no watermark, no AI gloss.
SAMPLE 01Stage dance — 15s, 4 cuts (standard multi-shot)Stage dance — 15s, 4 cut (multi-shot chuẩn)›
[GLOBAL STYLE] Cinematic live-stage performance film, 15 seconds, 16:9, shot on ARRI Alexa 35 with 35mm anamorphic lens, teal-and-amber grade, 24fps, 180-degree shutter, natural motion blur, coherent physics, fine film grain. One single color grade across all four shots. [REFERENCE MAP — STRICT] Image 1, Image 2 and Image 3 are the front, three-quarter and back views of ONE single dancer. They are the SAME person, not three people. Do not blend them into a new face. [CHARACTER IDENTITY LOCK — identical in EVERY shot] The dancer is a young woman in her early twenties. Oval face, almond dark-brown eyes, straight nose, full lips, defined eyebrows, warm light-tan skin, a small mole under the left eye. Hair: long, jet black, center-parted, tied in a high ponytail — identical in every shot. Slim athletic build. Maintain the exact same facial features, bone structure, hairstyle and skin tone in every frame and across every cut. No identity drift, no face morphing. Natural skin texture, visible pores, no beauty-filter smoothing. [WARDROBE & PROP LOCK] Outfit unchanged from first to last frame: cropped black satin bomber jacket worn open, white ribbed tank top, high-waisted black wide-leg trousers, white low-top sneakers, small silver hoop earrings, thin silver watch on the LEFT wrist. No wardrobe change, no accessory swapping between wrists. [CONTINUITY RULES — RACCORD] All four shots happen on the SAME stage within the SAME continuous 15 seconds. One 180-degree axis: the dancer always faces screen-right; the audience is always off-frame screen-right, the drum riser screen-left, the LED wall behind her. Never cross the axis. Key light from upper screen-left in every shot; shadows fall screen-right. Sweat sheen and loose strands of hair only increase, never reset. Motivated hard cuts on the beat only — no dissolves. [SHOT LIST] [00-03s] Shot 1 — Medium close-up, eye level. Camera: slow push-in on a gimbal, shallow depth of field. Action: The same dancer from Image 1 stands still, head down, then snaps her chin up on the first beat and locks eyes with the lens. Lighting: single hard spotlight from upper screen-left, haze in the air. [03-07s] Shot 2 — Full-body wide, low angle. Camera: smooth lateral dolly right, following her. Action: The same dancer from Image 1, same outfit, hits a body roll, an arm wave, then a clean heel pivot; her ponytail follows each move naturally. Lighting: same key from upper screen-left, purple rim light from behind. [07-11s] Shot 3 — Three-quarter medium shot, slightly high angle. Camera: 90-degree arc staying on the same side of the axis. Action: The same dancer from Image 1 drops into a floor freeze, holds two beats, rises with a controlled shoulder roll. Visible sweat sheen. Lighting: identical key direction; stage haze thickens. [11-15s] Shot 4 — Close-up, then slow pull back to medium. Camera: handheld with subtle organic sway, natural focus pull. Action: The same dancer from Image 1 lands the final pose on the last beat, chest rising with breath, a small confident smile. Lighting: spotlight narrows on her face from upper screen-left. Continuity: hair now visibly loosened and damp — never back to shot-1 state. [GLOBAL AUDIO] Deep bass-heavy hip-hop track, crisp snare accents, sneaker squeaks on the stage floor, fabric rustle, distant crowd cheer, natural stage reverb. All movement accents land exactly on the beat. [NEGATIVE] No face morphing, no identity drift, no face swap, no hairstyle change, no wardrobe change; no extra fingers, no warped hands, no extra limbs, no melting faces, no flickering, no artifacts; no crossing the 180-degree line, no dissolve transitions, no jump cuts; no text, no watermark, no AI gloss.
SAMPLE 02One-take street dance — 10s (safest for identity)One-take street dance — 10s (an toàn nhất cho identity)›
[GLOBAL STYLE] Photorealistic urban street film, 10 seconds, 9:16 vertical, ONE single continuous take with NO cuts, 24mm lens on a gimbal, golden-hour warm grade, 24fps, natural motion blur, coherent physics, subtle film grain. [REFERENCE MAP — STRICT] Image 1 and Image 2 are the front and three-quarter views of ONE single male dancer — the SAME person. Do not create a second person from them. [CHARACTER IDENTITY LOCK] Young man, early twenties, square jaw, dark almond eyes, thick straight eyebrows, medium-brown skin with natural texture and visible pores, short curly black fade haircut, lean athletic build. Maintain the exact same facial features, haircut and skin tone for the entire unbroken take. No identity drift, no morphing, no beauty-filter smoothing. [WARDROBE & PROP LOCK] Oversized washed-grey hoodie, black cargo pants, white high-top sneakers, black cap worn backwards, thin gold chain. Nothing changes during the take. [CONTINUITY RULES — RACCORD] One unbroken take: the camera travels one continuous physical path along the same sidewalk — it never teleports, no jump in time. The brick wall stays screen-left, street traffic stays screen-right, the low sun stays behind him creating a constant backlight rim. Time of day never changes. [ACTION + CAMERA — one continuous move] The camera starts on a low tracking shot behind his sneakers, rises smoothly to a full-body wide as the same man from Image 1 begins to pop and lock in the middle of the sidewalk. The camera orbits slowly to his three-quarter front while he executes a slide, a chest pop and a freeze — one move per beat — then pushes in to a medium close-up as he finishes with a grin and a nod to the lens. [GLOBAL AUDIO] Boom-bap beat with heavy kick, sneaker scuffs on concrete, passing cars, distant city chatter. Movement accents locked to the beat. [NEGATIVE] No cuts, no jump cuts; no face morphing, no identity drift, no hairstyle or wardrobe change; no extra fingers, no warped hands, no duplicated limbs, no flickering; no text, no watermark, no AI gloss.
SAMPLE 03Duet, two characters (hardest case: no face blending)Duet 2 nhân vật (case khó nhất: chống lai mặt)›
[GLOBAL STYLE] Cinematic contemporary dance film, 15 seconds, 16:9, 50mm lens, soft warm practical lighting, filmic low-contrast grade, 24fps, natural motion blur, coherent physics. [REFERENCE MAP — HIGHEST PRIORITY] Image 1 and Image 2 are two views of PERSON A (the woman) — the SAME woman. Image 3 and Image 4 are two views of PERSON B (the man) — the SAME man. Person A and Person B are two DIFFERENT people. Never blend the face of Person A with Person B. Never swap their faces, hairstyles or outfits. Person A is always the woman; Person B is always the man, in every frame. [IDENTITY LOCK — PERSON A] Woman, mid-twenties, heart-shaped face, large dark eyes, fair skin with natural texture, shoulder-length wavy auburn hair, slim build. Wears a flowing burgundy midi dress, bare feet. Identical in every shot. [IDENTITY LOCK — PERSON B] Man, late twenties, angular face, deep-set brown eyes, olive skin, short black hair swept back, tall lean build. Wears a loose white linen shirt, sleeves rolled to the elbow, black trousers. Identical in every shot. [CONTINUITY RULES — RACCORD] Same dance studio, same continuous minute. One 180-degree axis: PERSON A always occupies screen-LEFT, PERSON B always screen-RIGHT — never reversed across a cut. The tall window is always behind them; light always from screen-right, shadows to screen-left. Eyelines are always toward each other and consistent across cuts. During lifts, hands never lose contact — each hand has one clear role, no overlapping limbs. Motivated match cuts on movement only. [SHOT LIST] [00-04s] Shot 1 — Wide two-shot, static, eye level. Both dancers stand apart. Person A (screen-left) extends her hand; Person B (screen-right) steps toward her. Backlight from the window. [04-09s] Shot 2 — Medium two-shot, slow orbit on the same side of the axis. The same two people continue a partner sequence: one lift, one controlled descent — one movement per phrase of music. Person A remains screen-left. [09-15s] Shot 3 — Close-up on Person A, slow pull back to a full-body two-shot. Person A breathes hard, eyes toward Person B off-frame screen-right. The camera pulls back to reveal Person B catching her in the final pose. Her hair is now loose and slightly damp — never reset to the shot-1 state. [GLOBAL AUDIO] Solo piano with swelling strings, bare feet sliding on wooden floor, fabric movement, breath, natural room reverb. [NEGATIVE] No face swap between the two characters, no blending of the two faces, no identity drift, no gender change, no outfit swap; no extra limbs, no tangled or duplicated arms during lifts, no warped hands, no melting faces; no position reversal across cuts, no crossing the 180-degree line, no dissolves; no text, no watermark.
Debug guideQuy trình debug
When the output still fails: diagnose by symptom → jump back to the right part.Khi generate ra vẫn lỗi: chẩn đoán theo triệu chứng → quay lại đúng phần xử lý.
| SYMPTOMTRIỆU CHỨNG | BELONGS TOLỖI THUỘC | FIX, IN ORDERXỬ LÝ THEO THỨ TỰ |
|---|---|---|
| Face becomes a different person on a later cutMặt đổi thành người khác ở cut sau | PART 02 | ① Reduce to 2 cuts → ② add a ref image for the failing angle → ③ move [IDENTITY LOCK] to the very top → ④ add more specific hard anchors① Giảm còn 2 cut → ② bổ sung ảnh ref góc bị lỗi → ③ đẩy [IDENTITY LOCK] lên đầu prompt → ④ thêm hard anchor cụ thể hơn |
| Face/hands break apart within a shotMặt/tay vỡ hình ngay trong shot | PART 03 | ① Cut moves in that shot to 2-3 max → ② slow the turn: "over one full beat" → ③ avoid hand close-ups → ④ add the anatomy negatives① Bớt động tác trong shot đó (tối đa 2-3) → ② cho xoay chậm lại "over one full beat" → ③ tránh close-up tay → ④ bổ sung negative anatomy |
| Scene/direction mutates across cutsBối cảnh/hướng đi tự đổi giữa các cut | PART 04 | ① Write the [CONTINUITY RULES] block if missing → ② declare left/right landmarks + light direction → ③ change transitions to "hard cut on action"① Viết block [CONTINUITY RULES] nếu chưa có → ② khai báo mốc trái/phải + hướng sáng → ③ đổi transition thành "hard cut on action" |
| Clip is 90% good, only 1 element is wrongClip 90% ổn, chỉ sai 1 chi tiết | VIDEO EDITING | Element Manipulation (guide 2.5.1) — fix directly, do NOT regenerate: • Add: At [timestamp] and [spatial location] of Video 1, add [element].• Remove: Remove [element] from Video 1, keeping the rest of the video content unchanged.• Replace: Replace [element] in Video 1 with [element], with all original motions and camera work preserved.Element Manipulation (guide 2.5.1) — sửa trực tiếp, KHÔNG generate lại:• Thêm: At [timestamp] and [spatial location] of Video 1, add [element].• Xoá: Remove [element] from Video 1, keeping the rest of the video content unchanged.• Thay: Replace [element] in Video 1 with [element], with all original motions and camera work preserved. |
| Still drifting after everything aboveVẫn drift sau khi làm hết | STITCHINGKỸ THUẬT NỐI | Split into 2-3 generations, then join with Video Track Completion (guide 2.5.3) — max 3 clips, total ≤ 15s. Combine with Last Frame anchoring (Part 04-⑨) so the seams match.Tách thành 2-3 lần generate rồi nối bằng Video Track Completion (guide 2.5.3) — tối đa 3 clip, tổng ≤ 15s. Kết hợp Last Frame anchoring (Phần 04-⑨) để mối nối khớp raccord. |
| Clip is right but needs to be longerClip đúng ý nhưng cần dài hơn | STITCHINGKỸ THUẬT NỐI | Video Extension (guide 2.5.2): Extend Video 1 forward/backward + [description]. The model does NOT regenerate the original segment → the face is guaranteed to stay.Video Extension (guide 2.5.2): Extend Video 1 forward/backward + [mô tả]. Model KHÔNG re-generate đoạn gốc → mặt chắc chắn giữ nguyên. |
| Want exact choreography from a sample clipMuốn bê đúng vũ đạo từ clip mẫu | REFERENCE | Motion Reference (guide 2.4.1): Refer to the dance moves from Video 1... keeping the motion details consistent — no need to describe moves in words, and deformation drops too.Motion Reference (guide 2.4.1): Refer to the dance moves from Video 1... keeping the motion details consistent — đỡ phải tả động tác bằng chữ, giảm luôn lỗi biến dạng. |
| Want a camera path / effect from a sample clipMuốn copy đường máy / hiệu ứng từ clip mẫu | REFERENCE | Camera Motion Ref (2.4.2) / VFX Ref (2.4.3): Refer to the [camera movement / VFX effects] from Video 1... keeping the [cinematography / special effects] consistentCamera Motion Ref (2.4.2) / VFX Ref (2.4.3): Refer to the [camera movement / VFX effects] from Video 1... keeping the [cinematography / special effects] consistent |
10-second pre-generate checklistChecklist 10 giây trước khi bấm generate
remains in the bottom right corner throughout.
Tên tiếng Việt nhập không dấu khi dùng Text Rendering. Upload audio nhạc nền để sync nhịp nhảy — tăng chất lượng rõ rệt (nhớ: audio không upload một mình được, phải kèm ảnh/video). Video có branding: upload logo làm ảnh ref riêng + remains in the bottom right corner throughout.