MiniMax H3 提示詞指南:官方結構寫法,附真實片段實測
如何撰寫 MiniMax H3 提示詞:官方三段式結構、對白標籤、鏡頭時間軸與參考素材標籤,並用四段片段實際測試。
以 Markdown 閱讀MiniMax H3 的設計是讀取一種三段式的結構化提示詞:沿著時間軸展開的畫面與對白、整體聲音環境,以及背景音樂(如果有的話)。 MiniMax 把將一般請求轉換成這種格式的步驟稱為 H3-Context-IR,並表示它「critical to the quality of the final output」(對最終輸出品質至關重要)。[1] 以 MiniMax H3 打造的 SD Video 預設會在生成前改寫你的提示詞。在我們的四段測試片段中,一般提示詞和照官方結構寫的提示詞得到的結果同樣可用,而關閉改寫則改變了鏡頭和聲音混音。
本指南介紹官方結構,對白、切鏡和參考素材的寫法,以及測試的結果。
MiniMax H3 提示詞長什麼樣子?
在文字、圖片和關鍵影格模式下,官方指南依固定順序使用三個欄位:[2]
- integrated_multimodal_description:主體部分。依時間順序寫出視覺風格、構圖、主體、動作、運鏡、鏡頭切換、誰在說話、說了什麼,以及與畫面動作對應的聲音。
- overall_soundscape:貫穿整段片段的環境音、動作音效和非語言的人聲。
- non_diegetic_music:只有觀眾聽得到的背景音樂,沒有就寫
N/A。
以下是我們測試中照這種結構寫的一個簡短範例:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames a middle-aged baker in a flour-dusted apron behind the wooden counter of a small street bakery before sunrise, warm light from the ovens behind him. The camera pushes in with small amplitude at slow speed as the baker with a calm, slightly raspy voice (S1) places a fresh loaf on the counter, looks up at the camera and says: <d>[English] First batch of the morning.</d>
overall_soundscape: Trays clink softly inside the bakery and the doorbell rings once over a quiet street.
non_diegetic_music: N/A官方指南要求這些段落用英文撰寫,而對白、歌詞和畫面中可見的文字保留原本的語言。[3]
MiniMax H3 提示詞裡的對白怎麼寫?
給每位說話者一個固定的 ID,例如 (S1) 或 (S2),在標籤外描述聲音特質,<d>…</d> 裡只放語言標籤和要說的話:[2]
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>標籤裡的文字會逐字保留,所以要寫成它應該被說出來的樣子。同一位說話者在不同鏡頭中沿用同一個 ID,從不開口的角色不分配 ID。旁白的寫法,指南用的是「says in an off-screen voiceover」(以畫外音說),接著註明畫面中角色的嘴唇保持閉合。[2]
鏡頭和切鏡的時間怎麼控制?
把每個鏡頭標成 [Shot 1]、[Shot 2],並寫出切鏡時間:「[Shot 2] At 00:05.000, the camera cuts to a close-up of…」([Shot 2] 在 00:05.000 時,鏡頭切到……的特寫)。時間軸加總必須等於片段長度,也就是 4 到 15 秒。[2][3] 在我們的測試中,一般句子也有效:「At 3 seconds, cut to a close-up of steam rising from the bread」(在第 3 秒切到麵包冒出熱氣的特寫)在大約 3 秒處產生了一次切鏡(見下方片段 C)。
對白需要在所屬鏡頭裡留出時間。SD Video 自己的建議是每秒約 2.5 個字詞的對白,所以一段 5 秒的片段只放得下一句短句。[4]
如何在提示詞中引用圖片、影片和音訊?
參考素材透過標籤引用,依類型、依傳送順序編號:<Picture 1>、<Video 1>、<Audio 1>。官方指南要求每個標籤在提示詞的所有段落中保持一致。[3] 在完整參考模式下,官方改寫使用六個段落而不是三個:主體定義、摘要、保留分析(每個參考素材要保留什麼)、詳細描述、聲音環境和音樂。[3]
一個帶標籤的一般提示詞通常就夠了。這段片段使用了一支 3 秒的大提琴演奏參考影片,提示詞是「The cellist in <Video 1> keeps playing one long bow stroke, static camera, warm room tone, the cello note, no music bed.」(<Video 1> 中的大提琴手持續拉出一個長弓,固定鏡頭,溫暖的室內底噪,大提琴的音符,沒有背景音樂。):
SD Video 參考素材轉影片,480p,5 秒,一支參考影片。
使用多支參考影片會是什麼效果?
一支參考影片加一句簡短的提示詞。參考素材是一支 6 秒的溜冰者片段,提示詞是「The skaters in <Video 1> glide across an outdoor rink at dusk, the camera holds still, skates scraping the ice, distant chatter, no music.」(<Video 1> 中的溜冰者在黃昏的戶外溜冰場上滑行,鏡頭保持不動,冰刀刮過冰面,遠處有交談聲,沒有音樂。):
一支參考影片,480p,5 秒。
同樣的提示詞和參考素材,將 duration 設為 10:
一支參考影片,480p,10 秒。
兩支參考影片依傳送順序編號。這裡 <Video 1> 是溜冰者,<Video 2> 是一支 8.6 秒的舞蹈片段,提示詞是「The skaters in <Video 1> watch the dancer in <Video 2> perform on the ice at dusk, static camera, skates scraping, no music.」(<Video 1> 中的溜冰者在黃昏時觀看 <Video 2> 中的舞者在冰上表演,固定鏡頭,冰刀刮擦聲,沒有音樂。):
兩支參考影片,480p,5 秒。
這裡 <Video 1> 是那支 3 秒的大提琴手片段,<Video 2> 是舞蹈片段,提示詞是「The cellist in <Video 1> plays while the dancer in <Video 2> moves slowly beside her, static camera, cello note, footsteps, no music bed.」(<Video 1> 中的大提琴手演奏,<Video 2> 中的舞者在她身旁緩慢移動,固定鏡頭,大提琴的音符,腳步聲,沒有背景音樂。):
兩支參考影片,480p,5 秒。
aspect_ratio 維持 adaptive 時,影片會採用第一支參考影片的畫面形狀。舞蹈片段是直式的,所以這兩次的結果都是 480 × 832 的直式影片。兩次使用了相同的提示詞與相同的參考素材,且沒有固定 seed,因此每次執行都各自選了自己的 seed,兩次結果也不相同。提示詞是「The dancer in <Video 1> repeats the same moves in an empty studio, static camera, footsteps on wood, no music.」(<Video 1> 中的舞者在空蕩的舞蹈教室裡重複同樣的動作,固定鏡頭,木地板上的腳步聲,沒有音樂。):
同一個請求執行兩次,480p,5 秒,直式。
結構化格式比一般提示詞效果更好嗎?
我們在以 MiniMax H3 打造的 SD Video 上,用四種方式生成同一個麵包店場景:文字轉影片,480p,5 秒,16:9,seed 42。每個提示詞只跑一次,所以請把這些結果當作單次觀察,而不是平均值。
A. 一般提示詞。 麵包師傅放下麵包,在緩慢推近的鏡頭中對著鏡頭說出台詞。
A middle-aged baker in a flour-dusted apron stands behind the wooden counter of a small street bakery before sunrise. He places a fresh loaf on the counter, looks up at the camera and says in a calm, slightly raspy voice, "First batch of the morning." Medium shot, slow push-in, warm light from the ovens behind him. Trays clink softly, the doorbell rings once, no music.B. 官方結構。 也就是上方第一節展示的提示詞。結果與 A 接近:動作相同,台詞說得正確,畫面稍寬一點。
C. 帶定時切鏡的一般提示詞。 台詞在大約 2.5 秒處說完,鏡頭在 3 到 3.5 秒之間切到麵包特寫。
A middle-aged baker in a flour-dusted apron stands behind the wooden counter of a small street bakery before sunrise. He places a fresh loaf on the counter, looks up at the camera and says in a calm, slightly raspy voice, "First batch of the morning." Medium shot, warm light from the ovens. At 3 seconds, cut to a close-up of steam rising from the bread as he slices it. Trays clink softly, the knife crunches through the crust, no music.D. 關閉改寫的提示詞 A(prompt_enhancement: disabled)。鏡頭離得更遠、幾乎不動,門鈴聲先響,台詞在大約 3.5 秒處才開始。音軌安靜得多:平均音量比另外三段片段低約 20 dB。
這代表什麼:
| 提示詞 | 台詞是否說對 | 變化 |
|---|---|---|
| A. 一般 | 是 | 基準 |
| B. 官方結構 | 是 | 與 A 非常接近 |
| C. 一般,3 秒切鏡 | 是 | 切鏡落在要求的位置 |
| D. 一般,關閉改寫 | 是 | 更寬的固定鏡頭,台詞較晚,音訊安靜得多 |
服務會改寫你的提示詞時,一個清楚的一般提示詞就夠了。當你關閉改寫,或在沒有 H3-Context-IR 的情況下執行開源權重時,就要自己照結構撰寫,因為這時沒有任何環節會把你的文字轉換成模型預期的格式。
寫好 MiniMax H3 提示詞的技巧
- 讓時間軸與片段長度相符;不要為 5 秒的片段描述 10 秒的動作。[3]
- 使用具體的視覺和聲音細節,而不是「cinematic」(電影感)或「beautiful」(漂亮)這類字眼。[3]
- 描述聲音:室內底噪、動作音效以及它們的遠近。如果不想要背景音樂,就寫
no music。[4] - 對白保持簡短,每幾秒一句:自然語言寫法放在引號裡,結構化寫法放進
<d>標籤。 - 反覆調整一個鏡頭時,固定 seed,每次只改一個地方。[4]
- 使用關鍵影格時,說明首格或末格如何與動作銜接。[3]
SD Video 頁面有可以試寫提示詞的 Playground,API 參考列出所有參數,包括 prompt_enhancement。
提示詞常見問題
MiniMax H3 提示詞要用英文寫嗎?
官方指南以英文撰寫描述性段落,對白、歌詞和畫面文字保留原本的語言。[3] 其他語言的對白放在 <d> 標籤裡,並加上對應的語言標籤。
H3-Context-IR 是什麼?
它是 MiniMax 的前處理步驟,會讀取你傳送的文字和媒體,並把它們改寫成 H3 用來生成的結構化表示。它不在開源釋出的內容中;MiniMax 以 API 形式提供它,並公開提示詞指南,讓開發者可以自己實作。[1][5]
SD Video 的 prompt_enhancement 有什麼作用?
它控制生成前對提示詞的改寫:turbo(預設)、quality,或 disabled,也就是照原文送出。改寫不會更動 <Picture 1> 這類參考素材標籤。[4]
MiniMax H3 提示詞應該寫多長?
長到足以涵蓋整條時間軸。MiniMax 自己改寫後的範例,單一片段就有好幾百個字詞,SD Video 的指南也指出細節寫得多不會有壞處。[1][4]
參考資料
- MiniMax. MiniMax-H3 模型卡. 2026 年 10 月 4 日擷取自 huggingface.co。
- MiniMax. h3-prompt-writing: base mode reference(references/base-en.txt). 2026 年 10 月 4 日擷取自 github.com。
- MiniMax. H3 Prompt Writing skill. 2026 年 10 月 4 日擷取自 github.com。
- SeedRouter. SD Video API reference. 2026 年 10 月 4 日擷取自 seedrouter.ai。
- MiniMax. Create H3-Context-IR Task. 2026 年 10 月 4 日擷取自 platform.minimax.io。



