MiniMax H3 prompt guide: the official structure, tested on real clips
How to write MiniMax H3 prompts: the official three-part structure, dialogue tags, shot timing and reference labels, tested on four clips.
Read as MarkdownMiniMax H3 is built to read prompts in a structured form with three parts: the picture and dialogue along a timeline, the overall soundscape, and any background music. MiniMax calls the step that turns a plain request into this form H3-Context-IR, and says it is "critical to the quality of the final output".[1] SD Video, which is built on MiniMax H3, rewrites your prompt by default before generating. In our four test clips, a plain prompt and a prompt written in the official structure gave equally usable results, while switching the rewrite off changed the shot and the sound mix.
This guide covers the official structure, how to write dialogue, cuts and references, and what the tests showed.
What does a MiniMax H3 prompt look like?
For text, image and keyframe modes, the official guide uses three fields in a fixed order:[2]
- integrated_multimodal_description: the main body. Visual style, composition, subjects, actions, camera, shot changes, who speaks and what they say, and the sounds tied to on-screen actions, in time order.
- overall_soundscape: ambience, action sounds and non-verbal human sounds across the whole clip.
- non_diegetic_music: background music that only the audience hears, or
N/Afor none.
A short example in that structure, from our tests:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames a middle-aged baker in a flour-dusted apron behind the wooden counter of a small street bakery before sunrise, warm light from the ovens behind him. The camera pushes in with small amplitude at slow speed as the baker with a calm, slightly raspy voice (S1) places a fresh loaf on the counter, looks up at the camera and says: <d>[English] First batch of the morning.</d>
overall_soundscape: Trays clink softly inside the bakery and the doorbell rings once over a quiet street.
non_diegetic_music: N/AThe official guidance asks for these sections in English, while dialogue, lyrics and visible text stay in their original language.[3]
How do you write dialogue in a MiniMax H3 prompt?
Give each speaker a stable ID such as (S1) or (S2), describe the voice outside the tag, and put only a language tag and the spoken words inside <d>…</d>:[2]
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>The words inside the tag are kept verbatim, so write them exactly as they should be spoken. A speaker keeps the same ID across shots, and characters who never speak get no ID. For a narrator, the guide uses the phrase "says in an off-screen voiceover" and then states that the on-screen character's lips stay closed.[2]
How do you time shots and cuts?
Mark each shot as [Shot 1], [Shot 2] and give the cut time: "[Shot 2] At 00:05.000, the camera cuts to a close-up of…". The timeline must add up to the clip length, which is 4 to 15 seconds.[2][3] A plain sentence also works in our test: "At 3 seconds, cut to a close-up of steam rising from the bread" produced a cut at about 3 seconds (clip C below).
Dialogue needs room inside its shot. SD Video's own guidance is about 2.5 words of dialogue per second, so a 5-second clip holds one short sentence.[4]
How do you reference images, videos and audio?
References are addressed by label, numbered per type in the order you send them: <Picture 1>, <Video 1>, <Audio 1>. The official guide asks you to keep each label the same across every section of the prompt.[3] For the full reference mode, the official rewrite uses six sections instead of three: subject definitions, a summary, a retention analysis (what to keep from each reference), the detailed description, the soundscape and the music.[3]
A plain prompt with one label is often enough. This clip used a 3-second reference video of a cellist and the prompt "The cellist in <Video 1> keeps playing one long bow stroke, static camera, warm room tone, the cello note, no music bed.":
Reference to video on SD Video, 480p, 5 seconds, one reference video.
What do more reference videos look like?
One reference video and a short prompt. The reference is a 6-second clip of skaters, and the prompt is "The skaters in <Video 1> glide across an outdoor rink at dusk, the camera holds still, skates scraping the ice, distant chatter, no music.":
One reference video, 480p, 5 seconds.
The same prompt and reference with duration set to 10:
One reference video, 480p, 10 seconds.
Two reference videos are labeled in the order they are sent. Here <Video 1> is the skaters and <Video 2> is an 8.6-second dance clip, with the prompt "The skaters in <Video 1> watch the dancer in <Video 2> perform on the ice at dusk, static camera, skates scraping, no music.":
Two reference videos, 480p, 5 seconds.
Here <Video 1> is the 3-second cellist clip and <Video 2> is the dance clip, with the prompt "The cellist in <Video 1> plays while the dancer in <Video 2> moves slowly beside her, static camera, cello note, footsteps, no music bed.":
Two reference videos, 480p, 5 seconds.
With aspect_ratio left at adaptive, the clip takes the shape of the first reference video. The dance clip is vertical, so these two takes came out vertical at 480 × 832. Both used the same prompt and the same reference without a fixed seed, so each run picked its own seed and the takes differ. The prompt was "The dancer in <Video 1> repeats the same moves in an empty studio, static camera, footsteps on wood, no music.":
The same request run twice, 480p, 5 seconds, vertical.
Does the structured format beat a plain prompt?
We ran the same bakery scene four ways on SD Video, which is built on MiniMax H3: text to video, 480p, 5 seconds, 16:9, seed 42. Each prompt ran once, so read these as single observations, not averages.
A. Plain prompt. The baker sets down the loaf and says the line to camera during a slow push-in.
A middle-aged baker in a flour-dusted apron stands behind the wooden counter of a small street bakery before sunrise. He places a fresh loaf on the counter, looks up at the camera and says in a calm, slightly raspy voice, "First batch of the morning." Medium shot, slow push-in, warm light from the ovens behind him. Trays clink softly, the doorbell rings once, no music.B. Official structure. The prompt shown in the first section above. The result is close to A: the same action, the line spoken correctly, a slightly wider frame.
C. Plain prompt with a timed cut. The line finishes at about 2.5 seconds and the shot cuts to the bread close-up between 3 and 3.5 seconds.
A middle-aged baker in a flour-dusted apron stands behind the wooden counter of a small street bakery before sunrise. He places a fresh loaf on the counter, looks up at the camera and says in a calm, slightly raspy voice, "First batch of the morning." Medium shot, warm light from the ovens. At 3 seconds, cut to a close-up of steam rising from the bread as he slices it. Trays clink softly, the knife crunches through the crust, no music.D. Prompt A with the rewrite switched off (prompt_enhancement: disabled). The camera sits further back and barely moves, the doorbell comes first, and the line starts at about 3.5 seconds. The soundtrack is much quieter: its average level measured about 20 dB below the other three clips.
What this suggests:
| Prompt | Line spoken correctly | What changed |
|---|---|---|
| A. Plain | Yes | Baseline |
| B. Official structure | Yes | Very close to A |
| C. Plain, cut at 3 s | Yes | Cut landed where asked |
| D. Plain, rewrite off | Yes | Wider static shot, line later, much quieter audio |
When the service rewrites your prompt, a clear plain prompt is enough. Write the structure yourself when you switch the rewrite off, or when you run the open weights without H3-Context-IR, because then nothing converts your text into the form the model expects.
Tips for better MiniMax H3 prompts
- Match the timeline to the clip length; do not describe 10 seconds of action for a 5-second clip.[3]
- Use concrete visual and audio details instead of words like "cinematic" or "beautiful".[3]
- Describe the sound: room tone, action sounds and how far away they are. Say
no musicif you do not want a music bed.[4] - Keep dialogue short, one sentence per few seconds: in quotes in a plain prompt, inside
<d>tags in the structured form. - Fix the seed and change one thing at a time when you iterate on a shot.[4]
- With keyframes, say how the first or last frame connects to the action.[3]
The SD Video page has a playground for trying prompts, and the API reference lists every parameter, including prompt_enhancement.
Prompting questions
Should MiniMax H3 prompts be in English?
The official guidance writes the descriptive sections in English and keeps dialogue, lyrics and on-screen text in their original language.[3] Dialogue in another language goes inside the <d> tag with its language tag.
What is H3-Context-IR?
It is MiniMax's preprocessing step that reads the text and media you send and rewrites them into the structured representation H3 generates from. It is not in the open-source release; MiniMax offers it as an API and publishes the prompt guidance so developers can build their own.[1][5]
What does prompt_enhancement do on SD Video?
It controls the rewrite of your prompt before generation: turbo (the default), quality, or disabled to send your text as written. The rewrite does not change reference labels such as <Picture 1>.[4]
How long should a MiniMax H3 prompt be?
Long enough to cover the whole timeline. MiniMax's own rewritten examples run to several hundred words for a single clip, and SD Video's guidance notes there is no penalty for detail.[1][4]
References
- MiniMax. MiniMax-H3 model card. Retrieved October 4, 2026 from huggingface.co.
- MiniMax. h3-prompt-writing: base mode reference (references/base-en.txt). Retrieved October 4, 2026 from github.com.
- MiniMax. H3 Prompt Writing skill. Retrieved October 4, 2026 from github.com.
- SeedRouter. SD Video API reference. Retrieved October 4, 2026 from seedrouter.ai.
- MiniMax. Create H3-Context-IR Task. Retrieved October 4, 2026 from platform.minimax.io.



