MiniMax H3
Generate 4–15 second clips with sound at 768P or 2K with MiniMax H3: text to video, first and last frame, and reference to video, in MiniMax's official request format, delivered as a task.
MiniMax H3 (Hailuo 03) is MiniMax's multimodal video model. One request makes the picture and the sound together: a 4–15 second clip at 768P or 2K with its own dialogue, ambience and effects. The request body is MiniMax's official format with the model ID minimax-h3. Send it, keep the returned task ID, and read the finished video from the task. Frames and references go in as URLs.
Model IDs
| Model ID | Inputs | Resolutions | Length |
|---|---|---|---|
minimax-h3 | text, first and last frame, reference images, videos and audio | 768P, 2K | 4–15 seconds |
See the model page for current prices.
Quick example
curl https://api.seedrouter.ai/v1/videos/generations \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "minimax-h3",
"content": [
{"type": "text", "text": "A small sailboat glides across a calm turquoise bay at sunrise, seagulls in the distance, the sound of water and wind. No text, no logos."}
],
"resolution": "768P",
"duration": 5,
"ratio": "16:9"
}'Endpoint
POST https://api.seedrouter.ai/v1/videos/generations| Header | Value |
|---|---|
| Authorization | Bearer YOUR_API_KEY |
| Content-Type | application/json |
The response is a task ({"id": "task_...", "status": "processing"}), not the finished video. Poll GET /v1/tasks/{task_id} for the result. Keep API keys in server-side code.
Parameters
| Field | Type | Default | Notes |
|---|---|---|---|
model | string | required | minimax-h3 |
content | array | required | The prompt and any media, as items described below. Exactly one non-empty text item. |
resolution | enum | required | 768P or 2K. |
duration | integer | required | Any integer from 4 to 15 seconds. |
ratio | enum | by input | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Required for text alone, where adaptive is not accepted. |
content_filter | boolean | true | SeedRouter's content filter, described below. Not sent to the model. |
The schema is strict: unknown fields are rejected rather than ignored. callback_url is not available; poll the task instead. extra belongs to MiniMax-H3-Max and is not accepted for this model.
Content items
type | Item | role |
|---|---|---|
text | {"type": "text", "text": "..."}, up to 7,000 characters | none |
image_url | {"type": "image_url", "image_url": {"url": "https://..."}, "role": "..."} | first_frame, last_frame or reference_image |
video_url | {"type": "video_url", "video_url": {"url": "https://..."}, "role": "reference_video"} | reference_video |
audio_url | {"type": "audio_url", "audio_url": {"url": "https://..."}, "role": "reference_audio"} | reference_audio |
A single image without a role is the first frame. With more than one image, every image needs a role.
Modes
The items in content decide the mode; there is no mode field.
| Mode | content holds | ratio |
|---|---|---|
| Text to video | one text item | required, not adaptive |
| Image to video | text plus a first_frame and/or a last_frame image | any value is treated as adaptive: the image sets the shape |
| Reference to video | text plus any mix of up to 9 reference_image, 3 reference_video and 3 reference_audio items | optional, default adaptive |
Frames and references cannot be combined in one request.
{
"model": "minimax-h3",
"content": [
{"type": "text", "text": "The person in the photo speaks to camera: Follow the wind. Voice follows the reference audio."},
{"type": "image_url", "image_url": {"url": "https://example.com/person.jpg"}, "role": "reference_image"},
{"type": "audio_url", "audio_url": {"url": "https://example.com/voice.mp3"}, "role": "reference_audio"}
],
"resolution": "2K",
"duration": 6
}Media inputs
Every image, video and audio item is a public HTTP(S) URL. Base64 data URIs and mm_file:// IDs are not accepted: upload the file to your own storage and pass its URL.
| Input | Format | Limits |
|---|---|---|
| Image | JPG, JPEG, PNG, WebP, HEIC, HEIF | 30 MB or less; each side 256–5760 px; width / height 0.4–2.5 |
| Video | MP4, MOV (H.264 or H.265) | 50 MB or less; 2–15 seconds each, 15 seconds in total; each side 256–5760 px; width / height 0.4–2.5; 23.976–60 fps |
| Audio | WAV, MP3 | 15 MB or less; 2–15 seconds each, 15 seconds in total |
Reference videos are measured when the request is accepted; one whose length cannot be read is refused with nothing charged. Other limits are checked before generation, and a request that breaks one fails without charge.
Content filter
content_filter is true unless you set it. With the filter on, the text, every image and three frames of every reference video (first, middle and last) are checked before the model runs. A flagged request fails with code 60001 and is not charged; so does a request that cannot be checked. Reference audio is not checked. Set content_filter to false to skip the check; the website Playground always keeps it on.
Pricing dimensions
Check the model pricing section for current rates. MiniMax H3 is billed per second of video, at a rate set by the output resolution, plus input images beyond the first five:
billed seconds = duration + ceil(Σ each reference video's seconds)
extra images = max(0, number of images − 5)
cost = billed seconds × rate per second + extra images × rate per imageFirst frames, last frames and reference images all count as images. Reference audio is not billed. Both quantities are known when the request is accepted, so the amount reserved is the amount charged. View final charges in your account usage history. Failed tasks are not charged.
Output schema
Submission returns the task:
{"id": "task_...", "model": "minimax-h3", "status": "processing", "created_at": 1789689600}Get the task
GET https://api.seedrouter.ai/v1/tasks/{task_id}Poll every 10–20 seconds until status is completed or failed. A network timeout while polling does not mean generation failed: keep the task ID and resume checking it. Do not create another task to check progress.
Completed task
{
"id": "task_...",
"model": "minimax-h3",
"status": "completed",
"created_at": 1789689600,
"finished_at": 1789689740,
"output": {
"video_url": "https://static.seedrouter.ai/media/tasks/task_example/0.mp4",
"resolution": "768P",
"ratio": "16:9",
"duration": 5
}
}ratio is the aspect ratio the clip was made at, including the one chosen for adaptive. In our test, a 4-second 768P clip at 16:9 came back as MP4 (H.264) at 1344 × 768 and 24 fps, with a stereo AAC track at 32 kHz.
Errors
Requests rejected before a task is created return an HTTP error with an error object and are not charged. A task that fails after acceptance returns HTTP 200 when queried, with status: "failed" and an error object.
See the shared error catalog for codes, HTTP statuses, and retry guidance.
{
"id": "task_...",
"model": "minimax-h3",
"status": "failed",
"error": {
"code": 60001,
"message": "the request was blocked by content moderation"
}
}If submission itself times out, check your tasks before submitting again: the first request may have been accepted.
Tips
- Describe the shot and its sound together: subject, place, camera, light, dialogue and effects.
- Write dialogue into the prompt and attach a reference audio clip when the voice should follow it.
- Draft at
768P, then render the final shot at2Kwith the same request. - Use a first and a last frame to control where a shot starts and ends; use references to keep a product, a person or a motion.
- Add
no text, no logosto keep invented marks out.
