Claude Opus 5.5 is live on SeedRouter
SeedRouter Docs

MiniMax H3

Generate 4–15 second clips with sound at 768P or 2K with MiniMax H3: text to video, first and last frame, and reference to video, in MiniMax's official request format, delivered as a task.

View Markdown

MiniMax H3 (Hailuo 03) is MiniMax's multimodal video model. One request makes the picture and the sound together: a 4–15 second clip at 768P or 2K with its own dialogue, ambience and effects. The request body is MiniMax's official format with the model ID minimax-h3. Send it, keep the returned task ID, and read the finished video from the task. Frames and references go in as URLs.

Model IDs

Model IDInputsResolutionsLength
minimax-h3text, first and last frame, reference images, videos and audio768P, 2K4–15 seconds

See the model page for current prices.

Quick example

curl https://api.seedrouter.ai/v1/videos/generations \
  -H "Authorization: Bearer $SEEDROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "minimax-h3",
    "content": [
      {"type": "text", "text": "A small sailboat glides across a calm turquoise bay at sunrise, seagulls in the distance, the sound of water and wind. No text, no logos."}
    ],
    "resolution": "768P",
    "duration": 5,
    "ratio": "16:9"
  }'

Endpoint

POST https://api.seedrouter.ai/v1/videos/generations
HeaderValue
AuthorizationBearer YOUR_API_KEY
Content-Typeapplication/json

The response is a task ({"id": "task_...", "status": "processing"}), not the finished video. Poll GET /v1/tasks/{task_id} for the result. Keep API keys in server-side code.

Parameters

FieldTypeDefaultNotes
modelstringrequiredminimax-h3
contentarrayrequiredThe prompt and any media, as items described below. Exactly one non-empty text item.
resolutionenumrequired768P or 2K.
durationintegerrequiredAny integer from 4 to 15 seconds.
ratioenumby inputadaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Required for text alone, where adaptive is not accepted.
content_filterbooleantrueSeedRouter's content filter, described below. Not sent to the model.

The schema is strict: unknown fields are rejected rather than ignored. callback_url is not available; poll the task instead. extra belongs to MiniMax-H3-Max and is not accepted for this model.

Content items

typeItemrole
text{"type": "text", "text": "..."}, up to 7,000 charactersnone
image_url{"type": "image_url", "image_url": {"url": "https://..."}, "role": "..."}first_frame, last_frame or reference_image
video_url{"type": "video_url", "video_url": {"url": "https://..."}, "role": "reference_video"}reference_video
audio_url{"type": "audio_url", "audio_url": {"url": "https://..."}, "role": "reference_audio"}reference_audio

A single image without a role is the first frame. With more than one image, every image needs a role.

Modes

The items in content decide the mode; there is no mode field.

Modecontent holdsratio
Text to videoone text itemrequired, not adaptive
Image to videotext plus a first_frame and/or a last_frame imageany value is treated as adaptive: the image sets the shape
Reference to videotext plus any mix of up to 9 reference_image, 3 reference_video and 3 reference_audio itemsoptional, default adaptive

Frames and references cannot be combined in one request.

{
  "model": "minimax-h3",
  "content": [
    {"type": "text", "text": "The person in the photo speaks to camera: Follow the wind. Voice follows the reference audio."},
    {"type": "image_url", "image_url": {"url": "https://example.com/person.jpg"}, "role": "reference_image"},
    {"type": "audio_url", "audio_url": {"url": "https://example.com/voice.mp3"}, "role": "reference_audio"}
  ],
  "resolution": "2K",
  "duration": 6
}

Media inputs

Every image, video and audio item is a public HTTP(S) URL. Base64 data URIs and mm_file:// IDs are not accepted: upload the file to your own storage and pass its URL.

InputFormatLimits
ImageJPG, JPEG, PNG, WebP, HEIC, HEIF30 MB or less; each side 256–5760 px; width / height 0.4–2.5
VideoMP4, MOV (H.264 or H.265)50 MB or less; 2–15 seconds each, 15 seconds in total; each side 256–5760 px; width / height 0.4–2.5; 23.976–60 fps
AudioWAV, MP315 MB or less; 2–15 seconds each, 15 seconds in total

Reference videos are measured when the request is accepted; one whose length cannot be read is refused with nothing charged. Other limits are checked before generation, and a request that breaks one fails without charge.

Content filter

content_filter is true unless you set it. With the filter on, the text, every image and three frames of every reference video (first, middle and last) are checked before the model runs. A flagged request fails with code 60001 and is not charged; so does a request that cannot be checked. Reference audio is not checked. Set content_filter to false to skip the check; the website Playground always keeps it on.

Pricing dimensions

Check the model pricing section for current rates. MiniMax H3 is billed per second of video, at a rate set by the output resolution, plus input images beyond the first five:

billed seconds = duration + ceil(Σ each reference video's seconds)
extra images   = max(0, number of images − 5)
cost = billed seconds × rate per second + extra images × rate per image

First frames, last frames and reference images all count as images. Reference audio is not billed. Both quantities are known when the request is accepted, so the amount reserved is the amount charged. View final charges in your account usage history. Failed tasks are not charged.

Output schema

Submission returns the task:

{"id": "task_...", "model": "minimax-h3", "status": "processing", "created_at": 1789689600}

Get the task

GET https://api.seedrouter.ai/v1/tasks/{task_id}

Poll every 10–20 seconds until status is completed or failed. A network timeout while polling does not mean generation failed: keep the task ID and resume checking it. Do not create another task to check progress.

Completed task

{
  "id": "task_...",
  "model": "minimax-h3",
  "status": "completed",
  "created_at": 1789689600,
  "finished_at": 1789689740,
  "output": {
    "video_url": "https://static.seedrouter.ai/media/tasks/task_example/0.mp4",
    "resolution": "768P",
    "ratio": "16:9",
    "duration": 5
  }
}

ratio is the aspect ratio the clip was made at, including the one chosen for adaptive. In our test, a 4-second 768P clip at 16:9 came back as MP4 (H.264) at 1344 × 768 and 24 fps, with a stereo AAC track at 32 kHz.

Errors

Requests rejected before a task is created return an HTTP error with an error object and are not charged. A task that fails after acceptance returns HTTP 200 when queried, with status: "failed" and an error object.

See the shared error catalog for codes, HTTP statuses, and retry guidance.

{
  "id": "task_...",
  "model": "minimax-h3",
  "status": "failed",
  "error": {
    "code": 60001,
    "message": "the request was blocked by content moderation"
  }
}

If submission itself times out, check your tasks before submitting again: the first request may have been accepted.

Tips

  • Describe the shot and its sound together: subject, place, camera, light, dialogue and effects.
  • Write dialogue into the prompt and attach a reference audio clip when the voice should follow it.
  • Draft at 768P, then render the final shot at 2K with the same request.
  • Use a first and a last frame to control where a shot starts and ends; use references to keep a product, a person or a motion.
  • Add no text, no logos to keep invented marks out.