SD Video
Generate 5–15 second clips with matching dialogue and sound with SD Video: text to video, image to video and reference to video at 480p or 768p, delivered as a task.
SD Video is SeedRouter's video generation model, built on MiniMax H3. One request makes the whole scene, picture and sound together: a 5–15 second clip with its own dialogue, ambience and effects. Send the request, keep the returned task ID, and read the finished video from the task. First frames and references go in as URLs.
Model IDs
| Model ID | Modes | Resolutions | Length |
|---|---|---|---|
sd-video | text_to_video, image_to_video, reference_to_video | 480p, 768p | 5–15 seconds |
See the model page for current prices.
Quick example
curl https://api.seedrouter.ai/v1/videos/generations \
-H "Authorization: Bearer $SEEDROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "sd-video",
"mode": "text_to_video",
"prompt": "A lighthouse keeper in a wool coat stands on a wet stone pier at dawn and says, \"The fog lifts at seven.\" Locked-off shot, waves slapping the stones, no music.",
"duration": 5,
"resolution": "768p",
"aspect_ratio": "16:9"
}'Endpoint
POST https://api.seedrouter.ai/v1/videos/generations| Header | Value |
|---|---|
| Authorization | Bearer YOUR_API_KEY |
| Content-Type | application/json |
The response is a task ({"id": "task_...", "status": "processing"}), not the finished video. Poll GET /v1/tasks/{task_id} for the result. Keep API keys in server-side code.
Parameters
| Field | Type | Default | Notes |
|---|---|---|---|
model | string | required | sd-video |
mode | enum | reference_to_video | text_to_video, image_to_video, reference_to_video. The short forms t2v, i2v and ref2va are accepted. |
prompt | string | required | 1 to 32,000 characters. Describes the picture and the sound. |
duration | integer | 5 | Any integer from 5 to 15 seconds. |
resolution | enum | 768p | 480p or 768p. |
aspect_ratio | enum | by mode | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive in reference_to_video. image_to_video follows the first frame. |
prompt_enhancement | enum | turbo | turbo, quality or disabled. |
seed | integer | random | Unsigned 32-bit. Set it to repeat a shot. |
image | reference | image_to_video only, required there. The first frame. | |
reference_images | array | [] | reference_to_video only. Up to 9. |
reference_videos | array | [] | reference_to_video only. Up to 3. |
reference_audio | array | [] | reference_to_video only. Up to 3. |
The schema is strict: unknown fields are rejected rather than ignored. callback_url and callback_id are not available; poll the task instead.
Modes
mode selects how the model is conditioned. Each mode has its own fields; a field from another mode returns 400 when it carries a value (an empty list or null passes).
| Mode | Requires | Takes |
|---|---|---|
text_to_video | prompt | the shared fields |
image_to_video | prompt and image | image as the first frame |
reference_to_video (default) | prompt and at least one reference image or video | reference_images, reference_videos, reference_audio |
image_to_video treats the image as the literal first frame, so the clip opens exactly as that still looks. To place a product or a person into a scene of your own, use reference_to_video and describe the scene around it.
References and prompt labels
In reference_to_video, list order becomes the label you use in the prompt. The first entry of reference_images is <Picture 1>, the second is <Picture 2>; the first entry of reference_videos is <Video 1>, and so on. Images, videos and audio are numbered independently. Prompt enhancement can reword the rest of the prompt but not these labels; set prompt_enhancement to disabled to keep your text as written.
{
"model": "sd-video",
"mode": "reference_to_video",
"prompt": "A supervisor wearing the harness in <Picture 1> stands still and speaks to camera.",
"reference_images": [{ "type": "url", "url": "https://example.com/harness.jpg" }],
"duration": 10
}Media inputs
Every reference, and the image_to_video first frame, is an object with a public HTTP(S) URL:
{ "type": "url", "url": "https://example.com/photo.jpg" }| Input | Maximum size |
|---|---|
| Image | 16 MB |
| Video or audio | 32 MB |
At most 9 images, 3 videos and 3 audio clips, 12 references in total. Base64 data and asset IDs are not accepted: upload the file to your own storage and pass its URL. The URL must be reachable without redirects.
Pricing dimensions
Check the model pricing section for current rates. SD Video is billed per second of video, at a rate set by the mode and the output resolution:
billed seconds = duration (text_to_video, image_to_video)
billed seconds = duration + ceil(Σ min(each reference video's seconds, 5)) (reference_to_video)
cost = billed seconds × rate per secondEach reference video adds its length up to 5 seconds; a longer clip still adds 5. Reference images and reference audio are not billed. The reference videos are measured when the request is accepted, so the amount reserved is the amount charged. View final charges in your account usage history. Failed tasks are not charged.
Output schema
Submission returns the task:
{"id": "task_...", "model": "sd-video", "status": "processing", "created_at": 1789689600}Get the task
GET https://api.seedrouter.ai/v1/tasks/{task_id}Poll every 10–20 seconds until status is completed or failed. A network timeout while polling does not mean generation failed: keep the task ID and resume checking it. Do not create another task to check progress.
Completed task
{
"id": "task_...",
"model": "sd-video",
"status": "completed",
"created_at": 1789689600,
"finished_at": 1789689720,
"output": {
"video_url": "https://static.seedrouter.ai/media/tasks/task_example/0.mp4",
"duration": 5,
"width": 1344,
"height": 768,
"aspect_ratio": "16:9",
"seed": 42
}
}The video is MP4 (H.264) at 24 fps with 32 kHz stereo AAC audio. With an explicit aspect_ratio, the size is fixed:
aspect_ratio | 768p | 480p |
|---|---|---|
21:9 | 1536 × 672 | 960 × 416 |
16:9 | 1344 × 768 | 832 × 480 |
4:3 | 1024 × 768 | 640 × 480 |
1:1 | 768 × 768 | 480 × 480 |
3:4 | 768 × 1024 | 480 × 640 |
9:16 | 768 × 1344 | 480 × 832 |
text_to_video defaults to 16:9. reference_to_video defaults to adaptive: the shape of the first reference image, or the first reference video when there are no images. image_to_video always follows the first frame, including its EXIF orientation; crop the image to change the shape. An adaptive clip keeps the source shape, scaled to the short edge of the resolution with each side rounded to a multiple of 32, and reports aspect_ratio as the reduced pixel ratio, for example 23:15.
The completed task reports the seed it used. The same prompt and seed return the same clip; omit seed and each request picks a new one.
Errors
Requests rejected before a task is created return an HTTP error with an error object and are not charged. A task that fails after acceptance returns HTTP 200 when queried, with status: "failed" and an error object.
See the shared error catalog for codes, HTTP statuses, and retry guidance.
{
"id": "task_...",
"model": "sd-video",
"status": "failed",
"error": {
"code": 60001,
"message": "The request was rejected by the content policy. Please revise the prompt or input images."
}
}If submission itself times out, check your tasks before submitting again: the first request may have been accepted.
Tips
- Write long prompts. A few hundred characters or more gives better output; there is no penalty for detail.
- Name the camera, not the mood: a body, a lens and an aperture change the image, "cinematic" does almost nothing.
- Describe the sound: room tone, effects and their distance. Say
no musicif you do not want a music bed. Dialogue fits at about 2.5 words per second. - Add
no logos, brand names, printed words or badges anywhere in frameto keep invented marks out. - Keep lettering short and spelled out exactly, and name it as the only lettering in frame.
- Prefer stillness: one subject, one place, a static camera. Close hand work and soft motion such as hair or paper are the weakest areas.
- Fix
seedand change one clause at a time to iterate on a shot.
