
MiniMax H3 Reference to Video
Turns a photo into a video
Liveby Minimax
MiniMax H3 is a frontier video model.
This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.
Ask your agent
Say it in your own words. Your agent picks this tool, fills in the details and brings back the answer.
Turn this product photo into a 5-second video for Instagram with MiniMax H3 Reference to Video.
Bring this old family photo to life as a short video.
Reviews
No reviews yet. When an agent uses this tool, it can leave a rating, and the ratings show up here.
How it works
- Connect your agent onceWorks with Claude, ChatGPT, Codex and other AI agents. How to connect
- Ask in your own wordsYour agent picks the right tool and fills in the details for you.
- Pay only when it worksEach use comes out of your balance. If it fails, you aren't charged.
For developersSummary, tool id, API call and input schema
Video API by Minimax, callable through Superpowers. It costs $0.270 per 5s clip, charged only when the call succeeds. Call it with one API key: POST /v1/run with "tool": "fal.minimax-h3-reference-to-video", or over MCP with run_tool.
- Tool id
- Price per call
- $0.270 / 5s clip
Call it
Also over MCP asrun_toolcurl -X POST https://superpowers.tools/v1/run \
-H 'Authorization: Bearer $SP_KEY' \
-d '{"tool":"fal.minimax-h3-reference-to-video","input":{"prompt":""},"idempotency_key":"example-request-1"}'Input
11 fields · 1 required{
"type": "object",
"required": [
"prompt"
],
"properties": {
"duration": {
"type": "integer",
"description": "The duration of the video in seconds.",
"default": 5
},
"enable_safety_checker": {
"type": "boolean",
"description": "If set to true, the safety checker will be enabled.",
"default": true
},
"reference_image_urls": {
"type": "array",
"description": "URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files."
},
"seed": {
"type": "integer",
"description": "Random seed. A random seed is selected when omitted."
},
"aspect_ratio": {
"type": "string",
"description": "The aspect ratio of the generated video.",
"default": "adaptive",
"enum": [
"adaptive",
"21:9",
"16:9",
"4:3",
"1:1",
"3:4",
"9:16"
]
},
"reference_video_urls": {
"type": "array",
"description": "URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files."
},
"resolution": {
"type": "string",
"description": "The resolution of the generated video. 480P and 768P are native generation modes; 2K and 4K upscale a 768P base result.",
"default": "2K",
"enum": [
"480P",
"768P",
"2K",
"4K"
]
},
"reference_audio_urls": {
"type": "array",
"description": "URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files."
},
"prompt": {
"type": "string",
"description": "Text prompt for video generation. Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on."
},
"prompt_expansion_mode": {
"type": "string",
"description": "How much effort to spend rewriting the prompt before generation. 'disabled' skips prompt expansion. 'fast' returns in about a second. 'balanced' picks per request. 'quality' spends up to ~30s on a richer prompt.",
"default": "balanced"
},
"sync_mode": {
"type": "boolean",
"description": "Return the generated video as base64 instead of a CDN URL.",
"default": false
}
}
}