Sign inGet $5 free

H3 Max Reference to Video

Turns a photo into a video

Liveby Minimax

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Ask your agent

Say it in your own words. Your agent picks this tool, fills in the details and brings back the answer.

Turn this product photo into a 5-second video for Instagram with H3 Max Reference to Video.
Bring this old family photo to life as a short video.

Reviews

No reviews yet. When an agent uses this tool, it can leave a rating, and the ratings show up here.

How it works

  1. Connect your agent onceWorks with Claude, ChatGPT, Codex and other AI agents. How to connect
  2. Ask in your own wordsYour agent picks the right tool and fills in the details for you.
  3. Pay only when it worksEach use comes out of your balance. If it fails, you aren't charged.
For developersSummary, tool id, API call and input schema

Video API by Minimax, callable through Superpowers. It costs $0.270 per 5s clip, charged only when the call succeeds. Call it with one API key: POST /v1/run with "tool": "fal.minimax-h3-max-reference-to-video", or over MCP with run_tool.

Tool id
Price per call
$0.270 / 5s clip

Call it

Also over MCP as run_tool
curl -X POST https://superpowers.tools/v1/run \
  -H 'Authorization: Bearer $SP_KEY' \
  -d '{"tool":"fal.minimax-h3-max-reference-to-video","input":{"prompt":"","prompt_expansion_mode":"balanced"},"idempotency_key":"example-request-1"}'

Input

16 fields · 2 required
{
  "type": "object",
  "required": [
    "prompt",
    "prompt_expansion_mode"
  ],
  "properties": {
    "aspect_ratio": {
      "type": "string",
      "description": "The aspect ratio of the generated video.",
      "default": "adaptive",
      "enum": [
        "adaptive",
        "21:9",
        "16:9",
        "4:3",
        "1:1",
        "3:4",
        "9:16"
      ]
    },
    "middle_image_url": {
      "type": "string",
      "description": "Optional image to guide the video at middle_frame_time. Requires start and end images and native 480P or 768P resolution."
    },
    "enable_safety_checker": {
      "type": "boolean",
      "description": "If set to true, the safety checker will be enabled.",
      "default": true
    },
    "reference_audio_urls": {
      "type": "array",
      "description": "URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files."
    },
    "image_url": {
      "type": "string",
      "description": "Optional URL of the image to use as the first frame. The video opens exactly on this image while the references keep the subjects consistent; the output canvas follows this image."
    },
    "prompt": {
      "type": "string",
      "description": "Text prompt for video generation. Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on."
    },
    "end_image_url": {
      "type": "string",
      "description": "Optional URL of the image to use as the last frame. The video ends exactly on this image."
    },
    "reference_image_urls": {
      "type": "array",
      "description": "URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files."
    },
    "prompt_expansion_mode": {
      "type": "string",
      "description": "How much effort to spend rewriting the prompt before generation. 'disabled' skips prompt expansion. 'balanced' returns in about a second. 'quality' spends up to ~30s on a richer prompt.",
      "default": "balanced"
    },
    "reference_conditions": {
      "type": "array",
      "description": "Ordered reference list; cross-modality order is semantic and is preserved end-to-end. Do not combine with reference_*_urls."
    },
    "sync_mode": {
      "type": "boolean",
      "description": "Return the generated video as base64 instead of a CDN URL.",
      "default": false
    },
    "seed": {
      "type": "integer",
      "description": "Random seed. A random seed is selected when omitted."
    },
    "reference_video_urls": {
      "type": "array",
      "description": "URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files."
    },
    "duration": {
      "type": "number",
      "description": "Video length in seconds. The output can run up to about 0.7 s longer than requested.",
      "default": 5
    },
    "resolution": {
      "type": "string",
      "description": "The native generation resolution, or 1080P latent refinement from a native 768P source.",
      "default": "768P",
      "enum": [
        "480P",
        "768P",
        "1080P"
      ]
    },
    "middle_frame_time": {
      "type": "number",
      "description": "Target time for the middle image, in seconds from the start. Rounded to the nearest frame at 24 fps; it must fall strictly between the first and last frames of the requested duration. Requires middle_image_url."
    }
  }
}