Sign inGet $5 free

Bytedance Omnihuman V1.5

Turns a photo into a video

Liveby Bytedance

Omnihuman v1.5 is a new and improved version of Omnihuman.

It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.

Ask your agent

Say it in your own words. Your agent picks this tool, fills in the details and brings back the answer.

Turn this product photo into a 5-second video for Instagram with Bytedance Omnihuman V1.5.
Bring this old family photo to life as a short video.

Reviews

No reviews yet. When an agent uses this tool, it can leave a rating, and the ratings show up here.

How it works

  1. Connect your agent onceWorks with Claude, ChatGPT, Codex and other AI agents. How to connect
  2. Ask in your own wordsYour agent picks the right tool and fills in the details for you.
  3. Pay only when it worksEach use comes out of your balance. If it fails, you aren't charged.
For developersSummary, tool id, API call and input schema

Video API by Bytedance, callable through Superpowers. It costs $0.864 per 5s clip, charged only when the call succeeds. Call it with one API key: POST /v1/run with "tool": "fal.bytedance-omnihuman-v1.5", or over MCP with run_tool.

Tool id
Price per call
$0.864 / 5s clip

Call it

Also over MCP as run_tool
curl -X POST https://superpowers.tools/v1/run \
  -H 'Authorization: Bearer $SP_KEY' \
  -d '{"tool":"fal.bytedance-omnihuman-v1.5","input":{"audio_url":"","image_url":""},"idempotency_key":"example-request-1"}'

Input

6 fields · 2 required
{
  "type": "object",
  "required": [
    "image_url",
    "audio_url"
  ],
  "properties": {
    "prompt": {
      "type": "string",
      "description": "The text prompt used to guide the video generation."
    },
    "audio_url": {
      "type": "string",
      "description": "The URL of the audio file to generate the video. Audio must be under 30s long for 1080p generation and under 60s long for 720p generation."
    },
    "mask_url": {
      "type": "string",
      "description": "The URL of the mask image to apply to the image. Only the person in the white area of the mask will speak."
    },
    "image_url": {
      "type": "string",
      "description": "The URL of the image used to generate the video"
    },
    "resolution": {
      "type": "string",
      "description": "The resolution of the generated video. Defaults to 1080p. 720p generation is faster and higher in quality. 1080p generation is limited to 30s audio and 720p generation is limited to 60s audio.",
      "default": "1080p",
      "enum": [
        "720p",
        "1080p"
      ]
    },
    "turbo_mode": {
      "type": "boolean",
      "description": "Generate a video at a faster rate with a slight quality trade-off.",
      "default": false
    }
  }
}