Sign inGet $5 free

Cosmos 3 Super Image to Video

Turns a photo into a video

Liveby NVIDIA

Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs.

Ask your agent

Say it in your own words. Your agent picks this tool, fills in the details and brings back the answer.

Turn this product photo into a 5-second video for Instagram with Cosmos 3 Super Image to Video.
Bring this old family photo to life as a short video.

Reviews

No reviews yet. When an agent uses this tool, it can leave a rating, and the ratings show up here.

How it works

  1. Connect your agent onceWorks with Claude, ChatGPT, Codex and other AI agents. How to connect
  2. Ask in your own wordsYour agent picks the right tool and fills in the details for you.
  3. Pay only when it worksEach use comes out of your balance. If it fails, you aren't charged.
For developersSummary, tool id, API call and input schema

Video API by NVIDIA, callable through Superpowers. It costs $0.270 per 5s clip, charged only when the call succeeds. Call it with one API key: POST /v1/run with "tool": "fal.nvidia-cosmos-3-super-image-to-video", or over MCP with run_tool.

Tool id
Price per call
$0.270 / 5s clip

Call it

Also over MCP as run_tool
curl -X POST https://superpowers.tools/v1/run \
  -H 'Authorization: Bearer $SP_KEY' \
  -d '{"tool":"fal.nvidia-cosmos-3-super-image-to-video","input":{"image_url":"","prompt":""},"idempotency_key":"example-request-1"}'

Input

16 fields · 2 required
{
  "type": "object",
  "required": [
    "prompt",
    "image_url"
  ],
  "properties": {
    "num_inference_steps": {
      "type": "integer",
      "description": "Number of denoising steps. More steps yield higher quality but take longer.",
      "default": 28
    },
    "agentic_samples_per_iteration": {
      "type": "integer",
      "description": "Candidate videos to generate and judge per agentic iteration. The best candidate advances to the next rewrite stage.",
      "default": 2
    },
    "image_size": {
      "type": "string",
      "description": "The size of the generated video. The request is clamped and snapped to the nearest supported NVIDIA tier (256p/480p/720p) and aspect ratio.",
      "default": {
        "height": 480,
        "width": 832
      },
      "enum": [
        "square_hd",
        "square",
        "portrait_4_3",
        "portrait_16_9",
        "landscape_4_3",
        "landscape_16_9"
      ]
    },
    "enable_safety_checker": {
      "type": "boolean",
      "description": "Enable content moderation for the input prompt and image. Disabling it requires account authorization; unauthorized requests are always checked.",
      "default": true
    },
    "guidance_scale": {
      "type": "number",
      "description": "Classifier-free guidance scale. Higher values increase prompt adherence at the cost of diversity.",
      "default": 6
    },
    "enable_prompt_expansion": {
      "type": "boolean",
      "description": "If true, the Cosmos3-Nano Reasoner (a VLM that sees the first frame) rewrites the prompt into the dense caption Cosmos3 was trained on. The app starts a local Reasoner by default, or uses COSMOS_PROMPT_UPSAMPLER_BASE_URL when configured. Falls back to the raw prompt if expansion fails.",
      "default": true
    },
    "agentic_max_iterations": {
      "type": "integer",
      "description": "Maximum agentic prompt stages when agentic generation is enabled.",
      "default": 2
    },
    "seed": {
      "type": "integer",
      "description": "The same seed and prompt given to the same model version will produce the same video every time."
    },
    "num_frames": {
      "type": "integer",
      "description": "Number of frames to generate. More frames yield a longer video.",
      "default": 189
    },
    "enable_agentic_generation": {
      "type": "boolean",
      "description": "Enable the iterative Cosmos agentic loop: prompt upsampling, candidate video generation, VLM critique of sampled frames, and prompt rewrite. Each candidate is a full render, so this is substantially slower and costlier than a single generation.",
      "default": false
    },
    "image_url": {
      "type": "string",
      "description": "URL of the conditioning first-frame image for the video."
    },
    "agentic_early_stop": {
      "type": "boolean",
      "description": "Stop the agentic loop early when the critic score clears the strict quality threshold.",
      "default": true
    },
    "frames_per_second": {
      "type": "integer",
      "description": "Frames per second of the output video.",
      "default": 24
    },
    "prompt": {
      "type": "string",
      "description": "Text prompt describing the motion and scene of the video to generate."
    },
    "sync_mode": {
      "type": "boolean",
      "description": "If `True`, the video is returned as a data URI and the output data won't be available in the request history.",
      "default": false
    },
    "negative_prompt": {
      "type": "string",
      "description": "Content to steer the generation away from (artifacts, unwanted motion). Defaults to NVIDIA's recommended i2v negative prompt; pass an empty string to disable.",
      "default": "The video captures a series of frames showing macroblocking artifacts, chromatic aberration, high-frequency noise, and rolling shutter distortion. It includes static with no motion, motion blur, over-saturation, shaky footage, low resolution, grainy texture, pixelated images, poorly lit areas, underexposed and overexposed scenes, poor color balance, washed out colors, choppy sequences, jerky movements, low frame rate, bit-depth compression artifacts, color banding, unnatural transitions, outdated special effects, fake elements, unconvincing visuals, poorly edited content, jump cuts, hard cut, visual noise, and flickering. It features moiré patterns, edge halos, and temporal aliasing. Furthermore, the content defies common sense, generating illogical scenarios, nonsensical entities, absurd character behaviors, and conceptual paradoxes that violate basic human reasoning and everyday reality. The video looks like a surreal or glitchy hallucination. Overall, the video is of poor quality."
    }
  }
}