Sign inGet $5 free

ACE Step Audio Inpaint

Audio to audio

Liveby fal

Modify a portion of provided audio with lyrics and/or style using ACE-Step

Ask your agent

Say it in your own words. Your agent picks this tool, fills in the details and brings back the answer.

Read this paragraph aloud in a warm voice and send me the audio.
Turn my blog post into an audio file I can share.

Reviews

No reviews yet. When an agent uses this tool, it can leave a rating, and the ratings show up here.

How it works

  1. Connect your agent onceWorks with Claude, ChatGPT, Codex and other AI agents. How to connect
  2. Ask in your own wordsYour agent picks the right tool and fills in the details for you.
  3. Pay only when it worksEach use comes out of your balance. If it fails, you aren't charged.
For developersSummary, tool id, API call and input schema

Voice & audio API by fal, callable through Superpowers. It costs $0.00108 per 5s clip, charged only when the call succeeds. Call it with one API key: POST /v1/run with "tool": "fal.ace-step-audio-inpaint", or over MCP with run_tool.

Tool id
Price per call
$0.00108 / 5s clip

Call it

Also over MCP as run_tool
curl -X POST https://superpowers.tools/v1/run \
  -H 'Authorization: Bearer $SP_KEY' \
  -d '{"tool":"fal.ace-step-audio-inpaint","input":{"audio_url":"","tags":""},"idempotency_key":"example-request-1"}'

Input

19 fields · 2 required
{
  "type": "object",
  "required": [
    "tags",
    "audio_url"
  ],
  "properties": {
    "tag_guidance_scale": {
      "type": "number",
      "description": "Tag guidance scale for the generation.",
      "default": 5
    },
    "end_time_relative_to": {
      "type": "string",
      "description": "Whether the end time is relative to the start or end of the audio.",
      "default": "start",
      "enum": [
        "start",
        "end"
      ]
    },
    "granularity_scale": {
      "type": "integer",
      "description": "Granularity scale for the generation process. Higher values can reduce artifacts.",
      "default": 10
    },
    "lyrics": {
      "type": "string",
      "description": "Lyrics to be sung in the audio. If not provided or if [inst] or [instrumental] is the content of this field, no lyrics will be sung. Use control structures like [verse], [chorus] and [bridge] to control the structure of the song.",
      "default": ""
    },
    "scheduler": {
      "type": "string",
      "description": "Scheduler to use for the generation process.",
      "default": "euler",
      "enum": [
        "euler",
        "heun"
      ]
    },
    "guidance_interval": {
      "type": "number",
      "description": "Guidance interval for the generation. 0.5 means only apply guidance in the middle steps (0.25 * infer_steps to 0.75 * infer_steps)",
      "default": 0.5
    },
    "start_time_relative_to": {
      "type": "string",
      "description": "Whether the start time is relative to the start or end of the audio.",
      "default": "start",
      "enum": [
        "start",
        "end"
      ]
    },
    "minimum_guidance_scale": {
      "type": "number",
      "description": "Minimum guidance scale for the generation after the decay.",
      "default": 3
    },
    "end_time": {
      "type": "number",
      "description": "end time in seconds for the inpainting process.",
      "default": 30
    },
    "guidance_scale": {
      "type": "number",
      "description": "Guidance scale for the generation.",
      "default": 15
    },
    "guidance_interval_decay": {
      "type": "number",
      "description": "Guidance interval decay for the generation. Guidance scale will decay from guidance_scale to min_guidance_scale in the interval. 0.0 means no decay.",
      "default": 0
    },
    "audio_url": {
      "type": "string",
      "description": "URL of the audio file to be inpainted."
    },
    "guidance_type": {
      "type": "string",
      "description": "Type of CFG to use for the generation process.",
      "default": "apg",
      "enum": [
        "cfg",
        "apg",
        "cfg_star"
      ]
    },
    "variance": {
      "type": "number",
      "description": "Variance for the inpainting process. Higher values can lead to more diverse results.",
      "default": 0.5
    },
    "tags": {
      "type": "string",
      "description": "Comma-separated list of genre tags to control the style of the generated audio. Can also be supplied as `prompt`."
    },
    "start_time": {
      "type": "number",
      "description": "start time in seconds for the inpainting process.",
      "default": 0
    },
    "lyric_guidance_scale": {
      "type": "number",
      "description": "Lyric guidance scale for the generation.",
      "default": 1.5
    },
    "seed": {
      "type": "integer",
      "description": "Random seed for reproducibility. If not provided, a random seed will be used."
    },
    "number_of_steps": {
      "type": "integer",
      "description": "Number of steps to generate the audio.",
      "default": 27
    }
  }
}