Skip to main content
Dub transcribes a recording, translates its speech and renders a new spoken track. Input and output languages are independent. Read /healthz/dub for the currently available languages before submitting a job. Use Speak when you already have text to synthesize, or Cast when you want to direct a voiceover from a script. Dub starts with an existing recording and runs asynchronously. When /healthz/dub reports original_voices: true, source voices are matched automatically without substituting an unrelated speaker. Use at least 2 seconds of clear speech for each speaker. Voice similarity and accent can vary; review the result before publishing. Explicit speaker_voices overrides use the voices you selected.

Use Dubbing in the console

Open Dubbing, upload a recording, choose the input and output languages, and create a dub. In-app conversions use your signed-in workspace and credits. No API key setup is required. Track progress and download results from Saved jobs. Choose Edit passages on a completed dub to review its original transcript and translation. Correct text, expand Pronunciation when needed, and Save. Saving text does not change the audio. Select the passages to update and choose Regenerate to create a new version. The previous audio remains available; unselected passages reuse their existing speech. A correction to only the original transcript is translated again during regeneration. An edited translation is used as written. The pronunciation field replaces the spoken text of the entire passage; keep the readable translation in the translation field for subtitles. Load Original or Dubbed audio and use Play passage to compare the corresponding section. Export audio, available video, and SRT or VTT subtitles from the same workspace. Subtitles follow the currently rendered version, so regenerate before exporting corrections.

Pricing and credits

Dubbing costs USD 0.20 per source minute per output language, billed per second. A 10-minute recording dubbed into two languages costs USD 4.00 at the standard rate. Original voice matching, text editing, previews, downloads and subtitles are included. Regenerating a passage charges only for its selected source duration. Failed jobs charge zero. Each successful job is charged once, even if you download it repeatedly. The console shows an estimate before you start and a receipt when the job finishes. New dubs and revisions require available credits. Existing outputs remain downloadable when your balance runs out. Account terms apply; see pricing for current rates.

Use the API

In Console → API Keys, create or edit a product key and select Dub under Dub, speech-to-speech dubbing. The default Voice agent preset does not include Dub. Use a key with dub:render and check its scopes with GET /v1/me. A key without this scope returns 403; a live key may also require funded credit. Keep the key on your server. This example uses an English WAV. Supply source_lang=en and target_lang=hi. You do not need to transcribe or translate the recording yourself.

1. Submit the recording

The response is 202 Accepted, for example:
Keep the returned job_id. status_url is a relative path on https://api.pyai.com; it requires your product key. Send exactly one source: file or source_url. A file upload must be non-empty and no larger than 200 MiB. For a URL, send a publicly reachable media address in the source_url form field instead of file.

2. Poll until processing finishes

Poll every few seconds with backoff when the service is busy. Processing time depends on the recording and current capacity. This is not a streaming API. A failed job returns HTTP 200 with status: "error"; HTTP success alone does not establish that dubbing succeeded. A completed audio job includes fields such as:
source_seconds, dubbed_seconds and progress details may also be present. Treat stage labels and diagnostic fields as extensible; branch on status.

3. Download the audio

The result is a WAV file. Save it in your application and review it before publishing. Downloading before done returns 409 job_not_done. An expired output returns 410 and needs a new render. Do not rely on job output storage as your archive. Dubbing counts source time once per completed generation and output language, not the length of the translated audio. Previews and repeat downloads are included. When returned, billing.status reports pending, charged, or not_charged, and billing.charged_usd gives the confirmed charge. Failed generations are not charged. See the pricing page for current terms; this guide does not promise a fixed completion time.

Edit passages through the API

Completed jobs report editable: true when passage editing is available. All editing and export routes require the same dub:render scope and organization as the original job.
  1. Read GET /v1/dub/jobs/{job_id}/editor for the saved revision and passages.
  2. Save with PATCH /v1/dub/jobs/{job_id}/editor. Send the current revision and an edits array, with each passage’s id and changed source_text, translated_text, or pronunciation_text. A stale revision returns 409 dub_edit_conflict; fetch the latest editor before reconciling changes.
  3. Submit POST /v1/dub/jobs/{job_id}/revisions with the saved revision and segment_ids for 1–50 unique passages. Include an Idempotency-Key of 16–128 letters, digits, underscores, or hyphens. Reuse the same key and unchanged body after a lost response. A different operation needs a new key.
  4. Poll the returned job_id. The new job reports parent_job_id, and the original job remains available.
Fetch GET /v1/dub/jobs/{job_id}/source for original audio, or GET /v1/dub/jobs/{job_id}/subtitles.srt?track=translation for translated SRT. Use .vtt for WebVTT or track=source for the original transcript. Translated subtitle timing follows the rendered audio; source subtitles follow the source recording. Unsynthesized edits and pronunciation spellings are not exported. Editor reads, text saves, original-audio retrieval, subtitles, and audio/video exports are included. A revision counts only the source time covered by the selected passages; overlapping passages count once. The accepted rate is retained through processing and retries. Existing free jobs stay free.

Languages

Accepted input codes do not mean Dub can produce those languages. Speak’s language catalog does not describe Dub’s output-language availability.

Optional fields

All fields use multipart form data. JSON fields are JSON-encoded strings. The audio quickstart above does not require any of these options.

Video outputs

The submission endpoint also accepts video containers. A completed video job can include video_url or outputs.video; fetch that returned path with your product key. Audio-source jobs have no video output and return 404 no_video_output from /video. The source picture is retained with a dubbed audio track; this is not lip synchronization. The quickstart and listenable sample on the Dub product page demonstrate audio dubbing.

Handle errors

See the API reference for request and response schemas, or errors and limits for shared API failures.