> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Create an async transcription job

> Submit audio for batch transcription. Provide **exactly one** source:
either `audio_url` (an https URL we fetch; the input is processed transiently and never written to durable input storage)
a multipart upload (`multipart/form-data` with an `audio` file part and the same fields as form fields), **or** `gpt_live: {session_id}` for a connected project (gated availability).
GPT-Live: an organization owner/admin first configures the project connection through the management API. Uses a finalized stored OpenAI recording, downloaded transiently with a 100 MiB limit. Requires an active secret key with `hear:transcribe` and `transcribe:jobs`. No source audio is retained by this connector; normal result retention applies. OpenAI storage must be enabled and the session created with `store: true`. This is a post-call import, not an account-wide discovery service.
GPT-Live imports deduplicate by PyAI project, connection and session ID even without Idempotency-Key. Repeated submissions return the existing job; different normalized processing options return 409. Credential rotation retains the connection identity. Disconnect/reconnect creates a new identity. Failed jobs remain terminal; there is no automatic new billed job on resubmission.

Returns `202` immediately with a `queued` job; poll `GET /v1/transcription/jobs/{id}` or supply a `webhook_url` for a signed completion callback.
Completed JSON results include word- and segment-level offsets in decimal seconds from the decoded source-media timeline. Silence is not removed or compacted; resampling and internal chunking do not shift later offsets.
Set `channel: true` for stereo (dual-channel) recordings to get exact speaker separation per channel. Use `diarize: true` for model-derived speaker labels on mono audio; do not set both.
Set `vocabulary` only when this job has known names, brands, products, or distinctive terms. PyAI uses up to five sanitized terms. If stored vocabulary is enabled for `batch`, request terms come first and stored suggestions fill remaining slots.

**Input limits:** multipart uploads are limited to 1 GiB; `audio_url` downloads are limited to 512 MiB. There is no separate media-duration ceiling. Inputs must contain a decodable audio stream; the stable output formats are JSON, SRT, and VTT.

**Retention:** URL-fetched input bytes are not persisted. Uploaded input audio is retained for up to 7 days and result artifacts for up to 30 days. This endpoint has no `store: false` mode. `DELETE /v1/transcription/jobs/{id}` cancels queued/running work; it is not an erasure endpoint.
Requires the `transcribe:jobs` scope.



## OpenAPI

````yaml https://api.pyai.com/openapi.json post /v1/transcription/jobs
openapi: 3.1.0
info:
  title: PyAI API
  version: 2.9.0
  description: >-
    Telephony-native Voice AI behind one bearer key:


    - **Hear**, speech-to-text · `POST /v1/audio/transcriptions` (streaming +
    batch)

    - **Speak**, text-to-speech · `POST /v1/audio/speech`, `GET /v1/voices`

    - **Clone**, custom voices from a short clip · `/v1/voice/clones`

    - **Cast**, auto-directed expressive voiceovers · `/v1/cast`

    - **Dub**, asynchronous audio and video dubbing · `POST /v1/dub`
    ([guide](https://docs.pyai.com/guides/dub-overview))

    - **Cue configuration frames (reserved)**; Hear streams skip knowledge-base
    retrieval by default

    - **Omni**, full-duplex agentic voice (speech-to-speech, grounded in your
    knowledge bases + tools) · `/v1/omni`

    - **Knowledge Bases**, hosted grounding for Omni: create bases, add
    documents (file, URL, or text), crawl a public website, bind to agents or
    org defaults · `/v1/knowledgebases`

    - **AMD API**, answering-machine detection: know *who or what* answered a
    call (human, voicemail, IVR, iPhone/Google screening, dead number) with the
    reason it decided · `wss …/v1/amd/stream` (Twilio Media Streams drop-in),
    `POST /v1/amd/config`, `GET /v1/amd/calls/{id}`

    - **Agents Beta**, the live console feature to create, configure, test, and
    connect Omni voice agents without code. Beta features and limits may change.


    ## Authentication


    Create a key in the [console](https://console.pyai.com) (it is shown once)
    and send it as a bearer token:


    ```

    Authorization: Bearer pyai_live_...

    ```


    Keys are environment-scoped: `pyai_live_...` (production) and
    `pyai_test_...` (sandbox). `POST /v1/sandbox/keys` creates an instant,
    short-lived test key without login or billing. Account signup also creates a
    sandbox key. Live keys consume prepaid credit; phone verification may unlock
    promotional credit under graduated-signup rules, but credit is not
    guaranteed at signup.


    Keys are self-validating signed tokens: they work on every PyAI surface the
    instant they are created, no activation or propagation delay. Treat them as
    opaque strings (up to 512 chars) and never parse their contents.


    WebSocket endpoints can't use request headers from a browser, so pass the
    key as a **subprotocol** instead:


    ```

    Sec-WebSocket-Protocol: pyai.v1, pyai-key.pyai_live_...

    ```


    (server-side clients may instead append `?api_key=...` to the URL). Your key
    is authenticated on the upgrade and never reaches the model.


    ## Quickstart, Hear (speech-to-text)


    ```

    curl https://api.pyai.com/v1/audio/transcriptions \
      -H "Authorization: Bearer $PYAI_API_KEY" \
      -F file=@audio.wav -F model=pyai-hear
    # -> { "text": "..." }

    ```


    ## Quickstart, Speak (text-to-speech)


    ```

    curl https://api.pyai.com/v1/audio/speech \
      -H "Authorization: Bearer $PYAI_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model":"pyai-speak","input":"Hello from PyAI.","voice":"voice_abc"}' \
      --output speech.wav
    ```


    `voice` is a stock voice id from `GET /v1/voices` (the curated prebuilt
    catalog with personas and avatars) or a cloned voice id from
    `/v1/voice/clones`. Omit it to use the platform default voice
    (`stock_dorit_en_us`).


    ## Quickstart, Omni (realtime voice agent)


    Omni is **zero-state, there is nothing to create first.** Open a WebSocket,
    pass your key as a subprotocol, and send the agent's behavior (voice,
    persona, knowledge endpoint) in the first `configure` frame:


    ```

    wss://api.pyai.com/v1/omni?session_label=support&format=pcm16&rate=24000
      Sec-WebSocket-Protocol: pyai.v1, pyai-key.$PYAI_API_KEY
    ```


    The session is authorized by your key's **organization**; `session_label` is
    an **optional, opaque** tag (echoed to your own knowledge endpoint for
    correlation), omit it or use any value. When `session_label` equals a
    **`/v1/agents` profile id**, the engine loads persona, voice, and **greeting
    message** from that profile (turn-0 playback). `format` and `rate` are
    load-bearing on the connect URL (the SDK sets them). Prefix each PCM16 frame
    with byte `0x01`; prefix control JSON with byte `0x03`. **Optional
    convenience:** pre-store config via `POST /v1/agents` (including `greeting`,
    `consent_line`, `recordings_enabled`) and pass its id as `session_label`, or
    send everything inline in the post-handshake `configure` frame. Not required
    to connect.


    For reproducible eval runs, determinism controls (`seed`/`temperature`) ride
    the Omni session's `configure` frame, which the gateway passes through
    unchanged, they are honored once the engine supports them; no platform
    change is required.


    ## Scopes


    | Scope | Grants |

    | --- | --- |

    | `hear:transcribe` | `POST /v1/audio/transcriptions` |

    | `hear:stream` | `GET /v1/audio/transcriptions/stream` (WebSocket) |

    | `hear:configure` | `GET`/`PUT /v1/hear/vocabulary` |

    | `speak:synthesize` | `POST /v1/audio/speech` (Speak) |

    | `speak:clone` | `/v1/voice/clones` (Clone) |

    | `speak:design` | `/v1/voice/design` (Speak) |

    | `omni:session` | `/v1/omni` and `POST /v1/omni/sessions` (mint a browser
    session token) |

    | `omni:read` | `/v1/omni/calls` (Omni post-call records) |

    | `kb:manage` | `/v1/knowledgebases/*` (hosted knowledge bases for Omni
    grounding) |

    | `transcribe:jobs` | `/v1/transcription/jobs` |

    | `trace:configure` | `/v1/trace/config`, `/v1/trace/rule-packs` (Trace
    management) |

    | `trace:read` | `/v1/trace/interactions`, `/violations`, `/findings`,
    `/exposure` (Trace reads) |

    | `recap:configure` | `/v1/recap/config` (Recap management) |

    | `recap:configure` | `/v1/recap/crm-config` (Salesforce field mapping) |

    | `recap:read` | `/v1/recap/calls` (Recap reads and speaker-role
    corrections) |

    | `amd:detect` | `wss …/v1/amd/stream` (AMD realtime detection, Twilio
    drop-in) |

    | `amd:configure` | `/v1/amd/config` (AMD operating-point dial + webhook) |

    | `amd:read` | `/v1/amd/calls` (AMD decision records) |

    | `telephony:manage` | `/v1/telephony/*` (managed numbers) |


    `GET /v1/models`, `GET /v1/voices`, and `GET /v1/me` need no specific scope,
    any active key may call them. Wildcards (`hear:*`, `speak:*`, …, and the
    global `*`) grant every scope in their family.


    ## Canonical endpoints


    One row per product surface, endpoint, auth, required scope, and lifecycle
    status. **live** = generally available; **beta** = available now with
    features or limits that may change; **unavailable** = reserved in the
    contract but not active on the serving route.


    | Product | Endpoint | Auth | Scope | Status |

    | --- | --- | --- | --- | --- |

    | Identity | `GET /v1/me` | Bearer | _any active key_ | live |

    | Models | `GET /v1/models` | Bearer | _any active key_ | live |

    | Voices | `GET /v1/voices`, `GET /v1/voices/{id}` | Bearer | _any active
    key_ | live |

    | Hear (batch) | `POST /v1/audio/transcriptions` | Bearer |
    `hear:transcribe` | live |

    | Hear (vocabulary settings) | `GET`/`PUT /v1/hear/vocabulary` | Bearer |
    `hear:configure` | live |

    | Hear (streaming) | `GET /v1/audio/transcriptions/stream` (WS) |
    Subprotocol | `hear:stream` | live |

    | Cue | `GET /v1/audio/transcriptions/stream` + grounding (WS) | Subprotocol
    | `hear:stream` | unavailable |

    | Hear (async batch) | `POST`/`GET /v1/transcription/jobs` | Bearer |
    `transcribe:jobs` | live |

    | Speak (TTS) | `POST /v1/audio/speech` | Bearer | `speak:synthesize` | live
    |

    | Clone | `GET`/`POST /v1/voice/clones` | Bearer | `speak:clone` | live |

    | Speak (design) | `/v1/voice/design` | Bearer | `speak:design` | live |

    | Omni | `wss …/v1/omni?session_label=` | Subprotocol | `omni:session` |
    live |

    | Agent profiles (optional config) | `/v1/agents`, `/v1/agents/{id}` |
    Bearer | `omni:session` | live |

    | Knowledge Bases (hosted grounding) | `/v1/knowledgebases/*`, `PUT
    /v1/agents/{id}/knowledgebases` | Bearer | `kb:manage` (`omni:session` for
    the binding) | live |

    | Trace (config) | `/v1/trace/config`, `/v1/trace/rule-packs` | Bearer |
    `trace:configure` | beta |

    | Trace (reads) | `/v1/trace/interactions`, `/violations`, `/findings`,
    `/exposure` | Bearer | `trace:read` | beta |

    | Recap (config) | `/v1/recap/config` | Bearer | `recap:configure` | live |

    | Recap (CRM) | `/v1/recap/crm-config` | Bearer | `recap:configure` | live |

    | Integrations (Zapier) | `/v1/integrations/events`,
    `/v1/integrations/zapier/hooks` | Bearer | _any active key_ | live |

    | Recap (reads) | `/v1/recap/calls` | Bearer | `recap:read` | live |

    | Omni call records | `/v1/omni/calls`, `/v1/omni/calls/{id}` | Bearer |
    `omni:read` | live |

    | AMD (stream) | `wss …/v1/amd/stream` (Twilio Media Streams drop-in) |
    TwiML `<Parameter name="api_key">` (from Twilio) or subprotocol
    (server-side) | `amd:detect` | live |

    | AMD (config) | `GET`/`POST /v1/amd/config` | Bearer | `amd:configure` |
    live |

    | AMD (reads) | `GET /v1/amd/calls`, `/v1/amd/calls/{id}` | Bearer |
    `amd:read` | live |

    | Telephony | `/v1/telephony/*` | Bearer | `telephony:manage` | live |

    | Agents (console builder) | `https://console.pyai.com/agents` | Console
    session |, | beta |


    WebSocket surfaces authenticate with the `Sec-WebSocket-Protocol: pyai.v1,
    pyai-key.<API_KEY>` subprotocol pair (or `?api_key=` server-side);
    everything else takes the `Authorization: Bearer` key. Managed-number calls
    return 404 until the PyAI network is enabled for the account.


    ## Rate limits & billing


    Every key has a per-second rate limit (with burst) and a cap on concurrent
    realtime sessions. Exceeding either returns `429` with a `Retry-After`
    header. Usage is metered per minute of audio, transcription minutes (Hear),
    synthesized audio minutes (Speak), and realtime session minutes (Omni), and
    billed against your plan and credits. List prices: Hear $0.001/min (async
    Transcribe $0.0005/min), Speak $0.04/min, Omni $0.05/min including speech
    plus brain, and Agents Live Beta $0.08/min. Managed telephony is separate at
    $0.01/min. English Natural (`en1`) is available on Speak (streaming and
    buffered) and Omni; Omni acknowledges the canonical id and `voice_tier:
    natural`. Hindi uses the Standard-tier voices `hi1`–`hi4` on Omni; the
    former Hindi Natural aliases (`hi5`–`hi8`) are retired and no longer in the
    catalog. Standard and Natural voices are included in their product's base
    rate with no voice-tier add-on. The AMD API bills per **answered** call, the
    first 5,000 answered calls each month are free, then $0.004/answered call
    (no-answers, busies, and failed calls are free; AMD bundled with PyAI
    telephony/Omni is included at no charge). AI products (Hear, Speak, Omni)
    bill **per second by default**, the pulse is applied once to each meter's
    invoice-period total, so many short sessions are summed and rounded a single
    time (never minute-rounded per call), and an empty/failed call bills
    nothing. Coarser pulses are available as an optional enterprise override.
    Managed telephony minutes keep a 1-minute pulse. Per-character Speak billing
    is available on enterprise contracts.
  contact:
    name: PyAI
    url: https://pyai.com
servers:
  - url: https://api.pyai.com
    description: Production
security:
  - apiKey: []
  - xApiKey: []
tags:
  - name: Dub
    description: >-
      Asynchronous dubbing: submit a recording, choose the languages, poll the
      job and download the output.
  - name: Omni
    description: >-
      The flagship: build an AI voice agent with one WebSocket (`GET /v1/omni`)
      and one `configure` frame, nothing to pre-create. This group also holds
      the optional browser-token mint and the post-call records.
  - name: Knowledge Bases
    description: >-
      Hosted knowledge bases for Omni grounding: create a base, add documents
      (file upload, URL fetch, or pasted text), then bind it to agent profiles
      or set org-wide defaults. Bound bases are retrieved per turn, no
      `kb_endpoint` of your own required.
  - name: Identity
    description: >-
      Introspect the calling key: org/project, env, granted scopes, and
      limits/credit posture. Use it to self-diagnose a 401/403/402.
  - name: Speech To Text (Hear)
    description: Speech-to-text (streaming + batch)
  - name: Text To Speech (Speak)
    description: Text-to-speech, stock voices, and prompt-to-voice design
  - name: Clone
    description: Enroll, list, and delete custom voices from a short reference clip
  - name: Models
    description: Model catalog
  - name: Sandbox
    description: >-
      Zero-friction onboarding for coding agents: mint a free, instant, no-card
      sandbox key with no human steps.
  - name: Startup Program
    description: >-
      PyAI for Startups: $20k to $100k in PyAI credit for early-stage voice
      teams. Public application endpoint; review and activation happen out of
      band.
  - name: Cast
    description: >-
      Auto-directed, expressive multi-line voiceover projects and asynchronous
      renders.
  - name: Transcription Jobs
    description: Async batch transcription
  - name: Agents
    description: >-
      Agent profiles used by the live Agents Beta console and available directly
      through the API. Store Omni session config (persona, greeting, voice,
      conversation knobs) and reference it by id instead of sending a full
      `configure` frame each call. Profiles remain optional for direct
      `/v1/omni` integrations.
  - name: Trace
    description: >-
      Compliance & guardrails: per-agent config, rule packs, and the exposure /
      violations / interaction-evidence read views
  - name: AMD
    description: >-
      Answering-machine detection: know who or what answered a call (human,
      voicemail, IVR, iPhone/Google screening, dead number), with the reason it
      decided. Twilio Media Streams drop-in over `wss …/v1/amd/stream`; one
      operating-point dial; billed per answered call.
  - name: Telephony
    description: >-
      Managed phone numbers: search, provision, route to an agent, and release.
      Call minutes bill on telephony.minutes ($0.01/min).
  - name: WhatsApp
    description: >-
      WhatsApp Business Calling: register a WhatsApp Business number, enable
      calling, and let an Omni agent answer (and, with the user's permission,
      place) WhatsApp voice calls. Requires the `telephony:manage` scope.
  - name: Call Integrations
    description: >-
      Signed provider webhooks that import completed calls into Hear, Recap, and
      offline Trace.
paths:
  /v1/transcription/jobs:
    post:
      tags:
        - Transcription Jobs
      summary: Create an async transcription job
      description: >-
        Submit audio for batch transcription. Provide **exactly one** source:

        either `audio_url` (an https URL we fetch; the input is processed
        transiently and never written to durable input storage)

        a multipart upload (`multipart/form-data` with an `audio` file part and
        the same fields as form fields), **or** `gpt_live: {session_id}` for a
        connected project (gated availability).

        GPT-Live: an organization owner/admin first configures the project
        connection through the management API. Uses a finalized stored OpenAI
        recording, downloaded transiently with a 100 MiB limit. Requires an
        active secret key with `hear:transcribe` and `transcribe:jobs`. No
        source audio is retained by this connector; normal result retention
        applies. OpenAI storage must be enabled and the session created with
        `store: true`. This is a post-call import, not an account-wide discovery
        service.

        GPT-Live imports deduplicate by PyAI project, connection and session ID
        even without Idempotency-Key. Repeated submissions return the existing
        job; different normalized processing options return 409. Credential
        rotation retains the connection identity. Disconnect/reconnect creates a
        new identity. Failed jobs remain terminal; there is no automatic new
        billed job on resubmission.


        Returns `202` immediately with a `queued` job; poll `GET
        /v1/transcription/jobs/{id}` or supply a `webhook_url` for a signed
        completion callback.

        Completed JSON results include word- and segment-level offsets in
        decimal seconds from the decoded source-media timeline. Silence is not
        removed or compacted; resampling and internal chunking do not shift
        later offsets.

        Set `channel: true` for stereo (dual-channel) recordings to get exact
        speaker separation per channel. Use `diarize: true` for model-derived
        speaker labels on mono audio; do not set both.

        Set `vocabulary` only when this job has known names, brands, products,
        or distinctive terms. PyAI uses up to five sanitized terms. If stored
        vocabulary is enabled for `batch`, request terms come first and stored
        suggestions fill remaining slots.


        **Input limits:** multipart uploads are limited to 1 GiB; `audio_url`
        downloads are limited to 512 MiB. There is no separate media-duration
        ceiling. Inputs must contain a decodable audio stream; the stable output
        formats are JSON, SRT, and VTT.


        **Retention:** URL-fetched input bytes are not persisted. Uploaded input
        audio is retained for up to 7 days and result artifacts for up to 30
        days. This endpoint has no `store: false` mode. `DELETE
        /v1/transcription/jobs/{id}` cancels queued/running work; it is not an
        erasure endpoint.

        Requires the `transcribe:jobs` scope.
      operationId: createTranscriptionJob
      parameters:
        - name: Idempotency-Key
          in: header
          required: false
          schema:
            type: string
            maxLength: 255
          description: >-
            Opt-in safe retry (JSON body path). Reusing the key with an
            identical body replays the original 202 response; reusing it with a
            different body returns 409. GPT-Live sources use connection/session
            deduplication instead and ignore this header.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              oneOf:
                - required:
                    - audio_url
                  not:
                    required:
                      - gpt_live
                - required:
                    - gpt_live
                  not:
                    required:
                      - audio_url
              properties:
                gpt_live:
                  type: object
                  additionalProperties: false
                  required:
                    - session_id
                  properties:
                    session_id:
                      type: string
                      minLength: 1
                      maxLength: 256
                      description: >-
                        Opaque completed GPT-Live session ID belonging to the
                        connected OpenAI project.
                audio_url:
                  type: string
                  format: uri
                  description: >-
                    HTTPS URL of the audio to transcribe. PyAI fetches the input
                    transiently without writing it to durable input storage.
                    Maximum response body: 512 MiB.
                model:
                  type: string
                  default: pyai-hear-telephony
                channel:
                  type: boolean
                  default: false
                  description: >-
                    Dual-channel (stereo) separation. Channel 0 is labelled
                    `speaker_1`, channel 1 `speaker_2`; labels are neutral and
                    do not infer agent/customer roles. Do not combine with
                    `diarize`.
                diarize:
                  type: boolean
                  default: false
                  description: >-
                    Model-derived speaker separation for mono audio. Labels
                    identify turns within this result, not stable people across
                    separate jobs. Use `channel` instead for stereo recordings.
                numerals:
                  type: boolean
                  description: >-
                    Tri-state inverse-text normalization for English **final**
                    transcripts (never interim partials). `true` renders spoken
                    numbers as digits (phones, currency, dates, ordinals).
                    `false` keeps those spans in spoken form. Omitted keeps the
                    live engine default: number formatting is ON for finals.
                    Independent of `smart_format`.
                smart_format:
                  type: boolean
                  default: false
                  description: >-
                    Opt-in English punctuation and sentence capitalization on
                    **final** transcripts only. Interim partials are never
                    formatted. May change only case and punctuation; any failure
                    returns the unformatted final. Default `false`. Independent
                    of `numerals`. Non-English requests are unchanged.
                dictation:
                  type: boolean
                  default: false
                  description: >-
                    Opt-in spoken punctuation commands on English **final**
                    transcripts only: `period`, `comma`, `new paragraph`, and
                    `question mark`. Separate from `smart_format` and off by
                    default. Interim partials are never rewritten.
                drop_fillers:
                  type: boolean
                  default: false
                  description: >-
                    Opt-in stripping of filled pauses (`um`, `uh`, `umm`, `uhh`,
                    `er`) on English **final** transcripts. Off by default. Do
                    not enable on legal or compliance audio by default. Interim
                    partials are never rewritten.
                vocabulary:
                  type: array
                  items:
                    type: string
                    minLength: 4
                    maxLength: 64
                  maxItems: 5
                  description: >-
                    Optional per-job terms for known names, brands, products,
                    and other distinctive phrases. PyAI trims entries, removes
                    case-insensitive duplicates, and keeps the first spelling
                    and order. Entries shorter than 4 characters, longer than 64
                    characters, sentence-shaped input, non-string entries, and
                    conservative common words are ignored. At most the first 5
                    valid terms are used. Invalid entries do not reject the job.
                    When stored vocabulary is enabled for `batch`, request terms
                    come first and stored suggestions fill any remaining slots.
                    The effective list applies only to this job and does not
                    select transcription language.
                output_formats:
                  type: array
                  items:
                    type: string
                    enum:
                      - json
                      - srt
                      - vtt
                  default:
                    - json
                webhook_url:
                  type: string
                  format: uri
                  description: >-
                    HTTPS URL for `transcription.job.completed` or
                    `transcription.job.failed`. PyAI POSTs `{type, created,
                    data}` and signs the exact body in `X-PyAI-Signature:
                    t=<unix_seconds>,v1=<hex>`, where `v1` is HMAC-SHA256 over
                    `<t>.<rawBody>`.
                trace:
                  type: boolean
                  default: false
                  description: >-
                    Trace compliance add-on: deterministic PII scan + redaction
                    over the final transcript (SSN, card numbers,
                    CVV-in-context, email, US phone — the pii_v0 entity set).
                    The result carries the redacted transcript plus a `trace`
                    summary (verdict, PII count). Patterns run on the formatted
                    transcript, so pair with the default `numerals` (digits) for
                    full effect. Requires the org's Trace entitlement on this
                    surface (else `402`); bills one Trace call. Not supported
                    together with `diarize`/`channel`.
                rule_pack:
                  type: object
                  additionalProperties: true
                  description: >-
                    Optional Trace rule pack (only used when `trace` is true).
                    `entities` (or `redact`) narrows the scan to a subset of:
                    ssn, credit_card, cvv, email, us_phone; unknown names are
                    ignored and an empty selection means the full set.
                call_id:
                  type: string
                  description: >-
                    Optional stable call identifier for Recap; defaults to the
                    transcription job id.
                pack_id:
                  type: string
                  pattern: ^[a-z0-9_]+$
                  description: Optional Recap pack id.
                call_direction:
                  type: string
                  enum:
                    - inbound
                    - outbound
                  description: Optional Recap call direction.
                customer_name:
                  type: string
                  description: Optional Recap customer label.
                language:
                  type: string
                  enum:
                    - en
                    - fr
                    - es
                    - de
                    - hi
                  description: >-
                    Optional Recap summarization language. It does not affect
                    transcription: async jobs transcribe all eight Hear
                    languages (`en`/`es`/`fr`/`de`/`hi`/`it`/`pt`/`nl`) and the
                    spoken language is auto-detected per call.
                crm_fields:
                  type: object
                  additionalProperties: true
                  description: >-
                    Optional CRM metadata delivered durably with the Recap
                    trigger.
          multipart/form-data:
            schema:
              type: object
              required:
                - audio
              properties:
                audio:
                  type: string
                  format: binary
                  description: >-
                    Audio file, maximum 1 GiB. Must contain a decodable audio
                    stream.
                model:
                  type: string
                channel:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                    - stereo
                diarize:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                numerals:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                  description: >-
                    Tri-state inverse-text normalization for English **final**
                    transcripts (never interim partials). `true` renders spoken
                    numbers as digits (phones, currency, dates, ordinals).
                    `false` keeps those spans in spoken form. Omitted keeps the
                    live engine default: number formatting is ON for finals.
                    Independent of `smart_format`.
                smart_format:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                  description: >-
                    Opt-in English punctuation and sentence capitalization on
                    **final** transcripts only. Interim partials are never
                    formatted. May change only case and punctuation; any failure
                    returns the unformatted final. Default `false`. Independent
                    of `numerals`. Non-English requests are unchanged.
                dictation:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                  description: >-
                    Opt-in spoken punctuation commands on English **final**
                    transcripts only: `period`, `comma`, `new paragraph`, and
                    `question mark`. Separate from `smart_format` and off by
                    default. Interim partials are never rewritten.
                drop_fillers:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                  description: >-
                    Opt-in stripping of filled pauses (`um`, `uh`, `umm`, `uhh`,
                    `er`) on English **final** transcripts. Off by default. Do
                    not enable on legal or compliance audio by default. Interim
                    partials are never rewritten.
                vocabulary:
                  type: string
                  maxLength: 1024
                  description: >-
                    Optional per-job terms for known names, brands, products,
                    and other distinctive phrases. PyAI trims entries, removes
                    case-insensitive duplicates, and keeps the first spelling
                    and order. Entries shorter than 4 characters, longer than 64
                    characters, sentence-shaped input, non-string entries, and
                    conservative common words are ignored. At most the first 5
                    valid terms are used. Invalid entries do not reject the job.
                    When stored vocabulary is enabled for `batch`, request terms
                    come first and stored suggestions fill any remaining slots.
                    The effective list applies only to this job and does not
                    select transcription language. Send a comma-separated list
                    or a JSON array string.
                output_formats:
                  type: string
                  description: Comma-separated, e.g. `json,srt`.
                webhook_url:
                  type: string
                  format: uri
                  description: >-
                    HTTPS completion/failure webhook. Signed as documented on
                    the JSON request.
                trace:
                  type: string
                  enum:
                    - 'true'
                    - 'false'
                  description: >-
                    Trace compliance add-on (see JSON body). Requires the Trace
                    entitlement; not supported with diarize/channel.
                rule_pack:
                  type: string
                  description: JSON-encoded Trace rule pack (only used when trace=true).
                call_id:
                  type: string
                pack_id:
                  type: string
                  pattern: ^[a-z0-9_]+$
                call_direction:
                  type: string
                  enum:
                    - inbound
                    - outbound
                customer_name:
                  type: string
                language:
                  type: string
                  enum:
                    - en
                    - fr
                    - es
                    - de
                    - hi
                crm_fields:
                  type: string
                  description: JSON-encoded CRM metadata delivered with the Recap trigger.
      responses:
        '202':
          description: Job accepted
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/TranscriptionJob'
        '400':
          $ref: '#/components/responses/BadRequest'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '402':
          $ref: '#/components/responses/PaymentRequired'
        '403':
          $ref: '#/components/responses/Forbidden'
        '409':
          description: >-
            Idempotency-Key reused with a different body (`code:
            idempotency_conflict`)
          content:
            application/problem+json:
              schema:
                $ref: '#/components/schemas/Problem'
        '429':
          $ref: '#/components/responses/RateLimited'
components:
  schemas:
    TranscriptionJob:
      type: object
      required:
        - job_id
        - status
        - created_at
        - updated_at
      properties:
        job_id:
          type: string
          example: job_aZ09...
        status:
          type: string
          enum:
            - queued
            - running
            - completed
            - failed
            - cancelled
        created_at:
          type: integer
          description: Unix ms.
        updated_at:
          type: integer
          description: Unix ms.
        result:
          $ref: '#/components/schemas/TranscriptionJobResult'
          description: >-
            Present on completed jobs (inline). Large results are offloaded to
            result_url instead.
        result_url:
          type: string
          format: uri
          description: Signed GET URL for an offloaded large result.
        error:
          type: string
          description: >-
            Human-readable normalized failure message on failed jobs. This field
            is not a stable machine-readable failure code.
    Problem:
      type: object
      description: >-
        RFC 7807 problem+json returned by the control plane (request-validation
        and resource errors such as 400/404/409). The stable code is the last
        path segment of `type`.
      required:
        - title
        - status
      properties:
        type:
          type: string
          description: Problem type URI; ends with the stable code.
        title:
          type: string
        status:
          type: integer
        detail:
          type: string
        request_id:
          type: string
    TranscriptionJobResult:
      type: object
      required:
        - text
        - speakers
        - audio_seconds
        - segments
        - words
      description: >-
        Completed async Hear result. The spoken language is auto-detected per
        call across the eight supported Hear languages and the transcript is
        returned in the detected language. This object does not currently return
        the detected language code.
      properties:
        text:
          type: string
          description: >-
            Complete transcript. Diarized results prefix segments with neutral
            speaker labels.
        speakers:
          type: integer
          minimum: 1
          description: Number of distinct speaker labels in this result.
        audio_seconds:
          type: number
          minimum: 0
          description: Decoded source duration in decimal seconds, including silence.
          example: 120
        segments:
          type: array
          items:
            $ref: '#/components/schemas/TranscriptionSegment'
          description: >-
            Timestamped subtitle-ready spans. Segment boundaries may change when
            speaker assignment or the one-second readable-cue gap changes.
        words:
          type: array
          items:
            $ref: '#/components/schemas/TranscriptionWord'
          description: >-
            Word-level timestamps when alignment is available. `smart_format`
            may attach case or punctuation without moving timestamps. Options
            that add or remove tokens (`dictation` and `drop_fillers`) clear
            this array when the rewritten words cannot be safely realigned.
        formats:
          type: object
          additionalProperties:
            type: string
            format: uri
          description: >-
            Requested SRT/VTT format names mapped to short-lived signed GET
            URLs. JSON is the result object itself.
        trace:
          type: object
          description: Present when the Trace add-on was requested.
          properties:
            verdict:
              type: string
            n_pii:
              type: integer
              minimum: 0
            redacted:
              type: boolean
      example:
        text: '[speaker_1] Hello everyone.'
        audio_seconds: 120
        speakers: 1
        words:
          - word: Hello
            start: 0.42
            end: 0.81
            confidence: 0.98
            speaker: speaker_1
            channel: 0
        segments:
          - id: 0
            text: Hello everyone.
            start: 0.42
            end: 1.8
            speaker: speaker_1
            channel: 0
    Error:
      type: object
      description: >-
        OpenAI-compatible error envelope returned by the gateway data plane
        (401/402/403/429). Control-plane request/resource errors use Problem
        (application/problem+json) instead.
      required:
        - error
      properties:
        error:
          type: object
          required:
            - message
          properties:
            message:
              type: string
              description: Human-readable explanation.
            type:
              type: string
              description: Error category, e.g. rate_limit_error.
            code:
              $ref: '#/components/schemas/ErrorCode'
            param:
              type: string
              nullable: true
              description: Offending parameter when applicable, else null.
    TranscriptionSegment:
      type: object
      required:
        - id
        - start
        - end
        - text
      properties:
        id:
          type: integer
          minimum: 0
          description: Zero-based segment index.
          example: 0
        start:
          type: number
          minimum: 0
          description: >-
            Offset in decimal seconds from the start of the decoded source-media
            timeline. Silence is not removed or compacted; resampling and
            internal chunking do not shift later offsets.
          example: 0.42
        end:
          type: number
          minimum: 0
          description: >-
            Offset in decimal seconds from the start of the decoded source-media
            timeline. Silence is not removed or compacted; resampling and
            internal chunking do not shift later offsets.
          example: 1.8
        text:
          type: string
          example: Hello everyone.
        speaker:
          type: string
          description: >-
            Neutral speaker label. With `channel: true`, `speaker_1` maps to
            channel 0 and `speaker_2` to channel 1; this is exact channel
            separation, not an inferred agent/customer role. With mono `diarize:
            true`, labels are model-derived and must not be treated as stable
            identities across separate jobs.
          example: speaker_1
        channel:
          type: integer
          minimum: 0
          description: 'Zero-based source channel. Present on `channel: true` results.'
          example: 0
    TranscriptionWord:
      type: object
      required:
        - word
        - start
        - end
      description: >-
        Word-level timestamps when alignment is available. `smart_format` may
        attach case or punctuation without moving timestamps. Options that add
        or remove tokens (`dictation` and `drop_fillers`) clear this array when
        the rewritten words cannot be safely realigned.
      properties:
        word:
          type: string
          description: >-
            Recognized token. Punctuation may be attached when final-transcript
            formatting is enabled.
          example: Hello
        start:
          type: number
          minimum: 0
          description: >-
            Offset in decimal seconds from the start of the decoded source-media
            timeline. Silence is not removed or compacted; resampling and
            internal chunking do not shift later offsets.
          example: 0.42
        end:
          type: number
          minimum: 0
          description: >-
            Offset in decimal seconds from the start of the decoded source-media
            timeline. Silence is not removed or compacted; resampling and
            internal chunking do not shift later offsets.
          example: 0.81
        confidence:
          type: number
          minimum: 0
          maximum: 1
          description: >-
            Recognition confidence when supplied by the active Hear model.
            Omitted when unavailable.
          example: 0.98
        speaker:
          type: string
          description: >-
            Neutral speaker label. With `channel: true`, `speaker_1` maps to
            channel 0 and `speaker_2` to channel 1; this is exact channel
            separation, not an inferred agent/customer role. With mono `diarize:
            true`, labels are model-derived and must not be treated as stable
            identities across separate jobs.
          example: speaker_1
        channel:
          type: integer
          minimum: 0
          description: 'Zero-based source channel. Present on `channel: true` results.'
          example: 0
        entity:
          type: string
          description: >-
            Normalized entity category when supplied, such as `phone_number`,
            `account_id`, or `date`.
    ErrorCode:
      type: string
      description: >-
        Stable, machine-readable error code. Branch on this rather than the
        human `message`.
      enum:
        - invalid_request_error
        - invalid_session_label
        - unauthorized
        - forbidden
        - origin_not_allowed
        - credit_exhausted
        - key_budget_exceeded
        - insufficient_quota
        - rate_limit_exceeded
        - concurrency_limit_exceeded
        - daily_cap_exceeded
  responses:
    BadRequest:
      description: Invalid request (bad field, unsupported value)
      content:
        application/problem+json:
          schema:
            $ref: '#/components/schemas/Problem'
    Unauthorized:
      description: 'Missing or invalid API key (`code: unauthorized`)'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    PaymentRequired:
      description: >-
        Billing gate: org out of prepaid credit, per-key budget hit, or plan
        quota exhausted (`code: credit_exhausted | key_budget_exceeded |
        insufficient_quota`). Do not retry; add credit or raise the limit. A
        brand-new key may see this on its first call until the account is
        funded.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    Forbidden:
      description: >-
        Key lacks the required scope or the origin is not allow-listed (`code:
        forbidden | origin_not_allowed`)
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    RateLimited:
      description: >-
        Too many requests or too many concurrent realtime sessions; see
        Retry-After header (`code: rate_limit_exceeded |
        concurrency_limit_exceeded | daily_cap_exceeded`)
      headers:
        Retry-After:
          schema:
            type: integer
          description: Seconds to wait.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    apiKey:
      type: http
      scheme: bearer
      description: 'Use `Authorization: Bearer pyai_live_...` (or `pyai_test_...`).'
    xApiKey:
      type: apiKey
      in: header
      name: x-api-key
      description: >-
        Header alias for bearer auth on HTTP endpoints. WebSocket auth uses the
        subprotocol pair `pyai.v1, pyai-key.<API_KEY>`.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.