> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Answering-machine detection (WebSocket)

> Realtime answering-machine detection over a WebSocket. **This surface speaks Twilio's Media Streams protocol natively** (`start` / `media` / `stop` frames, G.711 μ-law 8 kHz base64, ~20 ms), so migrating from Twilio AMD is a one-line-TwiML change, point the call's media at PyAI, keep your carrier and your code.

```xml
<Response><Start>
  <Stream url="wss://api.pyai.com/v1/amd/stream">
    <Parameter name="api_key" value="YOUR_PYAI_KEY"/>
    <Parameter name="aggressiveness" value="0.25"/>
    <Parameter name="decision_timeout_ms" value="3000"/>
    <Parameter name="lead_id" value="lead-42"/>
    <Parameter name="webhook" value="https://you/amd-events"/>
  </Stream>
</Start>
  <!-- your existing call flow continues here -->
</Response>
```

Use `<Start><Stream>`, NOT `<Connect><Stream>`. `<Start>` forks the audio and TwiML continues to your next verb, so the call still goes where it was going; `<Connect>` hands the media path to the socket and blocks TwiML until the stream ends, and because AMD is listen-only and never sends audio back the caller would hear dead air and a dialer would never reach the agent. (`<Connect>` is correct for Omni, which is a two-way voice agent.) Drop `machineDetection` from the call and keep your carrier.

On a Twilio-originated stream TWILIO owns the socket and relays only media/mark frames, so the pushed `amd` event does not reach you: a Twilio integration must read the decision from the `webhook` `<Parameter>` (or `GET /v1/amd/calls/{id}`). The socket push is for clients that drive the socket themselves.

From Twilio, authenticate with the `api_key` `<Parameter>` shown above (Twilio strips query strings from the `<Stream>` URL and cannot send headers; PyAI verifies the key from the stream's `start` frame before processing any audio, and closes connections that never present a valid key). Server-side clients may instead authenticate at the handshake with the `Sec-WebSocket-Protocol: pyai.v1, pyai-key.<API_KEY>` subprotocol pair or `?api_key=`. Requires the `amd:detect` scope. Mid-call, PyAI pushes an `amd` decision event on the socket (and to the per-call TwiML `webhook` parameter): `answered_by` (the routing class: `human`, `machine`, `sit_invalid`, `unknown`), `answered_by_twilio` (Twilio's exact `AnsweredBy` enum for drop-in routing parity), `subtype` (when available), `party_detected`, `voicemail_ready`, `confidence`, `decision_ms`, and a human-readable `reason`. `party_detected` is true for human or machine classification and false for unknown or invalid-number outcomes. `voicemail_ready` is false on classification events: detecting voicemail does not establish that recording has started or authorize dropping a message. `decision_ms` measures processed inbound audio through the decision, not elapsed time from carrier answer. A `machine` decision can carry a `subtype` (`voicemail`, `ivr`, `screening`, `music`); the stored call record folds that subtype into `answered_by`. Read the stored record with `GET /v1/amd/calls/{id}` or receive it on the account-wide `amd.call.completed` webhook (`webhook_url` in `POST /v1/amd/config`). The per-call `aggressiveness` `<Parameter>` overrides the account default from `POST /v1/amd/config`.


Decision events also include `engine_version` (decision-policy identity), `rule_id` (stable rule identifier), `speech_ms` (cumulative voiced audio excluding pauses, null if untracked), `silence_ms` (current contiguous silence), `speech_elapsed_ms` (audio since first voiced frame including pauses, null without onset), and `thresholds` (effective window, human dwell, early-yield hold and rule flags). These optional diagnostics also appear on new stored call details and the account-wide completion webhook; older records may omit them. Historical turn-yield reasons called elapsed time including trailing silence `speech`; use the new structured fields for comparisons. Confidence is a rule score, not a calibrated probability. Temporal human decisions require at least 200 ms of measured speech; isolated pulses cannot qualify just by waiting in silence. English introductions allow 2500 ms of silence for continuation, while complete recognized greetings can qualify after 500 ms of silence. Recognized call-progress tones are excluded from speech evidence. Without an explicit decision_window_ms override, the English 5000 ms base window can extend once by up to 2000 ms for recent speech or pending recognition; thresholds.effective_deadline_ms records the resulting deadline.

English streams may set `decision_window_ms` in `start.customParameters` (or a TwiML Parameter) to an integer or decimal integer string from 1000 to 15000. Omitting it keeps the configured default. Non-English overrides and invalid values emit an `error` event with code `invalid_decision_window` and close with code 1008. A longer window allows late evidence; it does not force unknown calls to a binary result. Continue real-time media pacing. Replay only the callee channel; never stereo-downmix the rep and callee. Sales Dialer Predictive recordings use channel 0 for the callee, Auto/Dynamic use channel 1; verify the source format before replaying.

Set `decision_timeout_ms` in `start.customParameters` to an integer or decimal integer string from 1000 to 15000 (for example `3000` or `5000`) to cap elapsed decision time in any supported language. The clock starts when the authenticated start frame and its parameters are accepted, before recognizer startup. Decisive evidence returns a result earlier. At the cutoff, AMD uses decisive evidence already received by that time; otherwise it emits `unknown` with `rule_id=decision_timeout`, including when no media arrives or recognition is still pending. It does not force a human or machine guess. This elapsed cap also bounds the final recognition wait and is not extended by the adaptive audio window. `decision_window_ms` remains a separate processed-audio limit; either limit can finish the decision first. Omit `decision_timeout_ms` to preserve existing behavior. Invalid values emit `invalid_stream_parameters` and close with code 1008. Timing diagnostics `decision_timeout_ms` and `decision_elapsed_ms` accompany opted-in results in socket events, per-call webhooks, stored details and account completion webhooks. The timeout bounds the decision budget, not network delivery or receiver acknowledgement; allow transport time for your own fallback timer.

Customer correlation fields such as `lead_id` and `campaign_id` supplied as additional TwiML Parameters / `start.customParameters` are returned under `custom_parameters` in the decision event, both webhook paths and the stored call detail. Fields remain nested and cannot override `call_id`, `answered_by` or other AMD output. Names must match `[A-Za-z0-9][A-Za-z0-9_.-]{0,63}`; values must be strings of at most 1024 UTF-8 bytes. At most 32 correlation fields and 8192 UTF-8 bytes of combined names and values are accepted. Invalid correlation fields emit `invalid_stream_parameters`. Reserved authentication, routing and AMD configuration parameters and names beginning with `_` are excluded from the echo. Do not send secrets as correlation fields. Query parameters on the per-call `webhook` URL are preserved in that URL; they are not copied into the JSON body.

Business introductions, generic requests for the reason for calling, unfinished requests to record a name, and recording disclosures are not decisive machine evidence by themselves. The default policy also abstains on tonal audio without recognized words; historical music results remain readable. These cases can return unknown while stronger evidence is absent. Screening and IVR may lead to a human: preserve that routing opportunity rather than treating every machine result as authorization to end the call. AMD emits one final classification per stream; it does not predict a later human pickup.

### When the decision arrives

Measured over ~1,850 real answered calls (US telephony, 8 kHz μ-law), streamed at real time:

| verdict | typical | 9 in 10 by |
|---|---|---|
| `human` | ~1.4 s | ~3.0 s |
| `machine` | ~2.2 s | ~3.2 s |

These historical latency measurements are not a guarantee for the current policy. The English default window is 5 seconds of processed audio with a bounded extension to 7 seconds for eligible calls. Size your fallback timer beyond the applicable decision window plus transport time, and react to the decision event.

### What to do with each verdict

| `answered_by` | subtype | do |
|---|---|---|
| `human` | — | connect the agent |
| `machine` | `voicemail` | consider a message only after independently establishing recording readiness |
| `machine` | `ivr` | a phone tree that may reach a person; navigate or route to an agent, and do not drop a message |
| `machine` | `screening` | an AI screener (iPhone/Google) is relaying to a person; treat as a live-ish path, not voicemail |
| `machine` | `music` | tonal audio and nothing transcribed — hold music or ringback, but also a greeting we failed to transcribe. Keep waiting; do **not** read it as a positive hold-music signal, and do not gate a drop on it |
| `sit_invalid` | — | dead/invalid number, stop retrying it |
| `unknown` | `silence` | answered but nothing came down the line, retry later rather than burning an agent slot |
| `unknown` | — | no decisive evidence; your default decides |

Read subtype on the wire, stored record, or completion webhook. A machine classification alone does not authorize hangup or voicemail drop. Screening and IVR can connect a real person. Music is a historical/legacy classification of tonal audio without text, not proof of an unreachable lead. Consider voicemail actions only for subtype voicemail and after establishing recording readiness separately; voicemail_ready is false on classification events.

### Choosing `aggressiveness`

The aggressiveness dial adjusts evidence timing thresholds. It does not manufacture machine evidence from a deadline: uncertain calls remain unknown at every setting. For live-agent dialers, retain the default and preserve the call when uncertain. Validate human-to-machine errors using all independently reviewed human calls as the denominator; confidence is a rule score, not a calibrated probability.

Billed per **answered** call (`amd.calls`), the first 5,000 answered calls/month are free, then $0.004/answered call; AMD bundled with PyAI telephony/Omni is included.



## OpenAPI

````yaml https://api.pyai.com/openapi.json get /v1/amd/stream
openapi: 3.1.0
info:
  title: PyAI API
  version: 2.9.0
  description: >-
    Telephony-native Voice AI behind one bearer key:


    - **Hear**, speech-to-text · `POST /v1/audio/transcriptions` (streaming +
    batch)

    - **Speak**, text-to-speech · `POST /v1/audio/speech`, `GET /v1/voices`

    - **Clone**, custom voices from a short clip · `/v1/voice/clones`

    - **Cast**, auto-directed expressive voiceovers · `/v1/cast`

    - **Dub**, asynchronous audio and video dubbing · `POST /v1/dub`
    ([guide](https://docs.pyai.com/guides/dub-overview))

    - **Cue configuration frames (reserved)**; Hear streams skip knowledge-base
    retrieval by default

    - **Omni**, full-duplex agentic voice (speech-to-speech, grounded in your
    knowledge bases + tools) · `/v1/omni`

    - **Knowledge Bases**, hosted grounding for Omni: create bases, add
    documents (file, URL, or text), crawl a public website, bind to agents or
    org defaults · `/v1/knowledgebases`

    - **AMD API**, answering-machine detection: know *who or what* answered a
    call (human, voicemail, IVR, iPhone/Google screening, dead number) with the
    reason it decided · `wss …/v1/amd/stream` (Twilio Media Streams drop-in),
    `POST /v1/amd/config`, `GET /v1/amd/calls/{id}`

    - **Agents Beta**, the live console feature to create, configure, test, and
    connect Omni voice agents without code. Beta features and limits may change.


    ## Authentication


    Create a key in the [console](https://console.pyai.com) (it is shown once)
    and send it as a bearer token:


    ```

    Authorization: Bearer pyai_live_...

    ```


    Keys are environment-scoped: `pyai_live_...` (production) and
    `pyai_test_...` (sandbox). `POST /v1/sandbox/keys` creates an instant,
    short-lived test key without login or billing. Account signup also creates a
    sandbox key. Live keys consume prepaid credit; phone verification may unlock
    promotional credit under graduated-signup rules, but credit is not
    guaranteed at signup.


    Keys are self-validating signed tokens: they work on every PyAI surface the
    instant they are created, no activation or propagation delay. Treat them as
    opaque strings (up to 512 chars) and never parse their contents.


    WebSocket endpoints can't use request headers from a browser, so pass the
    key as a **subprotocol** instead:


    ```

    Sec-WebSocket-Protocol: pyai.v1, pyai-key.pyai_live_...

    ```


    (server-side clients may instead append `?api_key=...` to the URL). Your key
    is authenticated on the upgrade and never reaches the model.


    ## Quickstart, Hear (speech-to-text)


    ```

    curl https://api.pyai.com/v1/audio/transcriptions \
      -H "Authorization: Bearer $PYAI_API_KEY" \
      -F file=@audio.wav -F model=pyai-hear
    # -> { "text": "..." }

    ```


    ## Quickstart, Speak (text-to-speech)


    ```

    curl https://api.pyai.com/v1/audio/speech \
      -H "Authorization: Bearer $PYAI_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model":"pyai-speak","input":"Hello from PyAI.","voice":"voice_abc"}' \
      --output speech.wav
    ```


    `voice` is a stock voice id from `GET /v1/voices` (the curated prebuilt
    catalog with personas and avatars) or a cloned voice id from
    `/v1/voice/clones`. Omit it to use the platform default voice
    (`stock_dorit_en_us`).


    ## Quickstart, Omni (realtime voice agent)


    Omni is **zero-state, there is nothing to create first.** Open a WebSocket,
    pass your key as a subprotocol, and send the agent's behavior (voice,
    persona, knowledge endpoint) in the first `configure` frame:


    ```

    wss://api.pyai.com/v1/omni?session_label=support&format=pcm16&rate=24000
      Sec-WebSocket-Protocol: pyai.v1, pyai-key.$PYAI_API_KEY
    ```


    The session is authorized by your key's **organization**; `session_label` is
    an **optional, opaque** tag (echoed to your own knowledge endpoint for
    correlation), omit it or use any value. When `session_label` equals a
    **`/v1/agents` profile id**, the engine loads persona, voice, and **greeting
    message** from that profile (turn-0 playback). `format` and `rate` are
    load-bearing on the connect URL (the SDK sets them). Prefix each PCM16 frame
    with byte `0x01`; prefix control JSON with byte `0x03`. **Optional
    convenience:** pre-store config via `POST /v1/agents` (including `greeting`,
    `consent_line`, `recordings_enabled`) and pass its id as `session_label`, or
    send everything inline in the post-handshake `configure` frame. Not required
    to connect.


    For reproducible eval runs, determinism controls (`seed`/`temperature`) ride
    the Omni session's `configure` frame, which the gateway passes through
    unchanged, they are honored once the engine supports them; no platform
    change is required.


    ## Scopes


    | Scope | Grants |

    | --- | --- |

    | `hear:transcribe` | `POST /v1/audio/transcriptions` |

    | `hear:stream` | `GET /v1/audio/transcriptions/stream` (WebSocket) |

    | `hear:configure` | `GET`/`PUT /v1/hear/vocabulary` |

    | `speak:synthesize` | `POST /v1/audio/speech` (Speak) |

    | `speak:clone` | `/v1/voice/clones` (Clone) |

    | `speak:design` | `/v1/voice/design` (Speak) |

    | `omni:session` | `/v1/omni` and `POST /v1/omni/sessions` (mint a browser
    session token) |

    | `omni:read` | `/v1/omni/calls` (Omni post-call records) |

    | `kb:manage` | `/v1/knowledgebases/*` (hosted knowledge bases for Omni
    grounding) |

    | `transcribe:jobs` | `/v1/transcription/jobs` |

    | `trace:configure` | `/v1/trace/config`, `/v1/trace/rule-packs` (Trace
    management) |

    | `trace:read` | `/v1/trace/interactions`, `/violations`, `/findings`,
    `/exposure` (Trace reads) |

    | `recap:configure` | `/v1/recap/config` (Recap management) |

    | `recap:configure` | `/v1/recap/crm-config` (Salesforce field mapping) |

    | `recap:read` | `/v1/recap/calls` (Recap reads and speaker-role
    corrections) |

    | `amd:detect` | `wss …/v1/amd/stream` (AMD realtime detection, Twilio
    drop-in) |

    | `amd:configure` | `/v1/amd/config` (AMD operating-point dial + webhook) |

    | `amd:read` | `/v1/amd/calls` (AMD decision records) |

    | `telephony:manage` | `/v1/telephony/*` (managed numbers) |


    `GET /v1/models`, `GET /v1/voices`, and `GET /v1/me` need no specific scope,
    any active key may call them. Wildcards (`hear:*`, `speak:*`, …, and the
    global `*`) grant every scope in their family.


    ## Canonical endpoints


    One row per product surface, endpoint, auth, required scope, and lifecycle
    status. **live** = generally available; **beta** = available now with
    features or limits that may change; **unavailable** = reserved in the
    contract but not active on the serving route.


    | Product | Endpoint | Auth | Scope | Status |

    | --- | --- | --- | --- | --- |

    | Identity | `GET /v1/me` | Bearer | _any active key_ | live |

    | Models | `GET /v1/models` | Bearer | _any active key_ | live |

    | Voices | `GET /v1/voices`, `GET /v1/voices/{id}` | Bearer | _any active
    key_ | live |

    | Hear (batch) | `POST /v1/audio/transcriptions` | Bearer |
    `hear:transcribe` | live |

    | Hear (vocabulary settings) | `GET`/`PUT /v1/hear/vocabulary` | Bearer |
    `hear:configure` | live |

    | Hear (streaming) | `GET /v1/audio/transcriptions/stream` (WS) |
    Subprotocol | `hear:stream` | live |

    | Cue | `GET /v1/audio/transcriptions/stream` + grounding (WS) | Subprotocol
    | `hear:stream` | unavailable |

    | Hear (async batch) | `POST`/`GET /v1/transcription/jobs` | Bearer |
    `transcribe:jobs` | live |

    | Speak (TTS) | `POST /v1/audio/speech` | Bearer | `speak:synthesize` | live
    |

    | Clone | `GET`/`POST /v1/voice/clones` | Bearer | `speak:clone` | live |

    | Speak (design) | `/v1/voice/design` | Bearer | `speak:design` | live |

    | Omni | `wss …/v1/omni?session_label=` | Subprotocol | `omni:session` |
    live |

    | Agent profiles (optional config) | `/v1/agents`, `/v1/agents/{id}` |
    Bearer | `omni:session` | live |

    | Knowledge Bases (hosted grounding) | `/v1/knowledgebases/*`, `PUT
    /v1/agents/{id}/knowledgebases` | Bearer | `kb:manage` (`omni:session` for
    the binding) | live |

    | Trace (config) | `/v1/trace/config`, `/v1/trace/rule-packs` | Bearer |
    `trace:configure` | beta |

    | Trace (reads) | `/v1/trace/interactions`, `/violations`, `/findings`,
    `/exposure` | Bearer | `trace:read` | beta |

    | Recap (config) | `/v1/recap/config` | Bearer | `recap:configure` | live |

    | Recap (CRM) | `/v1/recap/crm-config` | Bearer | `recap:configure` | live |

    | Integrations (Zapier) | `/v1/integrations/events`,
    `/v1/integrations/zapier/hooks` | Bearer | _any active key_ | live |

    | Recap (reads) | `/v1/recap/calls` | Bearer | `recap:read` | live |

    | Omni call records | `/v1/omni/calls`, `/v1/omni/calls/{id}` | Bearer |
    `omni:read` | live |

    | AMD (stream) | `wss …/v1/amd/stream` (Twilio Media Streams drop-in) |
    TwiML `<Parameter name="api_key">` (from Twilio) or subprotocol
    (server-side) | `amd:detect` | live |

    | AMD (config) | `GET`/`POST /v1/amd/config` | Bearer | `amd:configure` |
    live |

    | AMD (reads) | `GET /v1/amd/calls`, `/v1/amd/calls/{id}` | Bearer |
    `amd:read` | live |

    | Telephony | `/v1/telephony/*` | Bearer | `telephony:manage` | live |

    | Agents (console builder) | `https://console.pyai.com/agents` | Console
    session |, | beta |


    WebSocket surfaces authenticate with the `Sec-WebSocket-Protocol: pyai.v1,
    pyai-key.<API_KEY>` subprotocol pair (or `?api_key=` server-side);
    everything else takes the `Authorization: Bearer` key. Managed-number calls
    return 404 until the PyAI network is enabled for the account.


    ## Rate limits & billing


    Every key has a per-second rate limit (with burst) and a cap on concurrent
    realtime sessions. Exceeding either returns `429` with a `Retry-After`
    header. Usage is metered per minute of audio, transcription minutes (Hear),
    synthesized audio minutes (Speak), and realtime session minutes (Omni), and
    billed against your plan and credits. List prices: Hear $0.001/min (async
    Transcribe $0.0005/min), Speak $0.04/min, Omni $0.05/min including speech
    plus brain, and Agents Live Beta $0.08/min. Managed telephony is separate at
    $0.01/min. English Natural (`en1`) is available on Speak (streaming and
    buffered) and Omni; Omni acknowledges the canonical id and `voice_tier:
    natural`. Hindi uses the Standard-tier voices `hi1`–`hi4` on Omni; the
    former Hindi Natural aliases (`hi5`–`hi8`) are retired and no longer in the
    catalog. Standard and Natural voices are included in their product's base
    rate with no voice-tier add-on. The AMD API bills per **answered** call, the
    first 5,000 answered calls each month are free, then $0.004/answered call
    (no-answers, busies, and failed calls are free; AMD bundled with PyAI
    telephony/Omni is included at no charge). AI products (Hear, Speak, Omni)
    bill **per second by default**, the pulse is applied once to each meter's
    invoice-period total, so many short sessions are summed and rounded a single
    time (never minute-rounded per call), and an empty/failed call bills
    nothing. Coarser pulses are available as an optional enterprise override.
    Managed telephony minutes keep a 1-minute pulse. Per-character Speak billing
    is available on enterprise contracts.
  contact:
    name: PyAI
    url: https://pyai.com
servers:
  - url: https://api.pyai.com
    description: Production
security:
  - apiKey: []
  - xApiKey: []
tags:
  - name: Dub
    description: >-
      Asynchronous dubbing: submit a recording, choose the languages, poll the
      job and download the output.
  - name: Omni
    description: >-
      The flagship: build an AI voice agent with one WebSocket (`GET /v1/omni`)
      and one `configure` frame, nothing to pre-create. This group also holds
      the optional browser-token mint and the post-call records.
  - name: Knowledge Bases
    description: >-
      Hosted knowledge bases for Omni grounding: create a base, add documents
      (file upload, URL fetch, or pasted text), then bind it to agent profiles
      or set org-wide defaults. Bound bases are retrieved per turn, no
      `kb_endpoint` of your own required.
  - name: Identity
    description: >-
      Introspect the calling key: org/project, env, granted scopes, and
      limits/credit posture. Use it to self-diagnose a 401/403/402.
  - name: Speech To Text (Hear)
    description: Speech-to-text (streaming + batch)
  - name: Text To Speech (Speak)
    description: Text-to-speech, stock voices, and prompt-to-voice design
  - name: Clone
    description: Enroll, list, and delete custom voices from a short reference clip
  - name: Models
    description: Model catalog
  - name: Sandbox
    description: >-
      Zero-friction onboarding for coding agents: mint a free, instant, no-card
      sandbox key with no human steps.
  - name: Startup Program
    description: >-
      PyAI for Startups: $20k to $100k in PyAI credit for early-stage voice
      teams. Public application endpoint; review and activation happen out of
      band.
  - name: Cast
    description: >-
      Auto-directed, expressive multi-line voiceover projects and asynchronous
      renders.
  - name: Transcription Jobs
    description: Async batch transcription
  - name: Agents
    description: >-
      Agent profiles used by the live Agents Beta console and available directly
      through the API. Store Omni session config (persona, greeting, voice,
      conversation knobs) and reference it by id instead of sending a full
      `configure` frame each call. Profiles remain optional for direct
      `/v1/omni` integrations.
  - name: Trace
    description: >-
      Compliance & guardrails: per-agent config, rule packs, and the exposure /
      violations / interaction-evidence read views
  - name: AMD
    description: >-
      Answering-machine detection: know who or what answered a call (human,
      voicemail, IVR, iPhone/Google screening, dead number), with the reason it
      decided. Twilio Media Streams drop-in over `wss …/v1/amd/stream`; one
      operating-point dial; billed per answered call.
  - name: Telephony
    description: >-
      Managed phone numbers: search, provision, route to an agent, and release.
      Call minutes bill on telephony.minutes ($0.01/min).
  - name: WhatsApp
    description: >-
      WhatsApp Business Calling: register a WhatsApp Business number, enable
      calling, and let an Omni agent answer (and, with the user's permission,
      place) WhatsApp voice calls. Requires the `telephony:manage` scope.
  - name: Call Integrations
    description: >-
      Signed provider webhooks that import completed calls into Hear, Recap, and
      offline Trace.
paths:
  /v1/amd/stream:
    get:
      tags:
        - AMD
      summary: Answering-machine detection (WebSocket)
      description: >-
        Realtime answering-machine detection over a WebSocket. **This surface
        speaks Twilio's Media Streams protocol natively** (`start` / `media` /
        `stop` frames, G.711 μ-law 8 kHz base64, ~20 ms), so migrating from
        Twilio AMD is a one-line-TwiML change, point the call's media at PyAI,
        keep your carrier and your code.


        ```xml

        <Response><Start>
          <Stream url="wss://api.pyai.com/v1/amd/stream">
            <Parameter name="api_key" value="YOUR_PYAI_KEY"/>
            <Parameter name="aggressiveness" value="0.25"/>
            <Parameter name="decision_timeout_ms" value="3000"/>
            <Parameter name="lead_id" value="lead-42"/>
            <Parameter name="webhook" value="https://you/amd-events"/>
          </Stream>
        </Start>
          <!-- your existing call flow continues here -->
        </Response>

        ```


        Use `<Start><Stream>`, NOT `<Connect><Stream>`. `<Start>` forks the
        audio and TwiML continues to your next verb, so the call still goes
        where it was going; `<Connect>` hands the media path to the socket and
        blocks TwiML until the stream ends, and because AMD is listen-only and
        never sends audio back the caller would hear dead air and a dialer would
        never reach the agent. (`<Connect>` is correct for Omni, which is a
        two-way voice agent.) Drop `machineDetection` from the call and keep
        your carrier.


        On a Twilio-originated stream TWILIO owns the socket and relays only
        media/mark frames, so the pushed `amd` event does not reach you: a
        Twilio integration must read the decision from the `webhook`
        `<Parameter>` (or `GET /v1/amd/calls/{id}`). The socket push is for
        clients that drive the socket themselves.


        From Twilio, authenticate with the `api_key` `<Parameter>` shown above
        (Twilio strips query strings from the `<Stream>` URL and cannot send
        headers; PyAI verifies the key from the stream's `start` frame before
        processing any audio, and closes connections that never present a valid
        key). Server-side clients may instead authenticate at the handshake with
        the `Sec-WebSocket-Protocol: pyai.v1, pyai-key.<API_KEY>` subprotocol
        pair or `?api_key=`. Requires the `amd:detect` scope. Mid-call, PyAI
        pushes an `amd` decision event on the socket (and to the per-call TwiML
        `webhook` parameter): `answered_by` (the routing class: `human`,
        `machine`, `sit_invalid`, `unknown`), `answered_by_twilio` (Twilio's
        exact `AnsweredBy` enum for drop-in routing parity), `subtype` (when
        available), `party_detected`, `voicemail_ready`, `confidence`,
        `decision_ms`, and a human-readable `reason`. `party_detected` is true
        for human or machine classification and false for unknown or
        invalid-number outcomes. `voicemail_ready` is false on classification
        events: detecting voicemail does not establish that recording has
        started or authorize dropping a message. `decision_ms` measures
        processed inbound audio through the decision, not elapsed time from
        carrier answer. A `machine` decision can carry a `subtype` (`voicemail`,
        `ivr`, `screening`, `music`); the stored call record folds that subtype
        into `answered_by`. Read the stored record with `GET /v1/amd/calls/{id}`
        or receive it on the account-wide `amd.call.completed` webhook
        (`webhook_url` in `POST /v1/amd/config`). The per-call `aggressiveness`
        `<Parameter>` overrides the account default from `POST /v1/amd/config`.



        Decision events also include `engine_version` (decision-policy
        identity), `rule_id` (stable rule identifier), `speech_ms` (cumulative
        voiced audio excluding pauses, null if untracked), `silence_ms` (current
        contiguous silence), `speech_elapsed_ms` (audio since first voiced frame
        including pauses, null without onset), and `thresholds` (effective
        window, human dwell, early-yield hold and rule flags). These optional
        diagnostics also appear on new stored call details and the account-wide
        completion webhook; older records may omit them. Historical turn-yield
        reasons called elapsed time including trailing silence `speech`; use the
        new structured fields for comparisons. Confidence is a rule score, not a
        calibrated probability. Temporal human decisions require at least 200 ms
        of measured speech; isolated pulses cannot qualify just by waiting in
        silence. English introductions allow 2500 ms of silence for
        continuation, while complete recognized greetings can qualify after 500
        ms of silence. Recognized call-progress tones are excluded from speech
        evidence. Without an explicit decision_window_ms override, the English
        5000 ms base window can extend once by up to 2000 ms for recent speech
        or pending recognition; thresholds.effective_deadline_ms records the
        resulting deadline.


        English streams may set `decision_window_ms` in `start.customParameters`
        (or a TwiML Parameter) to an integer or decimal integer string from 1000
        to 15000. Omitting it keeps the configured default. Non-English
        overrides and invalid values emit an `error` event with code
        `invalid_decision_window` and close with code 1008. A longer window
        allows late evidence; it does not force unknown calls to a binary
        result. Continue real-time media pacing. Replay only the callee channel;
        never stereo-downmix the rep and callee. Sales Dialer Predictive
        recordings use channel 0 for the callee, Auto/Dynamic use channel 1;
        verify the source format before replaying.


        Set `decision_timeout_ms` in `start.customParameters` to an integer or
        decimal integer string from 1000 to 15000 (for example `3000` or `5000`)
        to cap elapsed decision time in any supported language. The clock starts
        when the authenticated start frame and its parameters are accepted,
        before recognizer startup. Decisive evidence returns a result earlier.
        At the cutoff, AMD uses decisive evidence already received by that time;
        otherwise it emits `unknown` with `rule_id=decision_timeout`, including
        when no media arrives or recognition is still pending. It does not force
        a human or machine guess. This elapsed cap also bounds the final
        recognition wait and is not extended by the adaptive audio window.
        `decision_window_ms` remains a separate processed-audio limit; either
        limit can finish the decision first. Omit `decision_timeout_ms` to
        preserve existing behavior. Invalid values emit
        `invalid_stream_parameters` and close with code 1008. Timing diagnostics
        `decision_timeout_ms` and `decision_elapsed_ms` accompany opted-in
        results in socket events, per-call webhooks, stored details and account
        completion webhooks. The timeout bounds the decision budget, not network
        delivery or receiver acknowledgement; allow transport time for your own
        fallback timer.


        Customer correlation fields such as `lead_id` and `campaign_id` supplied
        as additional TwiML Parameters / `start.customParameters` are returned
        under `custom_parameters` in the decision event, both webhook paths and
        the stored call detail. Fields remain nested and cannot override
        `call_id`, `answered_by` or other AMD output. Names must match
        `[A-Za-z0-9][A-Za-z0-9_.-]{0,63}`; values must be strings of at most
        1024 UTF-8 bytes. At most 32 correlation fields and 8192 UTF-8 bytes of
        combined names and values are accepted. Invalid correlation fields emit
        `invalid_stream_parameters`. Reserved authentication, routing and AMD
        configuration parameters and names beginning with `_` are excluded from
        the echo. Do not send secrets as correlation fields. Query parameters on
        the per-call `webhook` URL are preserved in that URL; they are not
        copied into the JSON body.


        Business introductions, generic requests for the reason for calling,
        unfinished requests to record a name, and recording disclosures are not
        decisive machine evidence by themselves. The default policy also
        abstains on tonal audio without recognized words; historical music
        results remain readable. These cases can return unknown while stronger
        evidence is absent. Screening and IVR may lead to a human: preserve that
        routing opportunity rather than treating every machine result as
        authorization to end the call. AMD emits one final classification per
        stream; it does not predict a later human pickup.


        ### When the decision arrives


        Measured over ~1,850 real answered calls (US telephony, 8 kHz μ-law),
        streamed at real time:


        | verdict | typical | 9 in 10 by |

        |---|---|---|

        | `human` | ~1.4 s | ~3.0 s |

        | `machine` | ~2.2 s | ~3.2 s |


        These historical latency measurements are not a guarantee for the
        current policy. The English default window is 5 seconds of processed
        audio with a bounded extension to 7 seconds for eligible calls. Size
        your fallback timer beyond the applicable decision window plus transport
        time, and react to the decision event.


        ### What to do with each verdict


        | `answered_by` | subtype | do |

        |---|---|---|

        | `human` | — | connect the agent |

        | `machine` | `voicemail` | consider a message only after independently
        establishing recording readiness |

        | `machine` | `ivr` | a phone tree that may reach a person; navigate or
        route to an agent, and do not drop a message |

        | `machine` | `screening` | an AI screener (iPhone/Google) is relaying
        to a person; treat as a live-ish path, not voicemail |

        | `machine` | `music` | tonal audio and nothing transcribed — hold music
        or ringback, but also a greeting we failed to transcribe. Keep waiting;
        do **not** read it as a positive hold-music signal, and do not gate a
        drop on it |

        | `sit_invalid` | — | dead/invalid number, stop retrying it |

        | `unknown` | `silence` | answered but nothing came down the line, retry
        later rather than burning an agent slot |

        | `unknown` | — | no decisive evidence; your default decides |


        Read subtype on the wire, stored record, or completion webhook. A
        machine classification alone does not authorize hangup or voicemail
        drop. Screening and IVR can connect a real person. Music is a
        historical/legacy classification of tonal audio without text, not proof
        of an unreachable lead. Consider voicemail actions only for subtype
        voicemail and after establishing recording readiness separately;
        voicemail_ready is false on classification events.


        ### Choosing `aggressiveness`


        The aggressiveness dial adjusts evidence timing thresholds. It does not
        manufacture machine evidence from a deadline: uncertain calls remain
        unknown at every setting. For live-agent dialers, retain the default and
        preserve the call when uncertain. Validate human-to-machine errors using
        all independently reviewed human calls as the denominator; confidence is
        a rule score, not a calibrated probability.


        Billed per **answered** call (`amd.calls`), the first 5,000 answered
        calls/month are free, then $0.004/answered call; AMD bundled with PyAI
        telephony/Omni is included.
      operationId: amdStream
      responses:
        '101':
          description: >-
            WebSocket upgrade, Twilio Media Streams protocol; PyAI emits `amd`
            decision events.
        '401':
          $ref: '#/components/responses/Unauthorized'
        '403':
          $ref: '#/components/responses/Forbidden'
components:
  responses:
    Unauthorized:
      description: 'Missing or invalid API key (`code: unauthorized`)'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    Forbidden:
      description: >-
        Key lacks the required scope or the origin is not allow-listed (`code:
        forbidden | origin_not_allowed`)
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  schemas:
    Error:
      type: object
      description: >-
        OpenAI-compatible error envelope returned by the gateway data plane
        (401/402/403/429). Control-plane request/resource errors use Problem
        (application/problem+json) instead.
      required:
        - error
      properties:
        error:
          type: object
          required:
            - message
          properties:
            message:
              type: string
              description: Human-readable explanation.
            type:
              type: string
              description: Error category, e.g. rate_limit_error.
            code:
              $ref: '#/components/schemas/ErrorCode'
            param:
              type: string
              nullable: true
              description: Offending parameter when applicable, else null.
    ErrorCode:
      type: string
      description: >-
        Stable, machine-readable error code. Branch on this rather than the
        human `message`.
      enum:
        - invalid_request_error
        - invalid_session_label
        - unauthorized
        - forbidden
        - origin_not_allowed
        - credit_exhausted
        - key_budget_exceeded
        - insufficient_quota
        - rate_limit_exceeded
        - concurrency_limit_exceeded
        - daily_cap_exceeded
  securitySchemes:
    apiKey:
      type: http
      scheme: bearer
      description: 'Use `Authorization: Bearer pyai_live_...` (or `pyai_test_...`).'
    xApiKey:
      type: apiKey
      in: header
      name: x-api-key
      description: >-
        Header alias for bearer auth on HTTP endpoints. WebSocket auth uses the
        subprotocol pair `pyai.v1, pyai-key.<API_KEY>`.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.