> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Quickstart

> Mint an instant key and make your first PyAI call: Agents powered by Omni, Hear, Speak, Clone, Cast, AMD, or Trace.

PyAI is telephony-native Voice AI behind one API key: **Agents**, voice agents built through UI or API and run on the **Omni** realtime
endpoint; **Hear** for speech-to-text; **Speak** for
text-to-speech; **Clone** for a custom voice from a short clip; **Cast** for
expressive voice rendering; **AMD** for answering-machine detection; and
**Trace** for compliance guardrails. The REST API is OpenAI-compatible at
`https://api.pyai.com/v1`.

For a saved agent you can manage in both the UI and code, start with
[Build and run Agents](/agents/getting-started) or
[Create agents via API](/guides/create-agents-api). The direct Omni examples below
remain useful when you want inline session configuration without a saved profile.

## Step 1, get a key

**Fastest: an instant sandbox key, no signup, email, or card.** It works
immediately, skips the credit gate, and is bounded by daily caps:

```bash theme={null}
export PYAI_API_KEY="$(
  curl -sS -X POST https://api.pyai.com/v1/sandbox/keys \
    | python -c 'import json,sys; print(json.load(sys.stdin)["api_key"])'
)"
```

Verify it in 5 seconds (`GET /v1/me` echoes the org, scopes, and credit posture
your key resolves to):

```bash theme={null}
curl https://api.pyai.com/v1/me -H "Authorization: Bearer $PYAI_API_KEY"
```

A working sandbox key returns JSON like this. Read `scopes` before you call a
product. A sandbox mint includes Hear, Speak, Clone, Omni, Cast, AMD, Dub,
Recap, and Trace. It does **not** include voice design or hosted
knowledge-base management.

```json theme={null}
{
  "object": "identity",
  "key_id": "key_...",
  "org_id": "org_...",
  "project_id": "proj_...",
  "env": "test",
  "status": "active",
  "org_status": "active",
  "plan": "payg",
  "scopes": [
    "hear:transcribe",
    "hear:stream",
    "hear:configure",
    "transcribe:jobs",
    "speak:synthesize",
    "speak:clone",
    "omni:session",
    "omni:read",
    "amd:detect",
    "amd:configure",
    "amd:read",
    "cast:render",
    "dub:render",
    "recap:configure",
    "recap:read",
    "trace:configure",
    "trace:read",
    "text:cleanup"
  ],
  "limits": { "rps": 20, "burst": 40, "concurrency": 2, "daily_unit_cap": 10 },
  "credit": { "gated": false, "available_cents": 0 }
}
```

Install an SDK if you want retries and multipart handled for you:

```bash theme={null}
pip install pyai-sdk
# or
npm install @pyai/sdk
```

See [SDKs](/guides/sdks) for when to use the SDK vs raw HTTP.

**For production:** [create an account](https://console.pyai.com/playground?signup=1\&path=api\&utm_source=docs.pyai.com\&utm_medium=docs)
and mint a `pyai_live_` key in the console. Live usage is billed against
prepaid credit; see the [pricing page](https://pyai.com/pricing) for current
terms. HTTP auth accepts `Authorization: Bearer <key>` or the header
alias `x-api-key: <key>`.

<Note>
  **Fastest start:** scaffold a complete, runnable example in one command, no
  clone, no setup ceremony.

  ```bash theme={null}
  npm create pyai-app@latest                   # pick an example interactively
  npm create pyai-app@latest openai-drop-in     # already on OpenAI? migrate by changing the base URL
  ```

  Browse them all at [github.com/atomsai/pyai-examples](https://github.com/atomsai/pyai-examples).
</Note>

## Step 2, pick what you're building

Not sure which surface? Start at [Choose your path](/choose-your-path).

<CardGroup cols={2}>
  <Card title="AI voice agent (Omni)" href="/guides/omni-overview">Replace the speech-to-text, LLM, text-to-speech, VAD, and turn-detection cascade with one WebSocket.</Card>
  <Card title="Speech To Text (Hear)" href="/guides/hear-overview">Choose one-file, live-streaming, or timestamped async transcription.</Card>
  <Card title="Text To Speech (Speak)" href="/guides/speak-overview">Choose a voice, delivery mode, output format, and sample rate.</Card>
  <Card title="Clone a voice" href="/guides/voice-cloning">Enroll a consented voice from a short clip and use it on Speak or Omni.</Card>
  <Card title="Expressive voice (Cast)" href="/guides/cast-overview">Direct, preview, and render expressive multi-line speech.</Card>
  <Card title="Answering-machine detection (AMD)" href="#answering-machine-detection-amd">Know who or what answered, human, voicemail, IVR, screening, with the reason. Twilio drop-in.</Card>
  <Card title="Call summaries and actions (Recap)" href="#call-summaries-and-actions-recap">Turn completed transcripts into notes, action items, talk ratio, signals, and fields.</Card>
  <Card title="Compliance guardrails (Trace)" href="#compliance-guardrails-trace">Rule packs (TCPA, HIPAA, PII) and scorecards for eligible calls.</Card>
  <Card title="No code: Agents Live Beta" href="/agents/getting-started">Configure, test, and connect an Omni voice agent from the console.</Card>
</CardGroup>

### AI voice agent (Omni)

Omni is the complete voice agent behind one realtime endpoint. It hears,
reasons, calls tools, grounds answers in your knowledge, and speaks back. Your app
streams audio to one WebSocket instead of operating separate speech-to-text,
LLM, text-to-speech, VAD, turn-detection, and interruption components.

Connect and confirm that your key can configure a session:

```js theme={null}
import WebSocket from "ws"; // npm install ws

const ws = new WebSocket(
  "wss://api.pyai.com/v1/omni?format=pcm16&rate=24000",
  ["pyai.v1", `pyai-key.${process.env.PYAI_API_KEY}`],
);

ws.on("open", () => {
  const configure = Buffer.from(JSON.stringify({
    type: "configure",
    voice_id: "stock_dorit_en_us",
    persona: "You are a concise appointment assistant.",
  }));
  ws.send(Buffer.concat([Buffer.from([0x03]), configure]));
});

ws.on("message", (data, isBinary) => {
  if (!isBinary) return;
  const frame = Buffer.from(data);
  if (frame[0] !== 0x03) return;
  const event = JSON.parse(frame.subarray(1).toString("utf8"));
  console.log(event);
  if (event.event === "configured") ws.close();
});
```

Success is a control frame within about two seconds:

```json theme={null}
{ "event": "configured", "voice_id": "stock_dorit_en_us", "tools": 0 }
```

That proves authentication, framing, and configuration. To stream microphone
audio and play the reply, continue with the browser tutorial.

Start with the [Omni overview](/guides/omni-overview), follow the
[browser tutorial](/guides/browser-voice-agent), or configure and test an agent
without code in the [console](/agents/getting-started). For every frame and
event, use the [wire protocol reference](/realtime/omni-protocol).

<Warning>
  If you are updating an older Omni integration, follow the
  [Omni endpoint migration guide](/guides/migrate-omni-v2-chat) before changing
  your client.
</Warning>

### Speech-to-text (Hear)

One POST transcribes a file:

<CodeGroup>
  ```bash curl theme={null}
  curl https://api.pyai.com/v1/audio/transcriptions \
    -H "Authorization: Bearer $PYAI_API_KEY" \
    -F file=@audio.wav -F model=pyai-hear
  ```

  ```python Python theme={null}
  import os
  from pyai import PyAI  # pip install pyai-sdk

  pyai = PyAI(api_key=os.environ["PYAI_API_KEY"])
  result = pyai.audio.transcriptions.create(file=open("audio.wav", "rb"), model="pyai-hear")
  print(result["text"])
  ```

  ```ts Node theme={null}
  import PyAI from "@pyai/sdk"; // npm install @pyai/sdk
  import { readFile } from "node:fs/promises";

  const pyai = new PyAI({ apiKey: process.env.PYAI_API_KEY! });
  const result = await pyai.audio.transcriptions.create({
    file: new Blob([await readFile("audio.wav")]),
    filename: "audio.wav",
    model: "pyai-hear",
  });
  console.log(result.text);
  ```
</CodeGroup>

Expected response:

```json theme={null}
{ "text": "Your appointment is confirmed for Thursday at two." }
```

Need live partials (voice bots, captions, agent assist)? Stream over WebSocket
instead. PyAI Hear's first partial measured about 200 ms in-region:
[streaming STT guide](/guides/streaming-stt). For large backlogs, the async
[batch jobs guide](/guides/async-transcription-jobs) covers timestamped
recordings, speaker labels, SRT/VTT, signed webhooks, and retention.

### Text-to-speech (Speak)

<CodeGroup>
  ```bash curl theme={null}
  curl https://api.pyai.com/v1/audio/speech \
    -H "Authorization: Bearer $PYAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model":"pyai-speak","input":"Hello from PyAI.","voice":"stock_dorit_en_us"}' \
    --output hello.wav
  ```

  ```python Python theme={null}
  import os
  from pyai import PyAI  # pip install pyai-sdk

  pyai = PyAI(api_key=os.environ["PYAI_API_KEY"])
  audio = pyai.audio.speech(input="Hello from PyAI.", voice="stock_dorit_en_us")
  open("hello.wav", "wb").write(audio)
  ```

  ```ts Node theme={null}
  import PyAI from "@pyai/sdk"; // npm install @pyai/sdk

  const pyai = new PyAI({ apiKey: process.env.PYAI_API_KEY! });
  const audio = await pyai.audio.speech({ input: "Hello from PyAI.", voice: "stock_dorit_en_us" });
  // write `audio` (ArrayBuffer) to hello.wav
  ```
</CodeGroup>

`voice` is a stock voice id from `GET /v1/voices` (the curated prebuilt catalog
with personas and avatars) or a cloned voice id from
[Clone](/guides/voice-cloning). Omit it and Speak uses the default stock voice,
`stock_dorit_en_us`. Send `voice` explicitly in production code so a change to
the default can never change how your app sounds.
`pyai-speak` is the canonical model. OpenAI's `tts-1` and `tts-1-hd` model names
remain accepted for compatibility. The preset names `alloy`, `echo`, `fable`,
`onyx`, `nova`, and `shimmer` also remain drop-in aliases for PyAI stock voices.

See the [Speak guide](/guides/speak-overview) for the voice catalog, streaming
vs buffered delivery, output formats, sample rates, and telephony.

### Clone a voice

Enroll a consented English reference clip, poll until `status` is `ready`, then
pass the returned `voice_id` to Speak or Omni:

```bash theme={null}
curl https://api.pyai.com/v1/voice/clones \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -F "name=Ava, brand voice" \
  -F "file=@reference.wav"
```

Expected enroll response:

```json theme={null}
{ "id": "voice_abc", "name": "Ava, brand voice", "status": "pending" }
```

There is no GET-by-id for clones. Poll `GET /v1/voice/clones` (or
`pyai.clones.get` in the SDK, which filters that list) until `status` is
`ready` or `failed`. Then pass that `id` as `voice` on Speak or Omni. If the
first render answers `409 voice_propagating`, wait the `Retry-After` seconds
and send the same request again — do not re-enroll.

Clone requires the `speak:clone` scope. Signup keys and sandbox mint keys both
include it. The [Clone guide](/guides/voice-cloning) covers clip quality,
polling, Omni use, and why a clip gets rejected.

### Expressive voice rendering (Cast)

Cast lives only under `/v1/cast/*` and requires `cast:render`. Preview one line
without a console project:

```bash theme={null}
curl https://api.pyai.com/v1/cast/speech \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "We made it. Now let us show them what comes next.",
    "voice": "stock_dorit_en_us",
    "emotion": "warm",
    "intensity": "balanced"
  }' \
  --output preview.wav
```

Success is a WAV file. If the voice is not on Cast, you get
`422 unsupported_voice`. Refresh `GET /v1/cast/capabilities` and pick an id
from `voices[]`. The [Cast guide](/guides/cast-overview) covers direction,
capabilities, and long-form render jobs.

### Answering-machine detection (AMD)

Already on Twilio? Point a Media Stream at `wss://api.pyai.com/v1/amd/stream`
and get `answered_by` (`human`, `voicemail`, `ivr`, `screening`, …) plus the
`reason` it decided, in a fraction of the dead-air dwell. Set the
operating-point dial and webhook once:

```bash theme={null}
curl -X POST https://api.pyai.com/v1/amd/config \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"aggressiveness": 0.5, "webhook_url": "https://your-app.example.com/amd"}'
```

A mid-call decision looks like this:

```json theme={null}
{
  "event": "amd",
  "call_id": "C_123",
  "answered_by": "machine",
  "answered_by_twilio": "machine_start",
  "confidence": 0.96,
  "decision_ms": 720,
  "reason": "machine phrase: 'please leave a message' at 1.2s"
}
```

Sandbox keys already include the AMD scopes. AMD records usage for answered
calls; see the [pricing page](https://pyai.com/pricing) for current terms. Full
walkthrough: [AMD guide](/guides/amd-answering-machine-detection).

### Compliance guardrails (Trace)

Trace scans eligible calls against built-in rule packs (TCPA, HIPAA, PII,
brand-voice) and your own rules, then gives each scanned call a scorecard with
plain-English findings. Turn it on for your org (or one agent) with one PUT:

```bash theme={null}
curl -X PUT https://api.pyai.com/v1/trace/config \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "enabled": true,
    "rule_packs": { "tcpa": { "enabled": true }, "pii": { "enabled": true } },
    "guardrails": { "mode": "warn", "block_pii": { "patterns": ["ssn", "credit_card"] } }
  }'
```

Modes: `warn` logs only (never blocks), `modify` redacts PII / injects
disclosures, `block` suppresses, `human_handoff` escalates. Always fail-open.
Then read your exposure dashboard and per-call evidence:

```bash theme={null}
curl "https://api.pyai.com/v1/trace/exposure?window_days=30" \
  -H "Authorization: Bearer $PYAI_API_KEY"
```

Trace needs a key with the `trace:configure` / `trace:read` scopes. It's in
beta and must be enabled for your organization. Your console shows current
terms and any promotional treatment; see the
[pricing page](https://pyai.com/pricing). Full walkthrough:
[Trace guide](/guides/trace-guardrails).

### Call summaries and actions (Recap)

Sandbox keys include Recap scopes and mint with Recap enabled. Submit
speaker-labelled utterances, then read the typed record:

```bash theme={null}
curl -X POST https://api.pyai.com/v1/recap/calls/call_481 \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "call_direction": "outbound",
    "customer_name": "Acme",
    "utterances": [
      { "speaker_role": "agent", "text": "I will send the pricing sheet today.", "offset_s": 12.4, "duration_s": 3.1 },
      { "speaker_role": "customer", "text": "Please include the annual option.", "offset_s": 16.1, "duration_s": 2.3 }
    ]
  }'

curl https://api.pyai.com/v1/recap/calls/call_481 \
  -H "Authorization: Bearer $PYAI_API_KEY"
```

`POST` returns `202` with `status: "pending"`. Poll `GET` until `status` is
`complete`. A finished `record` looks like this:

```json theme={null}
{
  "object": "recap.call",
  "call_id": "call_481",
  "status": "complete",
  "headline": "Acme requested annual pricing and agreed to review it this week.",
  "record": {
    "format": "recap.record.v1",
    "tldr": "Acme requested annual pricing and agreed to review it this week.",
    "summary": "The call focused on pricing structure and the next review step.",
    "action_items": [{ "owner": "agent", "task": "Send annual pricing", "due": "today" }],
    "next_steps": "Email the annual option.",
    "talk_ratio": { "agent": 0.55, "customer": 0.45 },
    "signals": [],
    "fields": {}
  }
}
```

Full walkthrough: [Recap guide](/guides/recap-call-intelligence).

<Tip>
  Prefer the official [SDKs](/guides/sdks). They handle auth, retries,
  idempotency, and realtime: `npm install @pyai/sdk` or `pip install pyai-sdk`.
</Tip>

## Next steps

<CardGroup cols={2}>
  <Card title="Use-case build guides" href="/use-cases/overview">Build your own Gong, a voice dictation app, an AI receptionist, or call-center QA, step by step.</Card>
  <Card title="Build a browser voice agent" href="/guides/browser-voice-agent">Mic → Omni → speakers in \~10 minutes, all client-side.</Card>
  <Card title="Agent greeting messages" href="/guides/agent-greeting">Set an opening line on an agent profile in the console, spoken at turn 0.</Card>
  <Card title="Omni tools & function calling" href="/guides/omni-tools">Give your agent real actions: order lookups, bookings, transfers.</Card>
  <Card title="AMD guide" href="/guides/amd-answering-machine-detection">Twilio Media Streams drop-in, tuning, and webhooks.</Card>
  <Card title="SDKs" href="/guides/sdks">Official Python and TypeScript clients, plus LiveKit and Pipecat.</Card>
  <Card title="Authentication" href="/authentication">Keys, environments, rotation, revocation.</Card>
  <Card title="Production readiness" href="/production-readiness">Regions, languages, live-key billing, reconnect behavior, limits, and retention.</Card>
  <Card title="Pricing" href="https://pyai.com/pricing">Current rates, included usage, and plan availability.</Card>
  <Card title="Errors & limits" href="/errors-and-limits">Error codes, rate limits, idempotency.</Card>
  <Card title="API reference" href="/api-reference">Full request/response schemas, right here in the docs.</Card>
  <Card title="Runnable examples" href="https://github.com/atomsai/pyai-examples">Copy-paste apps, OpenAI drop-in, voice cloning, telephony, call analytics. `npm create pyai-app@latest`.</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.