> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Text To Speech (Speak)

> Choose a voice, synthesize streaming or buffered audio, select an output format, and move cleanly into telephony.

Speak turns text into audio with one OpenAI-compatible endpoint:

```text theme={null}
POST /v1/audio/speech
```

The canonical model is `pyai-speak`; the required scope is
`speak:synthesize`.

## Speaker diarization and speech generation

Speak generates audio from text. To identify who spoke when in an existing recording, use [speaker diarization with Hear](/guides/speaker-diarization). Hear async jobs produce speaker-labelled transcripts; your application can then use Speak for spoken output.

## Synthesize your first clip

Install an SDK if you want retries handled for you:

```bash theme={null}
pip install pyai-sdk
# or
npm install @pyai/sdk
```

<CodeGroup>
  ```bash curl theme={null}
  curl https://api.pyai.com/v1/audio/speech \
    -H "Authorization: Bearer $PYAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "pyai-speak",
      "input": "Your appointment is confirmed for Thursday at two.",
      "voice": "stock_dorit_en_us",
      "response_format": "wav"
    }' \
    --output confirmation.wav
  ```

  ```python Python theme={null}
  import os
  from pyai import PyAI

  pyai = PyAI(api_key=os.environ["PYAI_API_KEY"])
  audio = pyai.audio.speech(
      input="Your appointment is confirmed for Thursday at two.",
      voice="stock_dorit_en_us",
      response_format="wav",
  )
  open("confirmation.wav", "wb").write(audio)
  ```

  ```ts Node theme={null}
  import PyAI from "@pyai/sdk";
  import { writeFile } from "node:fs/promises";

  const pyai = new PyAI({ apiKey: process.env.PYAI_API_KEY! });
  const audio = await pyai.audio.speech({
    input: "Your appointment is confirmed for Thursday at two.",
    voice: "stock_dorit_en_us",
    response_format: "wav",
  });
  await writeFile("confirmation.wav", Buffer.from(audio));
  ```
</CodeGroup>

Success is a WAV file at the voice's native 24 kHz. Audio bytes are delivered
incrementally by default.

## Choose a voice from the catalog

Do not hardcode assumptions about voice language, product support, or delivery
mode. Read the catalog:

```bash theme={null}
curl "https://api.pyai.com/v1/voices?language=en&source=stock" \
  -H "Authorization: Bearer $PYAI_API_KEY"
```

For each row:

* `voice_id` is the canonical input.
* `aliases` are permanent convenience inputs on the advertised surfaces.
* `available_on` tells you whether the voice works on Speak, Omni, or both.
* `synthesis_modes` tells you which `stream` values the voice accepts. Every
  Speak voice accepts both; an empty list means the voice has no Speak surface
  at all.
* `speak_emotions` lists qualified emotion directions for this voice. An empty
  or absent list means directed emotion is unavailable.
* `tier` and `pricing` describe the customer-facing quality tier and its
  catalog treatment. Check the [pricing page](https://pyai.com/pricing) for
  current commercial terms.

Designed voices appear in the same catalog with `source: "design"`. Cloned
voices are managed separately under `/v1/voice/clones`.

## Natural voices

Natural is a selectable customer-facing voice tier, not a different endpoint:

* English `en1` (canonical id `stock_aria_en`) works on Speak in both streaming
  and buffered modes. The same id is available on Omni. It is the only
  Natural-tier voice in the catalog today; the former Hindi Natural aliases
  `hi5`–`hi8` are retired. Hindi is served by the Standard-tier voices, which
  work on both Speak and Omni.
* Check the [pricing page](https://pyai.com/pricing) for current voice-tier
  treatment.

Always inspect the live catalog. `available_on` is the product contract;
`synthesis_modes` is the delivery contract.

## Emotion direction

Emotion support is specific to each voice. Read `speak_emotions` from
`GET /v1/voices` before sending a direction. When the chosen voice advertises it,
pass `emotion` with `happy`, `sad`, `angry`, `fearful`, or `surprised`. This works
with both streaming and buffered delivery. Omit the field, or use `neutral`,
for the voice's ordinary delivery.

```ts theme={null}
const { data: voices } = await pyai.voices.list();
const voice = voices.find((voice) => voice.speak_emotions?.includes("happy"));
if (!voice) throw new Error("No voice currently advertises happy delivery.");
const audio = await pyai.audio.speechStream({
  input: "Your appointment is confirmed for Thursday at two.",
  voice: voice.voice_id,
  emotion: "happy",
});
```

A successful directed response includes `x-pyai-emotion` with the requested
name and `x-pyai-emotion-applied: true`. Unsupported voice/direction combinations
return `400 unsupported_parameter`; invalid direction names return
`400 invalid_emotion`. If a qualified direction is temporarily unavailable,
Speak returns `503 emotion_unavailable` before returning audio.

Speak does not offer an intensity control. Direction describes delivery;
keep pacing and check-in questions in the script itself.

## Streaming vs buffered delivery

`stream` defaults to `true`:

* Use `true` when playback should start as bytes arrive.
* Use `false` when your client requires a complete body and
  `Content-Length`.
  Every catalog voice accepts both values, so the documented request above works
  for any `voice_id` you find in `GET /v1/voices`. Some voices (today: Hindi and
  a few `es` / `fr` / `de` rows) are served from a fleet that has no streaming
  lane. Those are rendered on the blocking lane whichever value you send, and the
  response says so with `x-pyai-stream: buffered`: same status, same body, same
  `response_format`, higher time-to-first-byte. Nothing to retry, nothing to
  branch on unless you are measuring TTFB.

### Consume bytes as they arrive

An `ArrayBuffer` read followed by playback demonstrates buffered application
behavior even when the server is incremental. Consume the response body and pass
each chunk to your audio pipeline as it arrives. Abort a synthesis when the
caller interrupts, and do not retry after any audible bytes without a policy for
repeated speech.

```ts Node streaming consumption theme={null}
const controller = new AbortController();
const response = await fetch("https://api.pyai.com/v1/audio/speech", {
  method: "POST",
  signal: controller.signal,
  headers: {
    Authorization: `Bearer ${process.env.PYAI_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "pyai-speak",
    input: "Your appointment is confirmed for Thursday at two.",
    voice: "stock_dorit_en_us",
    response_format: "pcm",
    stream: true,
  }),
});

if (!response.ok || !response.body) throw new Error(`Speak failed: ${response.status}`);
console.info("effective delivery", response.headers.get("x-pyai-stream") ?? "incremental");

const reader = response.body.getReader();
try {
  for (;;) {
    const { done, value } = await reader.read();
    if (done) break;
    enqueuePcmForPlayback(value); // application-owned queue; returns immediately
  }
} catch (error) {
  // Some audio may already be audible. Record the partial result and use a
  // product-specific recovery policy instead of blindly retrying the sentence.
  handlePartialSynthesisFailure(error);
}

// On interruption: controller.abort(); clearQueuedAudio();
```

```bash theme={null}
curl https://api.pyai.com/v1/audio/speech \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "This response is returned as one complete file.",
    "voice": "stock_dorit_en_us",
    "response_format": "mp3",
    "stream": false
  }' \
  --output complete.mp3
```

## Input length

`input` is capped at **2000 characters** per request. Longer text answers
`400 input_too_long`, naming the limit and the length you sent — split the
script and concatenate the audio, or use
[Cast](/guides/cast-overview) for long-form work, which is built for it.

## Output formats and sample rates

| Format | Output | Typical use |
| - | - | - |
| `wav` | WAV container, 24 kHz default | General playback and files |
| `mp3`, `opus`, `aac`, `flac` | Encoded audio | Storage and distribution |
| `pcm` | Raw headerless PCM16 little-endian mono | Realtime frameworks and custom pipelines |
| `g711_ulaw`, `g711_alaw` | Raw headerless G.711, fixed 8 kHz mono | Phone networks and media streams |

For `pcm`, choose `sample_rate` from 8 kHz through 48 kHz. G.711 is always
8 kHz; omit `sample_rate` or pass exactly `8000`. A conflicting G.711 sample
rate is rejected.

The [telephony audio reference](/reference/telephony-audio) explains companding,
resampling, and the exact format to send to common phone transports.

## Stock, cloned, or designed

* **Stock:** select a row from `GET /v1/voices`; no enrollment required.
* **Clone:** enroll a voice you have explicit permission to use. See
  [Clone](/guides/voice-cloning).
* **Design:** create a new synthetic voice from a text description with
  `/v1/voice/design`, preview candidates, and save one into your voice library.

Voice is biometric data when it represents a real person. Establish consent
before cloning; prompt-designed synthetic voices are a separate workflow.

Start a design job with a stable idempotency key:

```bash theme={null}
curl https://api.pyai.com/v1/voice/design \
  -H "Authorization: Bearer $PYAI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: support-voice-v1" \
  -d '{
    "prompt": "A calm, clear support voice with measured pacing.",
    "candidates": 3,
    "sample_text": "Thanks for calling. How can I help?"
  }'
```

Poll `GET /v1/voice/design/{design_id}` until candidates are ready, preview
their signed URLs, then save one with
`POST /v1/voice/design/{design_id}/save`. Voice design requires
`speak:design`; cloning uses the separate `speak:clone` scope.

## Compatibility and unsupported controls

OpenAI model names are accepted as aliases for `pyai-speak`: `tts-1`,
`tts-1-hd`, `tts-1-1106`, `tts-1-hd-1106`, and `gpt-4o-mini-tts`. So are the
preset voice names `alloy`, `echo`, `fable`, `onyx`, `nova`, and `shimmer`. An
alias is a name and nothing else: it never changes the engine, the lane, or the
voice you would otherwise get. New integrations should prefer `pyai-speak` and
canonical catalog voice IDs.

`speed`, `seed`, and `temperature` are not active Speak controls. Sending them
returns `400 unsupported_parameter`; do not assume they were silently applied.

<CardGroup cols={2}>
  <Card title="Browse voices" href="/api-reference">Filter the live catalog and inspect surface/mode support.</Card>
  <Card title="Clone" href="/guides/voice-cloning">Enroll, test, use, and delete a consented voice.</Card>
  <Card title="Telephony audio" href="/reference/telephony-audio">PCM and G.711 formats, rates, and resampling.</Card>
  <Card title="Language support" href="/reference/language-support">Choose a voice that matches the required language and product.</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.