> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pyai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Speech To Text (Hear)

> Phone-tuned multilingual speech-to-text with about 200 ms first partials in-region, an OpenAI-compatible drop-in, and optional cleanup on English finals.

Hear is speech-to-text for real phone audio. First partials land in about 200 ms
in-region (revisable, early platform measurement). English finals can opt into cleanup: digits,
punctuation, spoken dictation, filler drop, and custom vocabulary.
Partials stay raw so the live UI stays fast.

Three surfaces. Pick by when the audio exists and what the result must contain:

| You need | Endpoint | Result | Scope |
| - | - | - | - |
| Plain text from one file now | `POST /v1/audio/transcriptions` | Synchronous text response | `hear:transcribe` |
| Partials and finals while someone speaks | `GET /v1/audio/transcriptions/stream` | WebSocket events | `hear:stream` |
| Timestamps, speakers, SRT/VTT, polling, or webhooks | `POST /v1/transcription/jobs` | Async job with word/segment offsets | `hear:transcribe` + `transcribe:jobs` |

<Warning>
  Hear synchronous and streaming transcription supports `en`, `es`, `fr`, `de`,
  `hi`, `it`, `pt`, and `nl`. Omit `language` for automatic detection, or send a
  published code to pin recognition. Streaming also accepts `language=auto`.
  Async jobs transcribe the same eight languages with the spoken language
  auto-detected per call; the detected language code is not returned in the
  result.
</Warning>

## Speaker diarization for recordings

Hear async jobs support [speaker diarization](/guides/speaker-diarization): use `diarize: true` for mono recordings, or `channel: true` when each participant is on a separate stereo channel. Read the returned mode before relying on speaker labels. This is a recording workflow, separate from synchronous and streaming transcription.

## Transcribe one file

Use the synchronous endpoint when the file is available now and your request can
wait for a text result.

```bash theme={null}
pip install pyai-sdk
# or
npm install @pyai/sdk
```

<CodeGroup>
  ```bash curl theme={null}
  curl https://api.pyai.com/v1/audio/transcriptions \
    -H "Authorization: Bearer $PYAI_API_KEY" \
    -F file=@audio.wav \
    -F model=pyai-hear \
    -F language=en
  ```

  ```python Python theme={null}
  import os
  from pyai import PyAI

  pyai = PyAI(api_key=os.environ["PYAI_API_KEY"])
  result = pyai.audio.transcriptions.create(
      file=open("audio.wav", "rb"),
      model="pyai-hear",
      language="en",
  )
  print(result["text"])
  ```

  ```ts Node theme={null}
  import PyAI from "@pyai/sdk";
  import { readFile } from "node:fs/promises";

  const pyai = new PyAI({ apiKey: process.env.PYAI_API_KEY! });
  const result = await pyai.audio.transcriptions.create({
    file: new Blob([await readFile("audio.wav")]),
    filename: "audio.wav",
    model: "pyai-hear",
    language: "en",
  });
  console.log(result.text);
  ```
</CodeGroup>

Expected response:

```json theme={null}
{ "text": "Your appointment is confirmed for Thursday at two." }
```

Supported file types are WAV, MP3, M4A, FLAC, and OGG. This compatibility
surface is best for a simple migration or a short file when you do not need a
durable job, speaker labels, subtitle artifacts, or contracted word offsets.

## Stream live speech

Use the WebSocket when a person is waiting for words to appear:

```text theme={null}
wss://api.pyai.com/v1/audio/transcriptions/stream
  ?protocol=pyai-hear-v1
  &language=auto
  &sample_rate=16000
  &encoding=pcm16
  &interim_results=true
```

Send mono little-endian PCM16. For 8 kHz telephony PCM, set
`sample_rate=8000`; Hear converts it before recognition and endpointing.
Other sample rates and compressed encodings receive `400 unsupported_audio_format`.

Render `partial` as revisable UI. Commit `final` as the corrected transcript.
Send `{"type":"commit"}` when your application must force the current utterance
to finish, and keep streaming silence during normal pauses so server
endpointing can advance.

The [streaming guide](/guides/streaming-stt) contains browser capture, framing,
event handling, endpointing, reconnect behavior, and a runnable example.

## Process a finished recording

Use async jobs when you need any of these:

* Word- or segment-level timestamps.
* Stereo channel separation or mono diarization. If diarization cannot be
  produced for a recording, the job completes with mode `transcript` and no
  speaker labels — check the `mode` field.
* SRT or VTT output.
* A result that survives the submission request.
* Polling or a signed completion webhook.
* Idempotent URL-based submission.

The [timestamped jobs guide](/guides/async-transcription-jobs) defines the
response schema, decoded source-media timeline behavior, speaker labels, size limits,
webhook signature, and retention.

## Format final transcripts

All three surfaces support final-text presentation options:

* `numerals`: digits or spoken-number form.
* `smart_format`: sentence capitalization and punctuation.
* `dictation`: spoken punctuation commands.
* `drop_fillers`: optional filled-pause removal.
* `vocabulary`: request-level or explicitly enabled stored terms for known vocabulary.

Formatting can change the relationship between text and word timestamps. Read
[Format Hear transcripts](/guides/hear-transcript-formatting) before enabling
token-changing options on subtitles, legal audio, or compliance evidence.

## Phone audio

Hear streaming accepts PCM16 or Opus. Decode an 8 kHz phone codec before sending
PCM16, and do not confuse companding with resampling. The
[telephony audio reference](/reference/telephony-audio) gives the exact G.711
and sample-rate path.

## OpenAI compatibility

`POST /v1/audio/transcriptions` is OpenAI-shaped, so an existing OpenAI client
needs only a new `base_url` and key. OpenAI model names are accepted as aliases
for `pyai-hear`: `whisper-1`, `gpt-4o-transcribe`, and `gpt-4o-mini-transcribe`.
An alias is a name and nothing else — the request is always served by Hear, so
the alias never changes the transcript you get.

<CardGroup cols={2}>
  <Card title="Stream live speech" href="/guides/streaming-stt">Partials, finals, endpointing, audio framing, and browser code.</Card>
  <Card title="Transcribe recordings" href="/guides/async-transcription-jobs">Timestamps, speakers, subtitles, jobs, and webhooks.</Card>
  <Card title="Format transcripts" href="/guides/hear-transcript-formatting">Numerals, punctuation, dictation, fillers, and vocabulary.</Card>
  <Card title="Language support" href="/reference/language-support">The Hear language matrix, quality status, and unsupported-language behavior.</Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.