allmodels.ioDocs
API ReferenceElevenLabs SDK

ElevenLabs SDK-compatible streaming-input TTS (WebSocket)

Streams text into TTS with the ElevenLabs WebSocket protocol. Use it to start generating audio before all of the text is available. The request must include `Upgrade: websocket`. **Client → server** frames (JSON): - `{ "text": " ", "voice_settings": { "speed": 1.0 }, "generation_config": {...} }` — optional init frame (single space). See the `VoiceSettings` schema for the supported keys; `speed` is honored across providers. - `{ "text": "Hello ", "try_trigger_generation": true }` — append text. - `{ "text": "world.", "flush": true }` — append and force generation. - `{ "text": "", "flush": true }` — flush without closing. - `{ "text": "" }` — close the stream. **Server → client** frames (JSON): - `{ "audio": "<base64 bytes>" }` — one frame per audio chunk. - `{ "audio": null, "error": "<code>", "code": "<status>" }` — error. Text is buffered by default to improve prosody across message boundaries. Two ElevenLabs-only options control buffering: - `auto_mode` (query param) — disables buffering; audio is generated per message rather than across the buffer. - `generation_config.chunk_length_schedule` (init-frame field) — tunes the buffer size, e.g. the init frame `{ "text": " ", "generation_config": { "chunk_length_schedule": [50] } }`. The `try_trigger_generation` and `flush` fields are supported by ElevenLabs, Deepgram, Grok, and Fish. MiniMax does not support mid-stream flushing. Deepgram WebSocket TTS supports only `pcm_*` and `ulaw_*`. Requesting `mp3_*` returns `400 invalid_output_format`. Use `provider_options` for provider-specific settings. You can also pass supported options directly as query parameters. If both forms set the same option, `provider_options[<name>]=<value>` takes precedence. Recognized enum and boolean option values are case-insensitive and are normalized to each provider's wire spelling; free-form values such as prompts, keyterms, and voice IDs retain their original case. The message-level protocol is documented at [the realtime WebSocket reference](/realtime-reference) (AsyncAPI spec: [/asyncapi.yaml](/asyncapi.yaml)).

GET/el/v1/text-to-speech/{voice_id}/stream-input

Streams text into TTS with the ElevenLabs WebSocket protocol. Use it to start generating audio before all of the text is available. The request must include Upgrade: websocket.

Client → server frames (JSON):

  • { "text": " ", "voice_settings": { "speed": 1.0 }, "generation_config": {...} } — optional init frame (single space). See the VoiceSettings schema for the supported keys; speed is honored across providers.
  • { "text": "Hello ", "try_trigger_generation": true } — append text.
  • { "text": "world.", "flush": true } — append and force generation.
  • { "text": "", "flush": true } — flush without closing.
  • { "text": "" } — close the stream.

Server → client frames (JSON):

  • { "audio": "<base64 bytes>" } — one frame per audio chunk.
  • { "audio": null, "error": "<code>", "code": "<status>" } — error.

Text is buffered by default to improve prosody across message boundaries. Two ElevenLabs-only options control buffering:

  • auto_mode (query param) — disables buffering; audio is generated per message rather than across the buffer.
  • generation_config.chunk_length_schedule (init-frame field) — tunes the buffer size, e.g. the init frame { "text": " ", "generation_config": { "chunk_length_schedule": [50] } }.

The try_trigger_generation and flush fields are supported by ElevenLabs, Deepgram, Grok, and Fish. MiniMax does not support mid-stream flushing.

Deepgram WebSocket TTS supports only pcm_* and ulaw_*. Requesting mp3_* returns 400 invalid_output_format.

Use provider_options for provider-specific settings. You can also pass supported options directly as query parameters. If both forms set the same option, provider_options[<name>]=<value> takes precedence. Recognized enum and boolean option values are case-insensitive and are normalized to each provider's wire spelling; free-form values such as prompts, keyterms, and voice IDs retain their original case.

The message-level protocol is documented at the realtime WebSocket reference (AsyncAPI spec: /asyncapi.yaml).

Authorization

xi-api-key<token>

Tenant key (ElevenLabs SDK convention).

In: header

Path Parameters

voice_id*string

Provider voice id (path segment). For Grok TTS, eve is the default voice.

Query Parameters

model_id?string

Model id: bare name or {author}/{modelName} slug (elevenlabs/eleven_flash_v2_5).

provider_order?array<string>

Comma-separated or repeated provider IDs in preferred order. Providers not listed remain eligible.

provider_only?array<string>

Comma-separated or repeated provider IDs. Only these providers may serve the request.

provider_ignore?array<string>

Comma-separated or repeated provider IDs that must not serve the request.

allow_fallbacks?boolean

Set to false to prevent fallback to another provider. Automatic provider retries are not currently supported.

output_format?string

Provider output format (case-insensitive), e.g. mp3_44100_128, pcm_16000, or ulaw_8000. Deepgram WS TTS only supports pcm_/ulaw_.

inactivity_timeout?string

Seconds of client inactivity before the ElevenLabs connection is closed.

auto_mode?string

ElevenLabs only. Generate audio for each message instead of buffering text across messages.

enable_logging?string

ElevenLabs only. Set to false to disable provider request logging.

language_code?string

ISO language code for speech generation.

apply_text_normalization?string

ElevenLabs text-normalization mode: auto, on, or off.

enable_ssml_parsing?string

ElevenLabs only. Parse SSML tags in the input text.

provider_options?|||||||||||

Provider-specific options in bracket notation, such as provider_options[encoding]=linear16. You can also pass supported options directly as query parameters. If both forms set the same option, the bracketed value takes precedence.

Header Parameters

Upgrade*string

Must be websocket to perform the protocol upgrade.

Response Body

application/json

application/json

application/json

application/json

application/json

Client
Language
import WebSocket from "ws";const ws = new WebSocket("wss://api.allmodels.io/el/v1/text-to-speech/eve/stream-input?model_id=grok/grok-tts&output_format=mp3_44100_128", {  headers: { Authorization: `Bearer ${process.env.ALLMODELS_API_KEY}` }});ws.on("open", () => {  ws.send(JSON.stringify({ text: "Hello from AllModels", flush: true }));});ws.on("message", (data, isBinary) => {  if (isBinary) process.stdout.write(data);  else console.log(JSON.parse(data.toString()));});ws.on("error", console.error);
Empty

{  "detail": {    "message": "Deepgram TTS WS only supports pcm_*/ulaw_* output formats.",    "status": "invalid_output_format",    "provider": "deepgram",    "output_format": "mp3_44100_128"  }}

{  "error": "missing_api_key"}

{  "detail": {    "loc": [      "body",      "text"    ],    "msg": "field required",    "type": "value_error.missing"  }}

{  "detail": {    "message": "websocket upgrade required",    "status": "websocket_upgrade_required"  }}
{  "detail": {    "message": "authentication was not initialized",    "status": "auth_not_initialized"  }}