Release Notes: Text To Speech#
This page tracks releases of the Text To Speech microservice. The most recent release is listed first; older entries are preserved for history.
v1.1.0#
Release Date: September 9, 2026
New:
Named voice selection through the
voiceparameter, replacing speaker-index based selection with human-readable voice identifiers.Voice discovery API (
GET /v1/audio/voices), including available voices and descriptions.Configurable default voice support via
models.tts.default_speaker.Speaking-style instructions for Qwen3-TTS through the
instructionsrequest field.
Improved:
Synthesis performance and speech quality: faster generation with more natural prosody across supported voices.
OpenAI API compatibility retained for supported request fields and voice handling.
Known Issues:
English-only synthesis; unsupported languages return HTTP
400.The
modelrequest parameter is accepted for API compatibility but the configured service model is always used.Unknown voice names return HTTP
400.
v1.0.0#
Release Date: June 2026
Initial release of the Text To Speech microservice: an OpenAI-API-compatible speech synthesis service with multi-runtime support and selectable models, built for edge deployment on Intel® hardware.
New:
OpenAI-compatible speech endpoint (
POST /v1/audio/speech) returning either rawaudio/wavor a JSON envelope with metadata and a base64-encoded WAV payload.Voice and model metadata endpoint (
GET /v1/audio/voices) for client discovery of available speakers.Multi-runtime TTS backends:
openvino(Intel®-optimized) andpytorch.Supported models: SpeechT5 (
microsoft/speecht5_tts) and Qwen3-TTS (Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice) withcustom_voiceandvoice_designvariants.Configurable device (
CPU,GPU) and precision (int8,int4,fp16,fp32) where supported by the runtime/model.Optional persistence of synthesized output to
storage/<session_id>/withX-Session-IDreturned in the response headers.Health endpoint (
GET /health) for readiness probes.Models are warm-loaded once per process and reused across requests to keep per-request synthesis latency low.
OpenVINO™ acceleration on Intel® CPUs and integrated/discrete GPUs.
Single
config.yamlshared by standalone and container runs, with env overrides viaTEXT_TO_SPEECH__....Docker Compose deployment exposing the API on port
8011; standalone Python mode binds127.0.0.1:8011on the host.Container runs as a non-root user (UID 1000).
Known issues:
English-only synthesis. Requests with any other language are rejected with HTTP
400.The
modelrequest field is accepted for OpenAI API compatibility but is ignored; the service always uses the model defined inconfig.yaml.For SpeechT5,
languageis accepted but ignored (English only). Thevoicefield selects one of the seven bundled speaker embeddings; an unknown name returns HTTP400.Compatibility with the Video Search and Summarization sample application will be added in a subsequent release.