Configuration#
kiosk-core and kiosk-ui are configured through environment variables
(see Environment Variables).
The three model-hosting services (audio-analyzer, text-to-speech,
rag-service) are configured through YAML files that the kiosk pins
and mounts into the containers. The most common changes are the
model and the inference device.
Model Selection#
Each model-hosting service reads the model identifier from the same pinned config file used for device selection:
Service |
File |
Model fields |
|---|---|---|
|
|
|
|
|
|
|
|
Use Hugging Face IDs where the field name is hf_id. Models are
downloaded and exported on first start into the per-service models/
directory; subsequent starts reuse the cache.
Supported / validated models#
The kiosk ships with the following defaults. These are the models the stack has been validated with — they are the recommended starting point. The Devices column lists the supported inference devices for each:
Service |
Field |
Default (validated) |
Other examples |
Devices |
|---|---|---|---|---|
|
|
|
|
|
|
|
|
other SpeechBrain emotion-recognition models |
|
|
|
|
|
|
|
|
|
other OpenVINO-exportable instruct LLMs |
|
|
|
|
|
|
|
|
|
|
|
[!IMPORTANT] Changing models is at your own discretion. The defaults above are the only combinations validated with this stack. Configuring models, variants, devices, or precisions other than the defaults may negatively affect the functionality, accuracy, latency, or stability of the application. You are responsible for ensuring the configuration you choose is correct and works for your use case — make changes only if you understand the implications.
In particular:
Some models do not function properly at aggressive quantization. If a model produces garbled, empty, or low-quality output at
int4, switch that model’sweight_format/dtypetoint8orfp16.A model must be exportable to OpenVINO IR for the OpenVINO backend; not every Hugging Face model is supported.
Larger models increase first-run download/export time, memory use, and per-request latency, and may not fit on the selected device.
After any change, restart the affected service and verify it loads and responds correctly before relying on it.
Inference Device#
Each model-hosting service reads its device from a pinned config file:
Service |
File |
Fields |
|---|---|---|
|
|
|
|
|
|
|
|
The supported devices for each model are listed in the Supported / validated models table above.
Use uppercase device names (CPU, GPU, and — for audio-analyzer ASR and
queue-service only — NPU). rag-service expects them as quoted strings;
audio-analyzer and text-to-speech unquoted.
[!IMPORTANT]
text-to-speechdoes not supportNPU.models.tts.deviceonly acceptsCPU/GPU(seeconfigs/text-to-speech/config.yaml); there is no NPU device mapping for this service indocker-compose.yml. Do not setmodels.tts.device: NPU— it is not a supported configuration.
After editing, restart the affected service and confirm OpenVINO picked the device:
docker compose up -d --build --force-recreate <service-name>
docker compose logs <service-name> | grep -i -E "device|compiling|GPU|CPU"
OpenVINO prints a Compiling model on <DEVICE> line on first load.
GPU execution is delegated to the OpenVINO backend used by each service. Whether a given model actually runs on GPU and how it performs depends on the OpenVINO version and operator coverage for that model.
Audio Analyzer ASR Provider/Device (config.yaml)#
Configure ASR in configs/audio-analyzer/config.yaml (single source of truth):
models.asr.providermodels.asr.device
make check-env validates this before startup and rejects unavailable
hardware early. For NPU, ACCEL_MOUNT_PATH must be set manually (in
.env or the shell) to the host NPU device node (/dev/accel/accel0)
before running make up or docker compose up — neither make up nor
make check-env auto-detects it. The checked-in docker-compose.yml
defaults ACCEL_MOUNT_PATH to /dev/null so CPU/GPU-only hosts stay
unaffected.
ASR on NPU: whisper-tiny/whisper-base only#
[!IMPORTANT] NPU works for ASR only with
models.asr.name: whisper-tinyorwhisper-base.whisper-small/medium/largefail to compile on NPU withCheck '!self_attn_nodes.empty()' failed.Why: OpenVINO NPU only supports static-shape models (see OpenVINO NPU docs). Whisper’s IR has dynamic shapes, so OpenVINO GenAI’s NPUW plugin pattern-matches attention blocks to make it static — a heuristic that succeeds for
tiny/baseand fails for larger models. Not fixable via config; re-test if you upgrade OpenVINO/GenAI or the audio-analyzer image.
models:
asr:
provider: openvino
device: NPU
name: whisper-base # or whisper-tiny — whisper-small+ fails to compile
weight_format: null
Other ASR Configurations#
# OpenAI + CPU
models: { asr: { provider: openai, device: CPU } }
# OpenVINO + GPU (any model size)
models: { asr: { provider: openvino, device: GPU } }
# OpenVINO + CPU (any model size)
models: { asr: { provider: openvino, device: CPU } }
openai supports CPU only — openai + GPU and openai + NPU are not supported.
ASR Support Matrix#
Provider |
CPU |
GPU |
NPU |
|---|---|---|---|
|
Yes |
No |
No |
|
Yes |
No |
No |
|
Yes |
Yes (Intel GPU required) |
|
If GPU is configured and unavailable on the host, make check-env fails before startup — no silent fallback.
Audio Analyzer Diarization Device (config.yaml)#
Diarization (models.diarization.device) is a separate component from ASR
(see ASR Support Matrix above) with its own,
more limited device support. Do not assume ASR’s CPU/GPU/NPU support
applies to diarization — it does not.
[!IMPORTANT] In the currently released Kiosk image, diarization only supports
CPU. The diarizer (pyannote/speaker-diarization-3.1, a PyTorch/SpeechBrain model) is loaded withtorch.device(<configured value>). PyTorch has no"gpu"device string (Intel GPU support in PyTorch requires"xpu", which this component does not use) and no"npu"device string at all. Settingdevice: GPUordevice: NPUis accepted by the config schema but fails at diarizer-load time with an error like:Expected one of cpu, cuda, ipu, xpu, ... device type at start of device string: gpuThis is non-fatal: the failure is caught, logged as a warning, and diarization is disabled for that session — the container stays healthy and ASR keeps working, but speaker labels are not produced.
Use
device: CPUfor diarization. Do not configureGPUorNPUformodels.diarization.device— they do not work in this image and will silently disable diarization rather than accelerate it.
An OpenVINO-backed diarization path that genuinely supports GPU/NPU
exists in a newer upstream edge-ai-libraries audio-analyzer checkout,
but is not part of the currently released Kiosk image covered by this
document. Do not configure GPU/NPU for diarization based on that
upstream code until a Kiosk image that includes it is released.
Queue Service Device (QUEUE_DEVICE)#
queue-service runs the YOLO26 person detector through a DLStreamer
(gvainference) pipeline. The inference device is controlled by
model.device in queue-service/conf/queue-config.yaml, and can be
overridden without editing that file via QUEUE_DEVICE in .env
(mapped by docker-compose.yml to QUEUE_SERVICE__MODEL__DEVICE, read by
queue-service/src/config_loader.py).
Default:
QUEUE_DEVICE=CPU— always available, no extra device mapping needed.QUEUE_DEVICE=GPU/QUEUE_DEVICE=NPU— supported, using the same/dev/driandACCEL_MOUNT_PATH-driven/dev/accelmapping described under Audio Analyzer ASR Provider/Device above. For NPU, setACCEL_MOUNT_PATHyourself to the host NPU device node before startingqueue-service— it is not auto-detected.
Verify the configured device actually reached the pipeline:
docker logs queue-service 2>&1 | grep "gvainference model"
# Expected: ... gvainference model=... device=NPU ... (or CPU/GPU, matching QUEUE_DEVICE)
Editing
queue-config.yaml’smodel.devicedirectly also works and takes precedence in the same way as any other YAML default —QUEUE_DEVICEonly needs to be set when you want to override it without touching the file.
Identity Service Device (IDENTITY_DEVICE)#
identity-service performs face detection/re-identification
(face-detection-retail-0005, face-reidentification-retail-0095) and
voice-print embedding (ecapa-tdnn-voice) — all three are OpenVINO IR
models, loaded via openvino.Core().compile_model(model, device), and
IDENTITY_DEVICE is correctly wired end-to-end from .env through
docker-compose.yml to the OpenVINO compile call. The device-selection
code itself has no bug and no model-format limitation.
[!NOTE] Face detection/re-identification support
CPU,GPU, andNPU. Theidentity-servicecontainer has NPU device passthrough via the sameACCEL_MOUNT_PATH//dev/accelmechanism used byaudio-analyzerandqueue-service. SetIDENTITY_DEVICE=NPUin.envtogether withACCEL_MOUNT_PATHpointing to the host NPU device node — this is not auto-detected and must be set manually — to run face detection/re-identification on the NPU.Voice-print embedding (
ecapa-tdnn-voice) does not supportNPU. Its OpenVINO IR contains an internal STFT reshape with an unbounded dynamic dimension (aten::view/Reshape), which the NPU compiler rejects at compile time (Got negative shape dim bound). WithIDENTITY_DEVICE=NPU, the service starts with face engine enabled and voice engine disabled (inference_ready=false, since voice verification requires both). UseIDENTITY_DEVICE=CPUorIDENTITY_DEVICE=GPUif voice authentication is required.The face/voice model files are not downloaded by default — run
./setup_models.sh --identityfirst. Without them, the face/voice engines stay disabled (inference_ready=false) regardless of the configured device.
OVMS-LLM Device (TARGET_DEVICE)#
TARGET_DEVICE controls the inference device for the ovms-llm container
only (the LLM served by OpenVINO Model Server for the ordering agent).
rag-service’s own embedding/reranker components have their own,
independent device configuration — see
RAG Service Embedding/Reranker Device
below.
Currently supported:
TARGET_DEVICE=CPU,TARGET_DEVICE=GPU,TARGET_DEVICE=NPU.
[!NOTE]
TARGET_DEVICE=NPUdevice passthrough works forovms-llm. Theovms-llmcontainer has NPU device passthrough via the sameACCEL_MOUNT_PATH//dev/accelmechanism used byaudio-analyzerandqueue-service. SetTARGET_DEVICE=NPUin.envtogether withACCEL_MOUNT_PATHpointing to the host NPU device node — this is not auto-detected and must be set manually — to compile the LLM for the NPU. OVMS logsAvailable devices for Open VINO: CPU, GPU, NPUand the model (Qwen3-4B-int8-ov) compiles successfully.NPU inference has been observed to fail at request time on at least one validated host, even after a successful compile — with two distinct symptoms seen: a short chat-completion request failed with
zeFenceHostSynchronize result: ZE_RESULT_ERROR_UNKNOWNinside OVMS’s LLM executor, and a longer RAG-augmented prompt (routed throughrag-service) failed withInput length exceeds the maximum allowed length. Both point to NPU driver/runtime or static-shape/context-length limitations for this model’s KV-cache/stateful execution graph, not a configuration issue. Validate end-to-end generation with realistic prompt lengths (not just/v3/modelsor container health) before relying onTARGET_DEVICE=NPUforovms-llmin production.Use
TARGET_DEVICE=CPUorTARGET_DEVICE=GPUif the host does not have an NPU,ACCEL_MOUNT_PATHis not set, or NPU generation requests fail as described above.
RAG Service Embedding/Reranker Device (RAG_EMBEDDING_DEVICE, RAG_RERANKER_DEVICE)#
rag-service’s embedding (BAAI/bge-large-en-v1.5) and reranker
(BAAI/bge-reranker-base) components are OpenVINO IR models exported
in-process by optimum-intel (rag-service/utils/ensure_model.py) and
loaded via OVModelForFeatureExtraction/equivalent
(rag-service/components/embedding_component.py,
reranker_component.py). Their device is set independently of
TARGET_DEVICE via RAG_EMBEDDING_DEVICE / RAG_RERANKER_DEVICE in
.env (mapped by docker-compose.yml to
SMART_KIOSK_RAG__MODELS__EMBEDDING__DEVICE /
SMART_KIOSK_RAG__RETRIEVAL__RERANKER__DEVICE).
Currently supported:
RAG_EMBEDDING_DEVICE/RAG_RERANKER_DEVICE=CPUorGPU. Default:GPU.Currently unsupported:
NPU.
[!IMPORTANT]
NPUis not supported for the embedding/reranker models.optimum-intel’s default export produces OpenVINO IR with dynamic (unbounded) sequence-length and batch shapes — required because queries and knowledge-base documents vary in length and the reranker batches multiple candidates per call (rag-service/config.yaml’smodels.embedding.batch_size). The NPU compiler rejects this IR (Missing upper bound for one or more nodes); setting either variable toNPUwill crashrag-serviceat startup.Forcing static/bounded shapes to work around this is not recommended: it would require padding every input to a fixed max length (wasting compute on short queries) and serializing what is currently a batched reranker call into one NPU invocation per candidate — likely slower overall than
GPU/CPU, for a component that is not the latency bottleneck (the LLM is).CPUis normally fast enough for embedding/reranking;GPUis the default to match prior behavior.
Environment Variables#
kiosk-core has no config file. All settings are controlled through environment variables.
kiosk-core API (main:app)#
Variable |
Default |
Description |
|---|---|---|
|
|
audio-analyzer transcription endpoint |
|
|
RAG query endpoint |
|
|
TTS speech synthesis endpoint |
|
|
Model name sent to the TTS service |
|
(unset) |
Voice name sent to the TTS service |
|
|
Language sent to the TTS service |
|
(unset) |
Optional style instructions for TTS |
|
|
Default audio sample rate in Hz |
|
|
Length of each audio chunk sent to audio-analyzer |
|
|
Silence duration after speech that ends a session |
|
|
Hard cap on session duration |
|
|
RMS threshold below which audio is treated as silence |
|
|
PortAudio capture block size |
|
|
Audio buffered before speech starts |
|
|
HTTP client timeout for downstream calls |
Gradio UI (gradio_app.py)#
Variable |
Default |
Description |
|---|---|---|
|
|
Base URL of the kiosk-core API |
|
|
Passed to start-file sessions as |
|
|
Passed to start-file sessions as |
|
|
Passed to start-file sessions as |
|
|
HTTP client timeout in the UI |
|
|
How often the UI polls for session state updates |
Kiosk UI runtime mode {#kiosk_ui_mode}#
The React kiosk UI (kiosk-ui/) ships as a single image that can serve
either of two screens, selected at container start — no rebuild:
Variable |
Default |
Description |
|---|---|---|
|
|
|
The value is written to /usr/share/nginx/html/config.js by
docker-entrypoint.sh (installed as an nginx docker-entrypoint.d
script) and read by the SPA before the React bundle loads. In
docker-compose.yml the two screens are separate containers
(kiosk-ui and kiosk-ui-customer) built from the same image/context,
published on different host ports (7860 and 7861 respectively) so
they can be shown on two separate monitors during a demo.
Compose Defaults#
When running with the top-level docker-compose.yml, the defaults are wired to the internal Compose network:
KIOSK_CORE_ANALYZER_URL=http://audio-analyzer:8010/v1/audio/transcriptionsKIOSK_CORE_RAG_URL=http://rag-service:8020/api/v1/queryKIOSK_CORE_TTS_URL=http://text-to-speech:8011/v1/audio/speechKIOSK_CORE_UI_BASE_URL=http://kiosk-core:8012
Most deployments should leave these values unchanged. Override them only when kiosk-core or kiosk-ui must call services outside the local Compose stack.
Session Parameters#
Session parameters (chunk duration, silence threshold, etc.) can also be provided per-request in the POST body for /api/v1/sessions/start and /api/v1/sessions/start-file. Per-request values take precedence over the environment variable defaults.
NPU Deployment Workflow#
Step-by-step workflow to run this stack with Intel NPU acceleration where supported.
NPU support by component:
Component |
NPU |
|---|---|
Queue Service ( |
✅ Yes |
Identity Service face/re-id ( |
✅ Yes |
Identity Service voice/ECAPA-TDNN ( |
❌ No — dynamic shape rejected by NPU compiler |
Audio Analyzer ASR ( |
⚠️ |
Audio Analyzer Diarization ( |
❌ No — CPU only |
OVMS-LLM ( |
⚠️ Compiles, but inference fails at runtime — not production-ready |
RAG Service embedding/reranker ( |
❌ No — dynamic shape rejected by NPU compiler |
Text-to-Speech ( |
❌ No |
Do not set NPU for a ❌ component — it will fail to start or silently disable the feature.
The steps below configure NPU for queue-service; adapt the device variable for other ✅/⚠️ components.
1 — System requirements#
Requirement |
Details |
|---|---|
Hardware |
Intel Core Ultra (Meteor Lake or later) with integrated NPU |
Host driver |
Intel NPU driver ( |
User-space runtime |
|
Host device |
|
OpenVINO |
Container image already bundles the correct runtime |
Verify the NPU device node is present before proceeding:
ls /dev/accel/
# Expected: accel0 accelmon0
2 — Install the Intel NPU driver (if not already installed)#
Refer to the Intel NPU driver repository: intel/linux-npu-driver. Installation varies by distribution. After installation:
# Verify kernel driver is loaded
lsmod | grep intel_vpu
# Verify device node exists
ls -la /dev/accel/accel0
3 — Set NPU device for queue-service (and optionally audio-analyzer ASR)#
# .env
QUEUE_DEVICE=NPU
Optionally, also enable NPU for ASR (whisper-tiny/whisper-base only) in configs/audio-analyzer/config.yaml:
models:
asr:
provider: openvino
device: NPU
name: whisper-base # whisper-tiny also works; whisper-small+ fails to compile
weight_format: null
4 — Set ACCEL_MOUNT_PATH and start the stack#
ACCEL_MOUNT_PATH is not auto-detected by make up or
make check-env — export it yourself before starting the stack:
export ACCEL_MOUNT_PATH=/dev/accel/accel0
cd smart-kiosk-assistant
make check-env
make up
Alternatively, set ACCEL_MOUNT_PATH=/dev/accel/accel0 directly in
.env so it’s picked up automatically on every make up /
docker compose up without exporting it each time.
5 — Verify NPU is active#
docker ps --filter "name=queue-service" --format "{{.Names}}\t{{.Status}}"
docker exec queue-service python3 -c "import openvino as ov; print(ov.Core().available_devices)" # expect: includes NPU
docker logs queue-service 2>&1 | grep -i "device=NPU\|npu"
If ASR is on NPU (whisper-tiny/whisper-base only):
docker logs audio-analyzer 2>&1 | grep -i "Loading Model" # expect: device=NPU
curl -s -X POST http://localhost:8010/v1/audio/transcriptions -F "file=@your_sample.wav"
If identity-service face/re-id is on NPU:
IDENTITY_DEVICE=NPU make up IDENTITY=true
docker logs identity-service 2>&1 | grep -i "face\|voice\|inference_ready"
6 — Troubleshooting#
Symptom |
Cause |
Fix |
|---|---|---|
Container unhealthy, |
NPU driver not loaded or |
Verify host driver and set |
|
NPU compiler not in image |
Rebuild the affected image with NPU user-space packages |
Slow first inference (20–60 s) |
NPU compiler cache is empty (cold start) |
Normal on first run; subsequent requests will be fast |
|
Model is |
Use |
Non-NPU containers unhealthy after NPU config change |
NPU-unrelated services picking up wrong env |
Only modify the specific component’s config (e.g. |
Cold-start note: The OpenVINO NPU compiler caches compiled kernels inside the container under
/tmp/ov_cache/. The first inference after a container restart takes significantly longer (20–60 s) while the cache warms up. This is expected behavior.