Release Notes: Video Search and Summarization Sample Application#
Upcoming updates (post 2026.1.0)#
Improved:
Search orchestration alignment for accelerator usage: Updated setup and compose behavior for Video Search so NPU device selections are preserved and propagated consistently for DataPrep and multimodal embedding services.
Simplified per-component device model (Compose): Retired the redundant
VDMS_DATAPREP_DEVICEbaseline knob. Device selection is now purely per-component —DATAPREP_EMBEDDING_DEVICE,DATAPREP_DETECTION_DEVICE, andMME_EMBEDDING_DEVICE(each defaults toCPU) — matching the Helm chart model.ENABLE_EMBEDDING_GPUis now a mode-aware embedding shortcut (sdk→DataPrep embedding GPU,api→MME embedding GPU).Helm accelerator support for search stack: Updated VSS Helm subcharts for
multimodal-embedding-msandvdms-dataprepto support NPU as an accelerator path (device-key validation, resource requests/limits, and/dev/accelmounts).Helm accelerator device permissions: Added
global.accelGroupIdsso the host gids owning/dev/dri(GPU) and/dev/accel(NPU) are injected into the podsupplementalGroups, letting the non-root container open the accelerator device (mirrors the Composegroup_addrender/video groups). Fixes NPU/GPU device initialization falling back to CPU-only.Helm OpenVINO™ model cache: Added a persistent OpenVINO™ cache (
ovCacheDir, default/app/ov_models/ov_cache) formultimodal-embedding-msandvdms-dataprep, plus a longer DataPrepstartupProbebudget, so GPU/NPU model compilation completes once and is reused across pod restarts instead of recompiling (avoids startup crash loops).Helm single-source image override:
global.registry,global.tag, andglobal.pullPolicynow apply across all VSS service images (pipeline-manager, video-ingestion, video-search, vss-ui, vdms-dataprep, multimodal-embedding-serving) from one place, with independent per-service PVCs for model/cache data.Clearer embedding error reporting: The video embedding flow now surfaces the real upstream DataPrep error instead of a misleading “Request timed out” message; only genuine timeouts (
408/504/connection aborts) are reported as timeouts.Search deployment documentation refresh: Added a dedicated Deployment Options for Video Search matrix (SDK/API with CPU/GPU/NPU combinations), including explicit
DATAPREP_EMBEDDING_DEVICE,MME_EMBEDDING_DEVICE, andDATAPREP_DETECTION_DEVICEexamples for accelerator-specific routing.Helm user-guide clarifications: Updated Helm guidance to include NPU device/key combinations and matching-device recommendations for shared PVC scheduling.
Version 2026.1.0#
Release Date: June 17, 2026
New:
EXPERIMENTAL - vLLM Intel® Arc™ Pro B-series GPU Support: Added initial experimental support for running vLLM on Intel® Arc™ Pro B-series GPUs (XPUs) via new
ENABLE_VLLM_GPUenvironment variable anddocker/compose.vllm.xpu.yamlDocker Compose overlay. This feature enables GPU-accelerated VLM captioning and LLM summarization on Intel® Arc™ Pro B-series hardware (e.g., B60, B65, B70).Note: This is an experimental feature in early stages and may require additional tuning and optimization. Performance characteristics are still being evaluated. Not recommended for production use.
Added a a new Dual UI mode with a new
--summary --searchCLI argument forsetup.shthat allows running both the summary and the search applications simultaneously at /summary and /search URI endpoints respectively.Added Dual UI support for Helm chart installations by allowing a values override file to be provided for summary and search modes simultaneously.
Improved:
Updated setup script and nginx configuration files to allow flexible UI routing for each existing mode of deployment (summary mode, search mode, Unified UI Mode) and the new Dual UI mode.
Refactored Helm chart to use a reusable
vssuisubchart with multi-mode nginx and consolidated embedding model config underglobal.embeddingModelName.Updated DLStreamer base image to
2026.1.0-ubuntu24-rc1for Video Ingestion Microservice.Setup Script: Updated the environment variable to setup embedding models. New
MULTIMODAL_EMBEDDING_MODELand the existingTEXT_EMBEDDING_MODELare used to provide embedding models in relevant modes.Docker Compose: Replaced
curlwith Pythonurllibpackage in the containerhealthcheckcommand for a lighter runtime footprint for Audio Analyzer.Docker Compose: Replaced environment variables with hard-coded mount paths. This helps in stopping containers without looking for preset variables.
Build Script: Removed Audio-Analyzer from the dependency build pipeline. A frozen version 1.3.3 will be used for the Audio Analyzer microservice for the current and all subsequent releases.
Setup Script: Removed unused environment variables and several environment variables being used as mount directories in Docker Compose files.
Version 1.3.3-rc1#
Release Date: 05 May 2026
Features:
Configurable final video summary: Added
PM_PRODUCE_FINAL_SUMMARYfeature flag to make the final LLM map-reduce video summary optional. When disabled, chunk-wise summaries are displayed chronologically instead. A per-video UI override checkbox is available in both upload flows. Audio transcript summarization is automatically skipped when the final summary is turned off.Audio transcript summarization: Added audio transcript summarization support and improved audio transcription accuracy.
OVMS-first architecture: Replaced the standalone
vlm-openvino-servingmicroservice with OpenVINO™ Model Server (OVMS) as the unified inference backend for both VLM captioning and LLM summarization. This is a breaking change; thevlm-inferencesubchart and container have been removed.Performance Optimizations (MME & VDMS-Data-Prep):
Refactored pre-processing and inference with
AsyncInferQueuebased OpenVINO™ inference and static shape model compilation for iGPU.Added ThreadPool for parallel open_clip image pre-processing with support for input tensor batching and padding for optimal OpenVINO™ inference paths.
Introduced PyAV-based video decode abstraction supporting keyframes and uniform sampled frames extraction with producer-consumer pattern for parallel decode and frame translation to PIL.
Enabled multiple/parallel decoder instances for file, RTSP stream, and bytes input sources.
Implemented frame batching for pipelined pre-processing and inference with integrated PyAV decoder in VDMS data-prep.
Search Timeout and Resource Management: Added
SEARCH_DATAPREP_TIMEOUT_MSconfiguration to prevent VSS-UI timing out during embedding creation. Added ulimit constraints with soft and hard limits to enable shared memory creation and define memory block allocation boundaries.
HW used for validation:
Intel® Xeon® 5 + Intel® Arc™ B580 GPU
Vanilla Kubernetes Cluster
Known Issues/Limitations:
This release includes only limited testing on EMT‑S and EMT‑D, some behaviors may not yet be fully validated across all scenarios.
HW sizing of the Video Search or Video Summarization pipeline is in progress. Optimization of the pipelines will follow HW sizing.
Known issues are internally tracked. Reference not provided here.
how-to-performancedocument is not updated yet. HW sizing details will be added to this section shortly.NPU support with OVMS is added as experimental feature and may not work for all models or configurations.
Version 1.3.2#
Release Date: 17 Feb 2026
Features:
In VSS search mode, users can now filter results by time range via:
Query parsing to infer time ranges (e.g., “person seen in last 5 minutes”).
Direct time range input from the UI.
Added live system performance metrics in the search UI (enable with
export ENABLE_VSS_COLLECTOR=true).Fixed the build script of the
vdms-dataprepmicroservice.Added telemetry collection of the application metrics for VDMS-dataprep microservice and VLM microservice at
/telemetryendpoint.
HW used for validation:
Intel® Xeon® 5 + Intel® Arc™ B580 GPU
Vanilla Kubernetes Cluster