gvaclassify#
Performs object classification. Accepts the ROIs or full frame as input and outputs classification results with metadata.
Pad Templates:
SINK template: 'sink'
Availability: Always
Capabilities:
video/x-raw
format: { (string)BGRx, (string)BGRA, (string)BGR, (string)NV12, (string)I420 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
video/x-raw(memory:DMABuf)
format: { (string)DMA_DRM }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
video/x-raw(memory:VASurface)
format: { (string)NV12 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
video/x-raw(memory:VAMemory)
format: { (string)NV12 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
SRC template: 'src'
Availability: Always
Capabilities:
video/x-raw
format: { (string)BGRx, (string)BGRA, (string)BGR, (string)NV12, (string)I420 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
video/x-raw(memory:DMABuf)
format: { (string)DMA_DRM }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
video/x-raw(memory:VASurface)
format: { (string)NV12 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
video/x-raw(memory:VAMemory)
format: { (string)NV12 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
Element has no clocking capabilities.
Element has no URI handling capabilities.
Pads:
SINK: 'sink'
Pad Template: 'sink'
SRC: 'src'
Pad Template: 'src'
Element Properties:
batch-size : Number of frames batched together for a single inference. If the batch-size is 0, then it will be set by default to be optimal for the device. Not all models support batching. Use model optimizer to ensure that the model has batching support.
flags: readable, writable
Unsigned Integer. Range: 0 - 1024 Default: 0
batch-timeout : Timeout (ms) for OpenVINO™ Automatic Batching. Waits for batch to accumulate inference requests before execution. If the number of frames collected reaches batch-size, inference is executed with a full batch and the timer is reset. If timeout occurs before collecting all frames specified by batch-size, inference is executed on collected frames individually (as if batch-size=1) and the timer is reset. If batch-timeout is set to 0, it operates as if batch-size were set to 1, executing inference on individual frames. Value -1 disables timeout, waiting indefinitely for full batch. Note: Not supported with VA backends (pre-process-backend=va or va-surface-sharing).
flags: readable, writable
Integer. Range: -1 - 2147483647 Default: -1
core-pinning : List or range of CPU cores to pin this inference element to (e.g., '0-3' or '0,2,3').
flags: readable, writable
String. Default: null
cpu-throughput-streams: Deprecated. Use ie-config=CPU_THROUGHPUT_STREAMS=<number-streams> instead
flags: readable, writable, deprecated
Unsigned Integer. Range: 0 - 4294967295 Default: 0
custom-postproc-lib : Path to the .so file defining custom model output converter. The library must implement the Convert function: void Convert(GstTensorMeta *outputTensors, const GstStructure *network, const GstStructure *params, GstAnalyticsRelationMeta *relationMeta);
flags: readable, writable
String. Default: null
custom-preproc-lib : Path to the .so file defining custom input image pre-processing
flags: readable, writable
String. Default: null
device : Target device for inference. Please see OpenVINO™ Toolkit documentation for list of supported devices.
flags: readable, writable
String. Default: "CPU"
gpu-throughput-streams: Deprecated. Use ie-config=GPU_THROUGHPUT_STREAMS=<number-streams> instead
flags: readable, writable, deprecated
Unsigned Integer. Range: 0 - 4294967295 Default: 0
ie-config : Comma separated list of KEY=VALUE parameters for Inference Engine configuration. See OpenVINO™ Toolkit documentation for available parameters
flags: readable, writable
String. Default: ""
inference-interval : Interval between inference requests. An interval of 1 (Default) performs inference on every frame. An interval of 2 performs inference on every other frame. An interval of N performs inference on every Nth frame.
flags: readable, writable
Unsigned Integer. Range: 1 - 4294967295 Default: 1
inference-region : Identifier responsible for the region on which inference will be performed
flags: readable, writable
Enum "InferenceRegionType3" Default: 1, "roi-list"
(0): full-frame - Perform inference for full frame
(1): roi-list - Perform inference for roi list
labels : Array of object classes. It could be set as the following example: labels=<label1,label2,label3>
flags: readable, writable
String. Default: null
labels-file : Path to .txt file containing object classes (one per line)
flags: readable, writable
String. Default: null
model : Path to inference model network file
flags: readable, writable
String. Default: null
model-instance-id : Identifier for sharing a loaded model instance between elements of the same type. Elements with the same model-instance-id will share all model and inference engine related properties
flags: readable, writable
String. Default: null
model-proc : Path to JSON file with description of input/output layers pre-processing/post-processing
flags: readable, writable
String. Default: null
name : The name of the object
flags: readable, writable
String. Default: "gvaclassify0"
nireq : Number of inference requests
flags: readable, writable
Unsigned Integer. Range: 0 - 1024 Default: 0
no-block : (Experimental) Option to help maintain frames per second of incoming stream. Skips inference on an incoming frame if all inference requests are currently processing outstanding frames
flags: readable, writable, deprecated
Boolean. Default: false
object-class : Filter for Region of Interest class label on this element input
flags: readable, writable
String. Default: null
ov-extension-lib : Path to the .so file defining custom OpenVINO operations.
flags: readable, writable
String. Default: null
parent : The parent of the object
flags: readable, writable
Object of type "GstObject"
pre-process-backend : Select a pre-processing method (color conversion, resize and crop), one of 'ie', 'opencv', 'va', 'va-surface-sharing, 'vaapi', 'vaapi-surface-sharing'. If not set, it will be selected automatically: 'va' for VAMemory and DMABuf, 'ie' for SYSTEM memory.
flags: readable, writable
String. Default: ""
pre-process-config : Comma separated list of KEY=VALUE parameters for image processing pipeline configuration
flags: readable, writable
String. Default: ""
qos : Handle Quality-of-Service events
flags: readable, writable
Boolean. Default: false
reclassify-interval : Determines how often to reclassify tracked objects. Only valid when used in conjunction with gvatrack.
The following values are acceptable:
- 0 - Do not reclassify tracked objects
- 1 - Always reclassify tracked objects
- 2:N - Tracked objects will be reclassified every N frames. Note the inference-interval is applied before determining if an object is to be reclassified (i.e. classification only occurs at a multiple of the inference interval)
flags: readable, writable
Unsigned Integer. Range: 0 - 4294967295 Default: 1
reshape : If true, model input layer will be reshaped to resolution of input frames (no resize operation before inference). Note: this feature has limitations, not all network supports reshaping.
flags: readable, writable
Boolean. Default: false
reshape-height : Height to which the network will be reshaped.
flags: readable, writable
Unsigned Integer. Range: 0 - 4294967295 Default: 0
reshape-width : Width to which the network will be reshaped.
flags: readable, writable
Unsigned Integer. Range: 0 - 4294967295 Default: 0
scale-method : Scale method to use in pre-preprocessing before inference. Only default and scale-method=fast (VAAPI based) supported in this element
flags: readable, writable
String. Default: null
scheduling-policy : Scheduling policy across streams sharing same model instance: throughput (select first incoming frame), latency (select frames with earliest presentation time out of the streams sharing same model-instance-id; recommended batch-size less than or equal to the number of streams)
flags: readable, writable
String. Default: "throughput"
skip-raw-tensors : Skip attaching raw classification output tensors to metadata. When false (default), converters may attach both the interpreted results (for example classification labels) and raw tensor payloads copied from the output layer (for example logits). If the add-tensor-data property of gvametaconvert is set to true, raw tensor data is included in the output JSON by gvametapublish. When true, converters still attach interpreted metadata but omit the raw payload, which helps avoid flooding the buffer and JSON output with large tensors such as depth maps.
flags: readable, writable
Boolean. Default: false
share-va-display-ctx: Whether to share VA Display context across inference elements: true (share context, default), false (do not share context)
flags: readable, writable
Boolean. Default: true
Special case: zero-shot classification with CLIP#
gvaclassify supports open-vocabulary (zero-shot) image classification. Rather than a model with a
fixed classification head, it runs a CLIP image encoder (vision tower + visual projection) and a
post-processing converter, clip_zeroshot, scores the resulting image embedding by cosine
similarity against precomputed text-label embeddings. The label set is supplied at runtime as a
.safetensors file, so classes can be added or changed by regenerating that file without retraining.
Pipeline#
flowchart LR
image["Image or ROI"]
encoder["CLIP image encoder<br/>Inference device"]
image_embedding["Image embedding<br/>1 x D"]
text_embeddings["Text-label embeddings<br/>Classes x D"]
converter["clip_zeroshot converter<br/>Host CPU"]
scoring["Cosine similarity<br/>Logit scaling and softmax"]
output["Top-k classification metadata<br/>Label, confidence and rank"]
image --> encoder --> image_embedding --> converter
text_embeddings --> converter
converter --> scoring --> output
Only the CLIP vision tower runs on the inference device. The similarity, temperature scaling and top-k are cheap host-side operations, which keeps the device graph static-shape (important for NPU).
Selecting zero-shot mode#
There is a single way to run zero-shot classification, and it has two halves:
model=- required, as for anygvaclassifyelement. It points at the CLIP image-encoder IR exported for zero-shot. That IR carriesmodel_type=clip_zeroshotin themodel_infosection of itsmodel.xml, and that is what selects theclip_zeroshotconverter.zeroshot-embeddings-file=<file>.safetensors- required, supplies the class bank to score against.
gvaclassify model=<clip-image-encoder>.xml zeroshot-embeddings-file=labels.safetensors
There are two CLIP converters and they must not be confused:
|
Converter |
Model output |
Used for |
|---|---|---|---|
|
|
Unprojected vision-tower output |
Image-to-image comparison outside the pipeline |
|
|
Projected image embedding (CLIP shared space) |
Zero-shot classification against text-label embeddings |
The zeroshot-embeddings-file property supplies the class bank; it does not select the
converter. Both mismatches are rejected at pipeline construction with an explicit error:
model_type=clip_zeroshotwith nozeroshot-embeddings-file: there is nothing to classify against.zeroshot-embeddings-fileset on a model that is notclip_zeroshot: an unprojectedclip_tokenembedding lives in a different vector space, so cosine similarities against text embeddings would be meaningless.
Zero-shot properties:
Property |
Meaning |
Default |
|---|---|---|
|
Path to the |
unset |
|
Number of ranked classes to attach per region. |
1 |
Embeddings artifact (.safetensors)#
The file contains a single 2-D tensor named embeddings (also accepted: label_embeddings,
text_embeddings) of shape [num_classes, embedding_dim], dtype F32 or F16, with rows aligned
to the configured labels. Optional file metadata includes:
logit_scale: the model’s CLIP temperature (logit_scale.exp()), used to calibrate the softmax.unknown_threshold: top-1 cosine similarity below which a result is labelledunknown.model,labels,prompt: informational values.
The file is read natively (no Python or PyTorch dependency at runtime).
Choosing .safetensors over a pickled .pth avoids arbitrary code execution when loading a file
that, per the threat model, is treated as untrusted.
Preprocessing#
CLIP requires specific normalization. The exported IR carries this in the model_info section of
model.xml (mean_values and scale_values are the CLIP mean/std multiplied by 255,
color_space=RGB, resize_type=crop). DL Streamer reads it and composes the input affine transform. The IR input keeps a fixed spatial shape
[N, 3, 224, 224].
Calibration and the unknown class#
logit_scale: CLIP scores arelogit_scale * cosine(the learned temperature is about 100). Without it, a softmax over raw cosine in[-1, 1]is nearly flat and the confidences, while correctly ordered, are not meaningful. The converter applieslogit_scalefrom the embeddings file metadata before the softmax. If it is absent, the converter falls back to1.0and warns.unknown_threshold: If the top-1 cosine similarity is below this optional embeddings-file metadata value, the result is labelledunknown(label_id=-1) rather than forced to the nearest class. Thresholding on raw cosine keeps the decision independent oflogit_scale. An omitted or negative value disables the check.
Output metadata#
Each emitted classification carries the usual label, label_id, confidence and rank, plus
zs_mode (true), zs_unknown (bool) and zs_model (the model name).
Tooling and sample#
Model preparation reuses the Hugging Face scripts in scripts/download_models:
download_hf_models.py --model <clip_id> --export-variant clip-zeroshotexports the CLIP image encoder (the projected image embedding) to OpenVINO IR withmodel_type=clip_zeroshotand writes preprocessing intomodel_info.clip_text_embeddings.pyturnslabels.txtintolabels.safetensorswithlogit_scalemetadata.
See the end-to-end pipeline in
samples/gstreamer/gst_launch/zero_shot_classification/.