gvaclassify#

Performs object classification. Accepts the ROIs or full frame as input and outputs classification results with metadata.

Pad Templates:
SINK template: 'sink'
   Availability: Always
   Capabilities:
      video/x-raw
               format: { (string)BGRx, (string)BGRA, (string)BGR, (string)NV12, (string)I420 }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]
      video/x-raw(memory:DMABuf)
               format: { (string)DMA_DRM }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]
      video/x-raw(memory:VASurface)
               format: { (string)NV12 }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]
      video/x-raw(memory:VAMemory)
               format: { (string)NV12 }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]

SRC template: 'src'
   Availability: Always
   Capabilities:
      video/x-raw
               format: { (string)BGRx, (string)BGRA, (string)BGR, (string)NV12, (string)I420 }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]
      video/x-raw(memory:DMABuf)
               format: { (string)DMA_DRM }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]
      video/x-raw(memory:VASurface)
               format: { (string)NV12 }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]
      video/x-raw(memory:VAMemory)
               format: { (string)NV12 }
                  width: [ 1, 2147483647 ]
                  height: [ 1, 2147483647 ]
            framerate: [ 0/1, 2147483647/1 ]

Element has no clocking capabilities.
Element has no URI handling capabilities.

Pads:
SINK: 'sink'
   Pad Template: 'sink'
SRC: 'src'
   Pad Template: 'src'

Element Properties:
batch-size          : Number of frames batched together for a single inference. If the batch-size is 0, then it will be set by default to be optimal for the device. Not all models support batching. Use model optimizer to ensure that the model has batching support.
                        flags: readable, writable
                        Unsigned Integer. Range: 0 - 1024 Default: 0
batch-timeout       : Timeout (ms) for OpenVINO™ Automatic Batching. Waits for batch to accumulate inference requests before execution. If the number of frames collected reaches batch-size, inference is executed with a full batch and the timer is reset. If timeout occurs before collecting all frames specified by batch-size, inference is executed on collected frames individually (as if batch-size=1) and the timer is reset. If batch-timeout is set to 0, it operates as if batch-size were set to 1, executing inference on individual frames. Value -1 disables timeout, waiting indefinitely for full batch. Note: Not supported with VA backends (pre-process-backend=va or va-surface-sharing).
                        flags: readable, writable
                        Integer. Range: -1 - 2147483647 Default: -1
core-pinning        : List or range of CPU cores to pin this inference element to (e.g., '0-3' or '0,2,3'). 
                        flags: readable, writable
                        String. Default: null
cpu-throughput-streams: Deprecated. Use ie-config=CPU_THROUGHPUT_STREAMS=<number-streams> instead
                        flags: readable, writable, deprecated
                        Unsigned Integer. Range: 0 - 4294967295 Default: 0
custom-postproc-lib : Path to the .so file defining custom model output converter. The library must implement the Convert function: void Convert(GstTensorMeta *outputTensors, const GstStructure *network, const GstStructure *params, GstAnalyticsRelationMeta *relationMeta);
                        flags: readable, writable
                        String. Default: null
custom-preproc-lib  : Path to the .so file defining custom input image pre-processing
                        flags: readable, writable
                        String. Default: null
device              : Target device for inference. Please see OpenVINO™ Toolkit documentation for list of supported devices.
                        flags: readable, writable
                        String. Default: "CPU"
gpu-throughput-streams: Deprecated. Use ie-config=GPU_THROUGHPUT_STREAMS=<number-streams> instead
                        flags: readable, writable, deprecated
                        Unsigned Integer. Range: 0 - 4294967295 Default: 0
ie-config           : Comma separated list of KEY=VALUE parameters for Inference Engine configuration. See OpenVINO™ Toolkit documentation for available parameters
                        flags: readable, writable
                        String. Default: ""
inference-interval  : Interval between inference requests. An interval of 1 (Default) performs inference on every frame. An interval of 2 performs inference on every other frame. An interval of N performs inference on every Nth frame.
                        flags: readable, writable
                        Unsigned Integer. Range: 1 - 4294967295 Default: 1
inference-region    : Identifier responsible for the region on which inference will be performed
                        flags: readable, writable
                        Enum "InferenceRegionType3" Default: 1, "roi-list"
                           (0): full-frame       - Perform inference for full frame
                           (1): roi-list         - Perform inference for roi list
labels              : Array of object classes. It could be set as the following example: labels=<label1,label2,label3>
                        flags: readable, writable
                        String. Default: null
labels-file         : Path to .txt file containing object classes (one per line)
                        flags: readable, writable
                        String. Default: null
model               : Path to inference model network file
                        flags: readable, writable
                        String. Default: null
model-instance-id   : Identifier for sharing a loaded model instance between elements of the same type. Elements with the same model-instance-id will share all model and inference engine related properties
                        flags: readable, writable
                        String. Default: null
model-proc          : Path to JSON file with description of input/output layers pre-processing/post-processing
                        flags: readable, writable
                        String. Default: null
name                : The name of the object
                        flags: readable, writable
                        String. Default: "gvaclassify0"
nireq               : Number of inference requests
                        flags: readable, writable
                        Unsigned Integer. Range: 0 - 1024 Default: 0
no-block            : (Experimental) Option to help maintain frames per second of incoming stream. Skips inference on an incoming frame if all inference requests are currently processing outstanding frames
                        flags: readable, writable, deprecated
                        Boolean. Default: false
object-class        : Filter for Region of Interest class label on this element input
                        flags: readable, writable
                        String. Default: null
ov-extension-lib    : Path to the .so file defining custom OpenVINO operations.
                        flags: readable, writable
                        String. Default: null
parent              : The parent of the object
                        flags: readable, writable
                        Object of type "GstObject"
pre-process-backend : Select a pre-processing method (color conversion, resize and crop), one of 'ie', 'opencv', 'va', 'va-surface-sharing, 'vaapi', 'vaapi-surface-sharing'. If not set, it will be selected automatically: 'va' for VAMemory and DMABuf, 'ie' for SYSTEM memory.
                        flags: readable, writable
                        String. Default: ""
pre-process-config  : Comma separated list of KEY=VALUE parameters for image processing pipeline configuration
                        flags: readable, writable
                        String. Default: ""
qos                 : Handle Quality-of-Service events
                        flags: readable, writable
                        Boolean. Default: false
reclassify-interval : Determines how often to reclassify tracked objects. Only valid when used in conjunction with gvatrack.
The following values are acceptable:
- 0 - Do not reclassify tracked objects
- 1 - Always reclassify tracked objects
- 2:N - Tracked objects will be reclassified every N frames. Note the inference-interval is applied before determining if an object is to be reclassified (i.e. classification only occurs at a multiple of the inference interval)
                        flags: readable, writable
                        Unsigned Integer. Range: 0 - 4294967295 Default: 1
reshape             : If true, model input layer will be reshaped to resolution of input frames (no resize operation before inference). Note: this feature has limitations, not all network supports reshaping.
                        flags: readable, writable
                        Boolean. Default: false
reshape-height      : Height to which the network will be reshaped.
                        flags: readable, writable
                        Unsigned Integer. Range: 0 - 4294967295 Default: 0
reshape-width       : Width to which the network will be reshaped.
                        flags: readable, writable
                        Unsigned Integer. Range: 0 - 4294967295 Default: 0
scale-method        : Scale method to use in pre-preprocessing before inference. Only default and scale-method=fast (VAAPI based) supported in this element
                        flags: readable, writable
                        String. Default: null
scheduling-policy   : Scheduling policy across streams sharing same model instance: throughput (select first incoming frame), latency (select frames with earliest presentation time out of the streams sharing same model-instance-id; recommended batch-size less than or equal to the number of streams)
                        flags: readable, writable
                        String. Default: "throughput"
skip-raw-tensors    : Skip attaching raw classification output tensors to metadata. When false (default), converters may attach both the interpreted results (for example classification labels) and raw tensor payloads copied from the output layer (for example logits). If the add-tensor-data property of gvametaconvert is set to true, raw tensor data is included in the output JSON by gvametapublish. When true, converters still attach interpreted metadata but omit the raw payload, which helps avoid flooding the buffer and JSON output with large tensors such as depth maps.
                        flags: readable, writable
                        Boolean. Default: false
share-va-display-ctx: Whether to share VA Display context across inference elements: true (share context, default), false (do not share context)
                        flags: readable, writable
                        Boolean. Default: true

Special case: zero-shot classification with CLIP#

gvaclassify supports open-vocabulary (zero-shot) image classification. Rather than a model with a fixed classification head, it runs a CLIP image encoder (vision tower + visual projection) and a post-processing converter, clip_zeroshot, scores the resulting image embedding by cosine similarity against precomputed text-label embeddings. The label set is supplied at runtime as a .safetensors file, so classes can be added or changed by regenerating that file without retraining.

Pipeline#

        flowchart LR
   image["Image or ROI"]
   encoder["CLIP image encoder<br/>Inference device"]
   image_embedding["Image embedding<br/>1 x D"]
   text_embeddings["Text-label embeddings<br/>Classes x D"]
   converter["clip_zeroshot converter<br/>Host CPU"]
   scoring["Cosine similarity<br/>Logit scaling and softmax"]
   output["Top-k classification metadata<br/>Label, confidence and rank"]

   image --> encoder --> image_embedding --> converter
   text_embeddings --> converter
   converter --> scoring --> output
    

Only the CLIP vision tower runs on the inference device. The similarity, temperature scaling and top-k are cheap host-side operations, which keeps the device graph static-shape (important for NPU).

Selecting zero-shot mode#

There is a single way to run zero-shot classification, and it has two halves:

  1. model= - required, as for any gvaclassify element. It points at the CLIP image-encoder IR exported for zero-shot. That IR carries model_type=clip_zeroshot in the model_info section of its model.xml, and that is what selects the clip_zeroshot converter.

  2. zeroshot-embeddings-file=<file>.safetensors - required, supplies the class bank to score against.

gvaclassify model=<clip-image-encoder>.xml zeroshot-embeddings-file=labels.safetensors

There are two CLIP converters and they must not be confused:

model_type

Converter

Model output

Used for

clip_token

clip_token

Unprojected vision-tower output

Image-to-image comparison outside the pipeline

clip_zeroshot

clip_zeroshot

Projected image embedding (CLIP shared space)

Zero-shot classification against text-label embeddings

The zeroshot-embeddings-file property supplies the class bank; it does not select the converter. Both mismatches are rejected at pipeline construction with an explicit error:

  • model_type=clip_zeroshot with no zeroshot-embeddings-file: there is nothing to classify against.

  • zeroshot-embeddings-file set on a model that is not clip_zeroshot: an unprojected clip_token embedding lives in a different vector space, so cosine similarities against text embeddings would be meaningless.

Zero-shot properties:

Property

Meaning

Default

zeroshot-embeddings-file

Path to the .safetensors class embeddings. Required when the model is clip_zeroshot.

unset

zeroshot-topk

Number of ranked classes to attach per region.

1

Embeddings artifact (.safetensors)#

The file contains a single 2-D tensor named embeddings (also accepted: label_embeddings, text_embeddings) of shape [num_classes, embedding_dim], dtype F32 or F16, with rows aligned to the configured labels. Optional file metadata includes:

  • logit_scale: the model’s CLIP temperature (logit_scale.exp()), used to calibrate the softmax.

  • unknown_threshold: top-1 cosine similarity below which a result is labelled unknown.

  • model, labels, prompt: informational values.

The file is read natively (no Python or PyTorch dependency at runtime). Choosing .safetensors over a pickled .pth avoids arbitrary code execution when loading a file that, per the threat model, is treated as untrusted.

Preprocessing#

CLIP requires specific normalization. The exported IR carries this in the model_info section of model.xml (mean_values and scale_values are the CLIP mean/std multiplied by 255, color_space=RGB, resize_type=crop). DL Streamer reads it and composes the input affine transform. The IR input keeps a fixed spatial shape [N, 3, 224, 224].

Calibration and the unknown class#

  • logit_scale: CLIP scores are logit_scale * cosine (the learned temperature is about 100). Without it, a softmax over raw cosine in [-1, 1] is nearly flat and the confidences, while correctly ordered, are not meaningful. The converter applies logit_scale from the embeddings file metadata before the softmax. If it is absent, the converter falls back to 1.0 and warns.

  • unknown_threshold: If the top-1 cosine similarity is below this optional embeddings-file metadata value, the result is labelled unknown (label_id=-1) rather than forced to the nearest class. Thresholding on raw cosine keeps the decision independent of logit_scale. An omitted or negative value disables the check.

Output metadata#

Each emitted classification carries the usual label, label_id, confidence and rank, plus zs_mode (true), zs_unknown (bool) and zs_model (the model name).

Tooling and sample#

Model preparation reuses the Hugging Face scripts in scripts/download_models:

  • download_hf_models.py --model <clip_id> --export-variant clip-zeroshot exports the CLIP image encoder (the projected image embedding) to OpenVINO IR with model_type=clip_zeroshot and writes preprocessing into model_info.

  • clip_text_embeddings.py turns labels.txt into labels.safetensors with logit_scale metadata.

See the end-to-end pipeline in samples/gstreamer/gst_launch/zero_shot_classification/.