gvamono3d#

Performs monocular 3D object detection (e.g. MonoDETR). Consumes a single camera image together with the camera calibration and produces, for each detected object, a 2D bounding box annotated with a full 3D cuboid (location, dimensions and orientation).

gvamono3d is an inference element built on the same engine as gvadetect, so it inherits batching, VA pre-processing, multi-stream model sharing and the standard inference properties. The model is driven directly from its rt_info (model_type=mono3d). Detections are attached as GstAnalyticsCamera3DODMtd metadata, serialized to JSON by gvametaconvert and rendered by gvawatermark3d.

The camera calibration (KITTI P2 and the original image size) is required for the 3D lift and is provided through the calibration-file property. Both a KITTI .txt calibration (line starting with P2:) and a JSON file with an intrinsic_matrix (3x3) or projection_matrix (3x4) and optional image_size are supported.

Pad Templates:
  SINK template: 'sink'
    Availability: Always
    Capabilities:
      video/x-raw
                format: { (string)BGRx, (string)BGRA, (string)BGR, (string)NV12, (string)I420 }
      video/x-raw(memory:DMABuf)
                format: { (string)DMA_DRM }
      video/x-raw(memory:VASurface)
                format: { (string)NV12 }
      video/x-raw(memory:VAMemory)
                format: { (string)NV12 }

  SRC template: 'src'
    Availability: Always
    Capabilities:
      video/x-raw
                format: { (string)BGRx, (string)BGRA, (string)BGR, (string)NV12, (string)I420 }
      video/x-raw(memory:DMABuf)
                format: { (string)DMA_DRM }
      video/x-raw(memory:VASurface)
                format: { (string)NV12 }
      video/x-raw(memory:VAMemory)
                format: { (string)NV12 }

Element Properties:
  model               : Path to inference model network file (OpenVINO™ IR .xml).
                        flags: readable, writable
                        String. Default: null
  calibration-file    : Path to camera calibration: KITTI .txt (P2 row) or JSON with "intrinsic_matrix" (3x3) / "projection_matrix" (3x4) and optional "image_size".
                        flags: readable, writable
                        String. Default: null
  device              : Target device for inference. Please see OpenVINO™ Toolkit documentation for list of supported devices.
                        flags: readable, writable
                        String. Default: "CPU"
  threshold           : Only detections with confidence above the threshold are attached to the frame.
                        flags: readable, writable
                        Float. Range: 0 - 1 Default: 0.5

In addition to the properties above, gvamono3d supports the common inference element properties (batch-size, nireq, ie-config, model-instance-id, pre-process-backend, inference-interval, reshape, …); see gvadetect for their descriptions.

Calibration file formats#

Set calibration-file to either a KITTI calibration text file or a JSON file. The projection matrix is stored in row-major order and maps homogeneous 3D camera coordinates [X, Y, Z, 1] to image coordinates.

KITTI text#

The P2: row must start at the beginning of a line and contain the 12 values of the 3x4 projection matrix in row-major order. Other rows are ignored.

P2: 7.215377e+02 0.000000e+00 6.095593e+02 4.485728e+01 0.000000e+00 7.215377e+02 1.728540e+02 2.163791e-01 0.000000e+00 0.000000e+00 1.000000e+00 2.745884e-03

Equivalently, the values after P2: are ordered as follows:

p00 p01 p02 p03 p10 p11 p12 p13 p20 p21 p22 p23

KITTI text files do not specify the image dimensions, so gvamono3d uses the negotiated input frame width and height.

JSON#

A JSON file can provide the full 3x4 projection matrix as nested arrays:

{
  "projection_matrix": [
    [721.5377, 0.0, 609.5593, 44.85728],
    [0.0, 721.5377, 172.854, 0.2163791],
    [0.0, 0.0, 1.0, 0.002745884]
  ],
  "image_size": [1242, 375]
}

Alternatively, provide a 3x3 intrinsic matrix. gvamono3d converts it to a projection matrix by appending a zero column (P2 = [K | 0]):

{
  "intrinsic_matrix": [
    [721.5377, 0.0, 609.5593],
    [0.0, 721.5377, 172.854],
    [0.0, 0.0, 1.0]
  ],
  "image_size": [1242, 375]
}

image_size is optional and is ordered as [width, height]. It must describe the frame dimensions to which the calibration applies, before the inference element’s internal resize. If omitted, the negotiated input frame dimensions are used. If both matrix fields are present, projection_matrix takes precedence. File names ending in .json (case-insensitive) are parsed as JSON; all other file names are parsed as KITTI text.

On GPU the execution mode defaults to ACCURACY automatically (MonoDETR’s backbone overflows FP16 under the default performance mode); this can be overridden with ie-config=EXECUTION_MODE_HINT=….

Examples#

# CPU
gst-launch-1.0 filesrc location=input.mp4 ! decodebin3 ! \
  gvamono3d model=monodetr.xml device=CPU calibration-file=000000.txt threshold=0.2 ! \
  gvawatermark3d calibration-file=000000.txt ! videoconvert ! autovideosink

# GPU (execution mode defaults to ACCURACY)
gst-launch-1.0 filesrc location=input.mp4 ! decodebin3 ! \
  gvamono3d model=monodetr.xml device=GPU calibration-file=000000.txt threshold=0.2 ! \
  gvametaconvert format=json ! gvametapublish method=file file-path=out.jsonl ! fakesink