Get Started - Deploy Surgical Instrument#
This is a deployment guide for the Docker Compose stack:
hls-si-endoscopy— single Python service that captures frames from a Basler camera, USB/V4L2 webcam, or video file; runs OpenVINO inference; and renders results with an OpenGL vsync presenter.
Model required first. The stack expects an OpenVINO IR at
models/yolo11n_polyp/best_openvino_model/best.xml. If you do not already have one, follow Model Preparation before continuing.
Prerequisites#
Before you start, refer to System Requirements to confirm your setup compatibility.
Host tools#
The app runs entirely in a container, so the host only needs a small set of tools:
Tool |
Required? |
Why |
|---|---|---|
Docker Engine + Compose plugin |
Yes |
The app runs via |
make |
Yes |
|
git |
To get the code |
Needed for |
python3 |
No |
Python runs inside the container; the host does not need it. |
The easiest way to satisfy these on Ubuntu is the bundled setup script, which
installs/verifies Docker + Compose + make + git, adds your user to the
docker group, installs the Intel client GPU stack (Level Zero + OpenCL +
iHD VA-API), and configures a proxy only if you provide one:
# No proxy (typical):
make setup-prerequisites
# Behind a corporate proxy, export it first:
HTTP_PROXY=http://your-proxy:port HTTPS_PROXY=http://your-proxy:port \
make setup-prerequisites
Log out and back in afterward (or run newgrp docker) so docker-group
membership takes effect. On non-Ubuntu hosts, install Docker manually per the
Docker docs.
Model Preparation#
The stack expects an OpenVINO IR at
models/yolo11n_polyp/best_openvino_model/best.xml and (for SOURCE=file) a
demo video at videos/polyp_test.mp4. There are two ways to get there:
Pull a prebuilt image from the registry — the default
make upflow. You can skip local training and export, but forSOURCE=fileruns you still needmake download-datasetto assemble the demo video below.Build the model locally — install host prerequisites, download the REAL-Colon dataset subset, train YOLO11n on the Intel iGPU, and export a FP16 OpenVINO IR:
make check-l0 # verify host GPU stack
make backend-venv # create .venv-backend (torch+xpu, Ultralytics, OpenVINO)
make download-dataset # 7-study REAL-Colon subset (~67 GB) from figshare 22202866
make prepare-dataset MAX_POS_PER_VIDEO=800 # take maximum 800 positive frames per video
make backend-bootstrap # dataset -> train -> FP16 OpenVINO IR (cache-first)
Generate the demo video (required). Fresh clones do not include
videos/polyp_test.mp4. Generate it from the surgical-instrument/ workdir
before running make doctor / make up:
.venv-backend/bin/python scripts/create_endoscopy_video.py \
--images-dir datasets/REAL-Colon/raw/001-001_frames \
--output videos/polyp_test.mp4 \
--seconds 60 --fps 60 --width 1920 --height 1080
make doctor # preflight all runtime prerequisites
See Model Preparation for the full end-to-end walkthrough, dataset options, and cache-reset instructions.
Models and videos#
The application does not ship with the trained model binaries or demo videos. You need to place these resources under the host paths that are bind-mounted into the container:
models/yolo11n_polyp/best_openvino_model/best.xml
models/yolo11n_polyp/best_openvino_model/best.bin
videos/polyp_test.mp4
The Makefile mounts ../models to /models and ../videos to /videos by
default. Override with MODELS_DIR and VIDEOS_DIR when the host layout is
different.
A quick sanity check for all runtime prerequisites (Docker, /dev/dri, cached
IR, demo video, Intel L0 stack):
make doctor
1. Discover the camera and CPU topology#
make list-cameras # prints Basler serial(s) + model -> use as SERIAL=
make show-cores # prints P-core / E-core CPU sets -> use with CPU_CAPTURE/CPU_INFERENCE/CPU_DISPLAY
2. Bring the stack up#
make up supports two image sources, controlled by the REGISTRY flag.
2a. Pull the image from the registry (default)#
REGISTRY=true is the default. make up pulls the prebuilt image at TAG
(default latest) from REGISTRY_URL and starts it with the current runtime
knobs applied through Docker Compose.
# Low-latency live Basler camera (recommended tuned defaults).
# vsync-phase-locked capture + fixed 2ms exposure + core-pinned, SCHED_FIFO
# capture/inference/display threads — the configuration validated for lowest
# camera-to-screen latency.
make up SOURCE=camera LOWLATENCY=1 CAMERA_TRIGGER=vsync VSYNC_DIVISOR=2 EXPOSURE_US=2000 \
DEVICE=GPU FRAME_SKIP=3 SERIAL=<SERIAL_NUMBER> \
CPU_CAPTURE=1 CPU_INFERENCE=2 CPU_DISPLAY=3 RT_PRIORITY=80
# Minimal low-latency (let the app pick exposure / no core pinning).
make up LOWLATENCY=1 CAMERA_TRIGGER=vsync VSYNC_DIVISOR=2 SERIAL=<SERIAL_NUMBER>
# Baseline free-running camera.
make up SOURCE=camera SERIAL=<SERIAL_NUMBER>
# Video file (loops on EOF).
make up SOURCE=file SOURCE_ARG=/videos/polyp_test.mp4
# USB / V4L2 webcam.
make up SOURCE=webcam DEVICE_INDEX=0
# Select the OpenVINO inference device (default: GPU).
make up SERIAL=<SERIAL_NUMBER> DEVICE=CPU
make up SERIAL=<SERIAL_NUMBER> DEVICE=NPU # requires /dev/accel on the host
2b. Build the image from source#
REGISTRY=false builds the image locally from docker/Dockerfile before
starting the stack.
make up LOWLATENCY=1 CAMERA_TRIGGER=vsync VSYNC_DIVISOR=2 SERIAL=<SERIAL_NUMBER> REGISTRY=false
3. Watch logs#
make logs
Per-stage latency CSV is streamed to stdout when LATENCY_TRACE=1 (implied by
LOWLATENCY=1):
clock_time, trigger_to_grab_ms, grab_to_display_ms, trigger_to_display_ms, infer_ms, disp_fps, cap_fps
4. Stop the stack#
make down # stop + remove the container
make clean # also remove the built image
Press ESC inside the display window to quit the app.
Command reference#
Command |
What it does |
|---|---|
|
Pull or build the image and start the stack with current runtime knobs. |
|
Stop and remove the container. |
|
Follow container logs. |
|
Stop the stack and remove the built image. |
|
List connected Basler cameras (serial + model). |
|
Show P-core / E-core CPU sets for |
|
List all targets. |
For every runtime knob, CLI flag, and low-latency variant, see Runtime Configuration.