Model Preparation#
Skip this page if you already have a compatible pre-trained model exported to OpenVINO IR. This page covers the local training + export flow that produces the model artifact under
models/yolo11n_polyp/best_openvino_model/.
The Docker Compose runtime described in Get Started expects a trained OpenVINO IR to already exist on the host. If you don’t have one, this page walks through building it end-to-end from source dataset to FP16 IR on an Intel iGPU / Arc GPU.
The bootstrap flow is cache-first: it checks for
models/yolo11n_polyp/best_openvino_model/best.xml + a .trained_ok marker
before doing any work, so re-running make backend-bootstrap after a successful
run is effectively a no-op.
0. Install host prerequisites#
make setup-prerequisites installs everything the training venv and the Docker Compose runtime
need: base tools, Docker Engine + Compose v2, and the Intel client GPU stack
(Level Zero + OpenCL + iHD VA-API) from the official intel-graphics apt repo.
make setup-prerequisites # interactive; apt may prompt for confirmation
make setup-prerequisites SETUP_ARGS=-y # assume-yes to apt
make setup-prerequisites SETUP_ARGS=--dry-run
Then verify:
make check-l0 # dpkg check for libze1, libze-intel-gpu1, libigc2,
# libigdgmm12, intel-opencl-icd, intel-media-va-driver-non-free
# and /dev/dri device node
Log out and back in (or reboot) if make setup-prerequisites newly added your user to the
render, video, or docker groups.
1. Fetch the dataset#
The application is validated on REAL-Colon (Cosmo Intelligent Medical
Devices, figshare article 22202866). The full corpus is 60 studies (~880 GB).
The training subset we use is 4 studies (~67 GB).
make download-dataset # 4 studies, ~67 GB, to datasets/REAL-Colon/raw/
make prepare-dataset MAX_POS_PER_VIDEO=800 # take maximum 800 positive frames per video
2. Create the training virtualenv#
make backend-venv
Creates .venv-backend/ with:
torch==2.7.1+torchvision==0.22.1(both+xpubuilds from thepytorch/whl/xpuindex) — the Intel iGPU device backend.ultralytics==8.4.75(YOLO11 training).openvino==2026.2.0(FP16 IR export).
The venv is host-side (not in a container) so training uses the host’s Level
Zero driver and Intel iGPU directly. Requires the L0 stack installed by
make setup-prerequisites / verified by make check-l0.
3. Train + export#
make backend-bootstrap
Under the hood this runs python -m backend.main_bootstrap, which:
Auto-extracts dataset archives if needed.
Detects the REAL-Colon
*_frames/+*_annotations/layout, converts Pascal VOC XML bounding boxes to YOLO labels, and writes a Linux-cleandata.yamlunderdatasets/REAL-Colon/(70/15/15 train/val/test split, deterministic seed).Trains YOLO11n on the Intel iGPU (
device: xpu) for 50 epochs with the hyperparameters inbackend/config/model.yaml. Typical wall time on Arc iGPU (Meteor Lake / Lunar Lake / Arrow Lake) is ~20 minutes.Exports the best checkpoint to a FP16 OpenVINO IR at
models/yolo11n_polyp/best_openvino_model/best.xml+best.bin.Writes a
.trained_okmarker so subsequent runs cache-hit.
All defaults are in backend/config/model.yaml; override any of them via
environment variables:
DATASETS_DIR=/data/rc MODELS_DIR=/opt/models make backend-bootstrap
Or edit backend/config/model.yaml directly (e.g. change train.epochs,
train.batch, train.device, or add extra Ultralytics args).
4. Generate the demo video (required)#
Fresh clones do not include videos/polyp_test.mp4. Generate it from the
surgical-instrument/ workdir before running make doctor / make up. The
generator stitches frames from the REAL-Colon subset into an H.264 demo clip:
.venv-backend/bin/python scripts/create_endoscopy_video.py \
--images-dir datasets/REAL-Colon/raw/001-001_frames \
--output videos/polyp_test.mp4 \
--seconds 60 --fps 60 --width 1920 --height 1080
5. Verify and continue#
make doctor # confirms docker, /dev/dri, cached IR, demo video, L0 stack
make up ... # continue with the runtime flow in Get Started
make doctor prints an [OK] / [MISSING] line for every prerequisite and
exits non-zero if any hard requirement is missing.
Reset the cache#
To rebuild the model from scratch (e.g. after a dataset change):
rm -rf models/yolo11n_polyp/best_openvino_model \
models/yolo11n_polyp/.trained_ok \
datasets/REAL-Colon/data.yaml
make backend-bootstrap
The raw dataset archives under datasets/REAL-Colon/raw/ are left untouched —
only the derived labels, splits, and trained artifacts are regenerated.
See Troubleshooting for the common training-side
failure modes (torch.xpu unavailable, zeInit errors, dataset auto-detect
failures, etc.).