Troubleshooting#
This guide covers common runtime failures observed from the source code, Docker configuration, and service startup logic.
Service Fails to Start — SeaweedFS Not Ready#
Symptom:
SeaweedFS not ready, retrying in 2s (attempt 1): ...
Failed to ensure bucket after retries: ...
Cause: The service attempts to create/verify the SeaweedFS bucket on startup and retries up to 5 times. If SeaweedFS is still initializing or unreachable, startup will fail after all retries.
Resolution:
Ensure the SeaweedFS service is running and healthy before starting the behavioral-analysis service.
In Docker Compose, the
depends_on: seaweedfs: condition: service_healthysetting handles this automatically; verify that the SeaweedFS health check is passing.Check that
SEAWEEDFS_ENDPOINTis correct and reachable from within the container (e.g.,http://seaweedfs:8333).
# Verify SeaweedFS is reachable
curl http://seaweedfs:8333
Service Fails to Start — YOLO Model Not Found#
Symptom:
ValueError: Expected .xml model path, got: /models/yolo_models/yolo26n-pose/yolo26n-pose.xml
FileNotFoundError: Model file not found
Cause: The YOLO-Pose OpenVINO IR model (.xml + .bin) is not present at the path specified by YOLO_POSE_MODEL.
Resolution:
Download the model and place it at
./models/yolo_models/yolo26n-pose/yolo26n-pose.xml(and.bin).Ensure the volume mount in Docker Compose is correct:
${DOWNLOADED_MODEL_PATH:-./models}:/models:ro.The model path inside the container is
/models/yolo_models/yolo26n-pose/yolo26n-pose.xml.
# Verify model file exists on the host
ls -la ./models/yolo_models/yolo26n-pose/
Service Health and Connectivity#
Symptom: Service starts but does not process requests.
Cause: SeaweedFS bucket is unavailable or MQTT connection failed.
Resolution:
Verify SeaweedFS is running:
curl http://seaweedfs:8333.Check the
SEAWEEDFS_ENDPOINTandSEAWEEDFS_BUCKETenvironment variables.Verify the MQTT broker is running and accessible.
Check the
MQTT_HOSTandMQTT_PORTenvironment variables.Review logs:
docker logs behavioral-analysisor application stdout.
VLM Analysis Not Running / vlm_confirmed Always null#
Symptom: vlm_confirmed is always null in responses.
Cause: VLM is disabled, or the OVMS endpoint is unreachable.
Resolution:
Check whether
VLM_ENABLED=trueis set when you want VLM enabled; this is the global runtime switch.If the global VLM flag is disabled, the service will skip VLM confirmation regardless of pattern-level YAML values. Pattern-level
vlm.enabledonly matters when the service-level flag is already enabled.Confirm the OVMS service is running and responding:
curl http://ovms-vlm:8001/v2/health/ready.Check that
VLM_ENDPOINTpoints to the correct host and port.In Docker Compose,
depends_on: ovms-vlm: condition: service_healthyensures OVMS is ready before the service starts.
VLM Circuit Breaker Open#
Symptom in logs:
VLM circuit breaker open — skipping analysis (cooldown: 30s remaining)
Cause: The VLM client encountered 3 or more consecutive failures and opened the circuit breaker. Requests are not sent to OVMS for the 30-second cooldown period.
Resolution:
Check OVMS health:
curl http://ovms-vlm:8001/v2/health/ready.Verify the model is loaded:
curl http://ovms-vlm:8001/v2/models.Once OVMS recovers, the circuit breaker automatically probes after 30 seconds.
MQTT Consumer Not Receiving Messages#
Symptom: Service starts but no analyses are triggered from ba/requests.
Cause: MQTT connection failure or incorrect topic configuration.
Resolution:
Check logs for:
BA queue consumer connected, subscribed to ba/requests.If you see
BA queue consumer MQTT connect failed, rc=..., the broker is unreachable.Verify
MQTT_HOSTandMQTT_PORTenvironment variables.Confirm the MQTT broker is running and accessible from the container.
Verify the upstream service is publishing to the same topic (
ba/requestsor the value ofBA_REQUEST_TOPIC).
# Test broker reachability (from inside container)
python3 -c "import socket; s=socket.create_connection(('broker.scenescape.intel.com', 1883), timeout=5); print('OK')"
Container Exits Immediately After Starting#
Symptom: Container exits with code 1 shortly after launch.
Resolution:
Inspect logs:
docker logs <container-name>.Common causes:
Missing
SEAWEEDFS_ENDPOINTenvironment variable.YOLO model file missing at the mounted path.
Python import error (missing dependency — rebuild the image).
docker logs behavioral-analysis --tail 50
Log Locations#
Context |
Location |
|---|---|
Docker Compose |
|
Running container |
|
Standalone |
Standard output (stdout); redirect with |
Debugging Steps#
Check the health endpoint:
curl http://localhost:8085/health.Check service logs for startup errors (model loading, bucket creation, MQTT connection).
Verify all required services are running: SeaweedFS, OVMS (if VLM enabled), MQTT broker.
Confirm that all environment variables are set (see Configuration).
Confirm the YOLO model files exist at the configured path.
Enable debug logging: set
LOG_LEVEL=DEBUGand restart.Use
POST /api/v1/analyzewith a known entity to isolate whether the issue is frame storage, pose extraction, or VLM.