Configuration Guide#
The Semantic Search Agent uses environment variables or the .env file at the project root folder, for all configurations. You do not need a YAML subscription file; data sources are JSON files under the config/ folder.
Load Order#
The service loads the configuration in the following order:
Environment Variables or the
.envfile: Loaded via Pydantic Settings on startup. Environment variables take precedence over.envfile values.Configuration JSON Files: The ComparisonEngine loads
config/inventory.jsonandconfig/orders.jsonlazily on the first use and caches them in memory for the lifetime of the process.
Environment Variables#
You can set all variables as real environment variables or place them in the .env file at the project root folder.
Service Settings#
Variable |
Default |
Description |
|---|---|---|
|
|
Service name reported in health and logs. |
|
|
Service version reported in health responses. |
|
|
Logging level: |
|
|
Port the FastAPI server listens on. |
|
|
Port for the Prometheus |
|
|
Enable or disable the Prometheus metrics mount. |
Matching Configuration#
Variable |
Default |
Description |
|---|---|---|
|
|
Matching strategy: |
|
|
Minimum VLM confidence score to consider a semantic match successful. |
|
|
Maximum retries for VLM inference calls on transient failures. |
VLM Backend Selection#
Variable |
Default |
Description |
|---|---|---|
|
|
VLM backend to use: |
OpenVINO Model Server Backend Settings#
Required when VLM_BACKEND=ovms and strategy is semantic or hybrid.
Variable |
Default |
Description |
|---|---|---|
|
(empty) |
Full base URL of the OpenVINO model server (e.g. |
|
(empty) |
Model name served by OpenVINO model server (e.g. |
|
|
HTTP request timeout in seconds for OpenVINO model server calls. |
OpenVINO Toolkit Local Backend Settings#
Required when VLM_BACKEND=openvino_local and strategy is semantic or hybrid.
Variable |
Default |
Description |
|---|---|---|
|
(empty) |
Path to the OpenVINO IR model directory on disk. ⚠️ Required. |
|
|
Inference device: |
|
|
Maximum tokens to generate per inference call. |
|
|
Sampling temperature (0.0 = deterministic). |
OpenAI Backend Settings#
Required when VLM_BACKEND=openai and strategy is semantic or hybrid.
Variable |
Default |
Description |
|---|---|---|
|
(empty) |
OpenAI API key. ⚠️ Required. |
|
|
OpenAI model identifier to use for inference. |
|
|
Maximum tokens to generate per API call. |
Cache Settings#
Variable |
Default |
Description |
|---|---|---|
|
|
Enable or disable response caching for semantic match results. |
|
|
Cache backend: |
|
|
Redis hostname (used when |
|
|
Redis port. |
|
|
Redis database index. |
|
|
Cache entry time-to-live in seconds. |
Proxy Settings#
Pass these when running the service or building the Docker image behind a corporate proxy:
Variable |
Default |
Description |
|---|---|---|
|
(empty) |
HTTP proxy URL (e.g. |
|
(empty) |
HTTPS proxy URL. |
|
(empty) |
Comma-separated list of hosts to bypass the proxy. |
Note
The OpenVINO model server backend sets trust_env=False on its HTTP client to bypass proxy settings for internal OpenVINO model server communication. This is intentional because OpenVINO model server hosts are typically on the same internal network.
Configuration JSON Files#
Two JSON data files drive the comparison engine’s data sources. Their paths can be overridden via environment variables; defaults point to the config/ directory in the project root.
config/inventory.json#
A flat JSON array of item name strings that represent the available inventory:
[
"apple",
"banana",
"milk",
"bread",
"eggs",
"butter"
]
config/orders.json#
A JSON object that maps order IDs to lists of expected items with names and quantities:
{
"order_001": [
{"name": "apple", "quantity": 3},
{"name": "milk", "quantity": 2}
]
}
Path Override Variables#
Variable |
Default |
Description |
|---|---|---|
|
|
Base directory for configuration files. |
|
|
Path to the orders JSON file. |
|
|
Path to the inventory JSON file. |
Matching Strategy Reference#
Strategy |
VLM Required |
Behavior |
|---|---|---|
|
No |
Normalizes both strings by converting to lowercase, trimming, and stripping special characters, and compares them directly. |
|
Yes |
Sends a structured YES/NO prompt to the VLM backend for every comparison. Result is cached. |
|
Yes |
Tries exact match first. If exact confidence ≥ |