Adaptive Token Compressor#
Adaptive Token Compressor is a pluggable Python library purpose-built for LLM agent systems. Through a single unified compressor interface it applies tailored compression to each part of an agent pipeline — system prompt (harness), conversation context, and tool schemas — to significantly reduce token usage and improve inference efficiency on edge hardware.
Use Cases#
Reducing LLM inference cost in agentic pipelines with large system prompts or tool catalogs.
Compressing tool schemas at runtime so only contextually relevant tools are included in the prompt.
Tracking per-compressor and cross-compressor token savings, latency, and compression ratio telemetry.
Drop-in integration into existing agent frameworks via a factory-based compressor API.
Key Capabilities#
Unified compressor API with two specialised types:
harness(conversation messages and system prompt) andtool(tool schema selection).HarnessCompressor — section-aware message compression using hybrid rule-based and LLMLingua-2 model-based compression; preserves instruction-critical content while reducing token cost.
ToolCompressor — LLM-guided tool selection that scores and filters tool schemas from conversation context, with configurable injection placements to trade off token savings against prefix-cache hit rate.
Lingua Server backends — supports PyTorch and OpenVINO backends for the Lingua compression service, including Intel GPU (XPU) acceleration.
CompressionManager — orchestrates multiple compressors in a single pipeline and provides cross-compressor metric aggregation (tokens saved, compression ratio, average latency per request).
Factory construction — create and inspect compressors by type name at runtime using
create_compressor(),available_compressor_types(), andconfig_schema().
Next Steps#
Deploy Lingua Server — set up the Lingua backend required by HarnessCompressor.
Deploy LLM for Tool Prediction — set up the predictor service required by ToolCompressor.
Guide — metrics collection, multi-compressor pipelines, configuration reference, and FAQ.
Release Notes — version history and changelog.