Adaptive Token Compressor#
Adaptive Token Compressor is a pluggable Python library purpose-built for LLM agent systems. Through a single unified compressor interface it applies tailored compression to each part of an agent pipeline — system prompt (harness), conversation context, and tool schemas — to significantly reduce token usage and improve inference efficiency on edge hardware.
Use Cases#
Reducing LLM inference cost in agentic pipelines with large system prompts or tool catalogs.
Compressing tool schemas at runtime so only contextually relevant tools are included in the prompt.
Tracking per-compressor and cross-compressor token savings, latency, and compression ratio telemetry.
Drop-in integration into existing agent frameworks via a factory-based compressor API.
Key Capabilities#
Unified compressor API with two specialized types:
harness(conversation messages and system prompt) andtool(tool schema selection).HarnessCompressor — section-aware message compression using hybrid rule-based and LLMLingua-2 model-based compression; preserves instruction-critical content while reducing token cost.
ToolCompressor — LLM-guided tool selection that scores and filters tool schemas from conversation context, with configurable injection placements to trade off token savings against prefix-cache hit rate.
Lingua Server backends — supports PyTorch and OpenVINO backends for the Lingua compression service, including Intel GPU (XPU) acceleration.
CompressionManager — orchestrates multiple compressors in a single pipeline and provides cross-compressor metric aggregation (tokens saved, compression ratio, average latency per request).
Factory construction — create and inspect compressors by type name at runtime using
create_compressor(),available_compressor_types(), andconfig_schema().
Get Started#
If you want to reduce token usage and improve inference efficiency, start here with the basics:
Installation — set up the library and required dependencies.
Quick Start — run your first compression workflow in just a few steps.
Once it is up and running, explore the full Get Started guide for compressor metrics, multi-compressor usage, workflow concepts, configuration options, testing, FAQ, and additional resources.
Next Steps#
Deploy Lingua Server — set up the Lingua backend required by HarnessCompressor.
Deploy LLM for Tool Prediction — set up the predictor service required by ToolCompressor.
Release Notes — version history and changelog.