AX Optimize.
Measurably reduce LLM token costs — open-source-based context compression, plus cost governance.
Agent runs balloon: tool outputs, logs, RAG chunks and long conversations flow into the context round after round — and every token costs money. AX Optimize reduces these costs measurably by compressing the context ballast before it reaches the model. In the real-world run above: 66.8% fewer tokens with answer quality preserved.
$ ax-optimize --file agent-session.json --price-per-mtok 3 --monthly-runs 5000 AX Optimize — Token Reduction Tokens: 16,049 → 5,321 (−66.8%) Saved: 10,728 tokens per run Transforms: log, smart_crusher (user messages protected) Cost estimate (input tokens): at €3.00/1M tokens, 5,000 runs/month: ~€160.92/month · ~€1,931.04/year # Reduction is input-dependent: tool outputs/logs high, dense prose low.
Honestly framed: the compression engine is Headroom (open source, Apache 2.0) — we did not invent it, we integrate and operate it. It protects your actual user messages and primarily compresses the machine-generated ballast. Our contribution (the AX delta) is the cost story: tokens saved translated into euros and projected onto your volume — plus model routing and cost caps for genuine budget governance. Local-first: the context does not leave your machine via us.
Context compression
Tool outputs, logs and RAG content are condensed via the Headroom engine — user messages remain protected.
Cost projection
Saved tokens are translated into euros and projected onto your runs per month — transparently labelled as an estimate.
Model routing & caps
Cheaper models for simple steps, hard cost caps per day/task — budget governance instead of a surprise on the invoice.
Local-first
Compression runs on your machine. The context does not flow out via us — relevant with sensitive data.
- Measurable token reduction per run — typically 60–95% with tool-/log-heavy context
- Euro projection onto your volume (month/year), transparently labelled as an estimate
- Integration via MCP tool or as a proxy/library into your agent pipeline
- Local-first: context stays on your infrastructure
- Honest OSS basis: the engine is Headroom (Apache 2.0), attribution included — our value is integration, routing, cost governance
Does answer quality suffer from the compression?
How much does it really save?
Does our context leave our own infrastructure?
What is your contribution compared to the open-source project?
Related roles & solutions
Goes together
AX Optimize — does this fit your situation?
In a free initial call we clarify whether this service addresses your problem — or whether another route gets there faster. Honest assessment included.
Book an initial call