To compress your prompt, LLMLingua first runs it through a 7-billion-parameter model.
That's not a criticism, it's the architecture. LLMLingua, the prompt-compression tool most cost-optimization checklists point to first, works by scoring every token's importance with a small causal language model, GPT2-small or LLaMA-7B, then dropping the low-information ones before the compressed prompt goes to whatever model you're actually trying to save money on. It's a real technique with real numbers behind it: Jiang et al.'s original paper reports up to 20x compression on in-context learning and reasoning tasks, a 1.7-5.7x end-to-end latency speedup on GSM8K, and only a 1.5-point performance loss at maximum compression. LLMLingua-2, published at ACL 2024, swapped the causal LLaMA-7B scorer for a smaller, distilled BERT-base encoder trained via GPT-4 data distillation, faster and lighter, but still a model doing a forward pass on every prompt that comes through.
Here's the question a July 2026 paper actually sat down and tested: does compressing a prompt require a model at all?
What "the compressor has its own bill" actually means
Every LM-based compressor, LLaMA-7B, BERT-base, doesn't matter which, has to run somewhere. In a hosted deployment that's an extra API call to a compression-serving endpoint, extra latency stacked in front of the request you're trying to speed up, and extra infrastructure to keep warm regardless of traffic. At high query volume, that overhead is a real, measurable cost that most compression writeups leave out of the arithmetic entirely. The pitch is always "cut your input tokens 3-5x"; the pitch rarely mentions that cutting them required standing up and serving a second model first.
The paper that tried removing the model entirely
Ma, Feng, Chersoni, and Chen, "Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors," arXiv:2607.25335 (submitted July 28, 2026), frames the question plainly in its own abstract: "It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules." Instead of learning token importance from a model, the authors run an offline evolutionary search over lexical, syntactic, semantic, and discourse-level rule seeds, essentially breeding combinations of grammar-level heuristics against held-out data until a competitive rule set emerges. The search happens once, offline. What ships to production is a deterministic set of rules: no LM forward pass, no GPU, CPU-only processing at deployment time.
Tested across short passages, multi-document reasoning, and dialogue-memory QA datasets, the evolved rule-based compressors reach "performance similar to that of recent advanced prompt-compression strategies," strongest at light-to-moderate compression and degrading, predictably, as the compression ratio climbs. The paper also documents a shift in strategy as compression gets more aggressive: at light compression the rules prune individual tokens, at heavier compression they switch to extracting whole sentences instead, because past a certain point, deciding which words to keep needs more context than a token-level rule can see.
| Compressor | Mechanism | Model inference at deployment | Reported compression | Deployment cost profile |
|---|---|---|---|---|
| LLMLingua | LLaMA-7B / GPT2-small importance scoring | Yes — 7B-parameter forward pass per prompt | Up to 20x, 1.5-pt loss at max ratio | GPU-hosted, adds its own latency |
| LLMLingua-2 | Distilled BERT-base token classifier | Yes — ~110M-parameter forward pass | Comparable ratio, 3-6x faster than v1 | Lighter GPU/CPU, still a model call |
| Linguistic rules (2607.25335) | Evolutionary-search grammar rules | No — deterministic rules only | Comparable at light-moderate ratios | CPU-only, zero marginal inference cost |
Figures as reported in the cited papers. LLMLingua-2's exact compression ratio and the linguistic-rule paper's numeric accuracy-retention figures were not published in the abstract; see sources for the full text.
Where each approach actually earns its place
None of this makes LM-based compression obsolete, and the paper doesn't claim it does. It makes the choice between them a real engineering decision instead of a default:
Reach for rule-based compression first, on the traffic where it's free. Chat history, log dumps, tool output, boilerplate documentation, anything that's plain, well-structured natural-language filler is exactly the case linguistic rules handle at light-to-moderate compression without a GPU in the loop. It costs nothing to run, so there's no ratio it has to clear to be worth it, unlike a model-based compressor whose own inference has to be paid for before the savings start counting.
Reserve LM-based compression for the prompts where token importance is genuinely context-dependent. Dense technical text, code, anything where "informative" can't be decided by grammar alone, is where LLMLingua's learned scoring earns its keep, and where higher compression ratios (10x, 20x) can outrun the cost of running the compressor itself. The break-even isn't compression ratio in isolation, it's compression ratio against what the compressor itself costs to run at your query volume.
Either way, compression is a bet that the model didn't need what got cut, and nothing about the compression step proves the bet paid off. A rule-based compressor and a learned one can both quietly drop the one clause the answer depended on; the rule-based one just does it for free. The compression ratio on the input side tells you nothing about whether the output on the other side still holds up.
A four-step way to actually decide
- Profile what's in the prompt before picking a compressor, not after. System prompt, retrieved context, tool schemas, and conversation history behave differently under compression; a single blanket compressor applied to all of it is how teams end up paying for a GPU-hosted scorer on tokens that were plain filler the whole time.
- Default to the free tier. Apply deterministic, rule-based pruning to natural-language filler first: it's zero marginal cost, and per this paper's own findings, it's already competitive with learned compressors at light-to-moderate ratios.
- Bring in a learned compressor only where the math works out, dense or technical spans where a higher ratio is achievable and that ratio's savings clear the cost of running the compressor itself at your actual request volume.
- Verify the output, because no compressor checks its own work. This is the step both compression tiers skip: neither a free rule-based cut nor a paid model-based one confirms the answer that comes back still holds up. Nadir's calibrated verifier scores the generated answer itself, after compression and routing have already happened, and escalates automatically when the score says something got lost, the one check a token count on the input side can't perform for you.
Compression and routing already compound rather than compete, and Nadir's own Context Optimize layer is built the way this paper's own conclusion points: as a deterministic transform rather than a per-query model call, which is also what keeps it from breaking a provider's prompt cache the way query-aware compressors can. None of that replaces outcome measurement. Decision-only calls return advisory economics without executing a model; managed non-streaming responses can optionally enter the verifier cascade, while streaming responses cannot. Start free and inspect the decision receipt on your own traffic.
Sources: [Ma, Feng, Chersoni, and Chen, "Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors," arXiv:2607.25335 (July 2026)](https://arxiv.org/abs/2607.25335). [Jiang, Wu, Lin, Yang, and Qiu, "LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models," arXiv:2310.05736 (2023)](https://arxiv.org/abs/2310.05736). [Pan, Wu, Jiang, Xia, Luo, Zhang, Lin, Rühle, Yang, Lin, Zhao, Qiu, and Zhang, "LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression," arXiv:2403.12968 (ACL 2024 Findings)](https://arxiv.org/abs/2403.12968). [Microsoft Research, "LLMLingua: Innovating LLM efficiency with prompt compression"](https://www.microsoft.com/en-us/research/blog/llmlingua-innovating-llm-efficiency-with-prompt-compression/). Figures current as of August 2026 and subject to change as the cited paper completes peer review.