The Linguist's Discount

LLMLingua, the prompt compressor most cost checklists point to first, runs every prompt through a 7-billion-parameter model before the model you pay for. LLMLingua-2 swapped that LLaMA-7B scorer for a lighter, distilled BERT-base encoder, still a forward pass on every call. A July 2026 paper, arXiv:2607.25335, asks the question neither tool stopped to answer: does compressing a prompt require a model at all? Its answer is an offline evolutionary search over lexical, syntactic, semantic, and discourse-level rules that ships as a deterministic, CPU-only compressor, zero GPU forward passes at deployment, reaching performance comparable to recent learned compressors at light-to-moderate ratios. Here's what LM-based compression actually costs once you count the compressor's own bill, what the rule-based alternative trades away to avoid it, and a four-step framework for picking between them instead of defaulting to whichever one a blog post mentioned first.

Published 2026-08-21 by Dor Amir on the Nadir blog.

Filed under Context & Compression.

To compress your prompt, LLMLingua first runs it through a 7-billion-parameter model.

That's not a criticism, it's the architecture. LLMLingua, the prompt-compression tool most cost-optimization checklists point to first, works by scoring every token's importance with a small causal language model, GPT2-small or LLaMA-7B, then dropping the low-information ones before the compressed prompt goes to whatever model you're actually trying to save money on. It's a real technique with real numbers behind it: Jiang et al.'s original paper reports up to 20x compression on in-context learning and reasoning tasks, a 1.7-5.7x end-to-end latency speedup on GSM8K, and only a 1.5-point performance loss at maximum compression. LLMLingua-2, published at ACL 2024, swapped the causal LLaMA-7B scorer for a smaller, distilled BERT-base encoder trained via GPT-4 data distillation, faster and lighter, but still a model doing a forward pass on every prompt that comes through.

Here's the question a July 2026 paper actually sat down and tested: does compressing a prompt require a model at all?

What "the compressor has its own bill" actually means

Every LM-based compressor, LLaMA-7B, BERT-base, doesn't matter which, has to run somewhere. In a hosted deployment that's an extra API call to a compression-serving endpoint, extra latency stacked in front of the request you're trying to speed up, and extra infrastructure to keep warm regardless of traffic. At high query volume, that overhead is a real, measurable cost that most compression writeups leave out of the arithmetic entirely. The pitch is always "cut your input tokens 3-5x"; the pitch rarely mentions that cutting them required standing up and serving a second model first.

Bar chart comparing what runs, per request, before a compressed prompt reaches the target model. LLMLingua's original compressor uses a roughly 7-billion-parameter LLaMA-7B model, its own GPU forward pass. LLMLingua-2 swapped in a smaller, roughly 110-million-parameter distilled BERT-base encoder, still a model call. A July 2026 paper's evolutionary-search linguistic-rule compressor uses zero parameters: CPU-only regex and part-of-speech rules, no forward pass at all.
Bar chart comparing what runs, per request, before a compressed prompt reaches the target model. LLMLingua's original compressor uses a roughly 7-billion-parameter LLaMA-7B model, its own GPU forward pass. LLMLingua-2 swapped in a smaller, roughly 110-million-parameter distilled BERT-base encoder, still a model call. A July 2026 paper's evolutionary-search linguistic-rule compressor uses zero parameters: CPU-only regex and part-of-speech rules, no forward pass at all.

The paper that tried removing the model entirely

Ma, Feng, Chersoni, and Chen, "Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors," arXiv:2607.25335 (submitted July 28, 2026), frames the question plainly in its own abstract: "It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules." Instead of learning token importance from a model, the authors run an offline evolutionary search over lexical, syntactic, semantic, and discourse-level rule seeds, essentially breeding combinations of grammar-level heuristics against held-out data until a competitive rule set emerges. The search happens once, offline. What ships to production is a deterministic set of rules: no LM forward pass, no GPU, CPU-only processing at deployment time.

Tested across short passages, multi-document reasoning, and dialogue-memory QA datasets, the evolved rule-based compressors reach "performance similar to that of recent advanced prompt-compression strategies," strongest at light-to-moderate compression and degrading, predictably, as the compression ratio climbs. The paper also documents a shift in strategy as compression gets more aggressive: at light compression the rules prune individual tokens, at heavier compression they switch to extracting whole sentences instead, because past a certain point, deciding which words to keep needs more context than a token-level rule can see.

CompressorMechanismModel inference at deploymentReported compressionDeployment cost profile
LLMLinguaLLaMA-7B / GPT2-small importance scoringYes — 7B-parameter forward pass per promptUp to 20x, 1.5-pt loss at max ratioGPU-hosted, adds its own latency
LLMLingua-2Distilled BERT-base token classifierYes — ~110M-parameter forward passComparable ratio, 3-6x faster than v1Lighter GPU/CPU, still a model call
Linguistic rules (2607.25335)Evolutionary-search grammar rulesNo — deterministic rules onlyComparable at light-moderate ratiosCPU-only, zero marginal inference cost

Figures as reported in the cited papers. LLMLingua-2's exact compression ratio and the linguistic-rule paper's numeric accuracy-retention figures were not published in the abstract; see sources for the full text.

Where each approach actually earns its place

None of this makes LM-based compression obsolete, and the paper doesn't claim it does. It makes the choice between them a real engineering decision instead of a default:

Reach for rule-based compression first, on the traffic where it's free. Chat history, log dumps, tool output, boilerplate documentation, anything that's plain, well-structured natural-language filler is exactly the case linguistic rules handle at light-to-moderate compression without a GPU in the loop. It costs nothing to run, so there's no ratio it has to clear to be worth it, unlike a model-based compressor whose own inference has to be paid for before the savings start counting.

Reserve LM-based compression for the prompts where token importance is genuinely context-dependent. Dense technical text, code, anything where "informative" can't be decided by grammar alone, is where LLMLingua's learned scoring earns its keep, and where higher compression ratios (10x, 20x) can outrun the cost of running the compressor itself. The break-even isn't compression ratio in isolation, it's compression ratio against what the compressor itself costs to run at your query volume.

Either way, compression is a bet that the model didn't need what got cut, and nothing about the compression step proves the bet paid off. A rule-based compressor and a learned one can both quietly drop the one clause the answer depended on; the rule-based one just does it for free. The compression ratio on the input side tells you nothing about whether the output on the other side still holds up.

A four-step way to actually decide

  1. Profile what's in the prompt before picking a compressor, not after. System prompt, retrieved context, tool schemas, and conversation history behave differently under compression; a single blanket compressor applied to all of it is how teams end up paying for a GPU-hosted scorer on tokens that were plain filler the whole time.
  2. Default to the free tier. Apply deterministic, rule-based pruning to natural-language filler first: it's zero marginal cost, and per this paper's own findings, it's already competitive with learned compressors at light-to-moderate ratios.
  3. Bring in a learned compressor only where the math works out, dense or technical spans where a higher ratio is achievable and that ratio's savings clear the cost of running the compressor itself at your actual request volume.
  4. Verify the output, because no compressor checks its own work. This is the step both compression tiers skip: neither a free rule-based cut nor a paid model-based one confirms the answer that comes back still holds up. Nadir's calibrated verifier scores the generated answer itself, after compression and routing have already happened, and escalates automatically when the score says something got lost, the one check a token count on the input side can't perform for you.

Compression and routing already compound rather than compete, and Nadir's own Context Optimize layer is built the way this paper's own conclusion points: as a deterministic transform rather than a per-query model call, which is also what keeps it from breaking a provider's prompt cache the way query-aware compressors can. None of that replaces outcome measurement. Decision-only calls return advisory economics without executing a model; managed non-streaming responses can optionally enter the verifier cascade, while streaming responses cannot. Start free and inspect the decision receipt on your own traffic.


Sources: [Ma, Feng, Chersoni, and Chen, "Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors," arXiv:2607.25335 (July 2026)](https://arxiv.org/abs/2607.25335). [Jiang, Wu, Lin, Yang, and Qiu, "LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models," arXiv:2310.05736 (2023)](https://arxiv.org/abs/2310.05736). [Pan, Wu, Jiang, Xia, Luo, Zhang, Lin, Rühle, Yang, Lin, Zhao, Qiu, and Zhang, "LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression," arXiv:2403.12968 (ACL 2024 Findings)](https://arxiv.org/abs/2403.12968). [Microsoft Research, "LLMLingua: Innovating LLM efficiency with prompt compression"](https://www.microsoft.com/en-us/research/blog/llmlingua-innovating-llm-efficiency-with-prompt-compression/). Figures current as of August 2026 and subject to change as the cited paper completes peer review.

More on context & compression

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir Auto is a terminal launcher for Claude Code or Codex that asks Nadir for the main-session model before each turn, with your configured model as the fallback and cost ceiling. Delegation integrations for Claude Code, Codex, or Cursor instead recommend a model tier for subagent work, and the agent decides. Either way inference runs on your own provider account, so Nadir holds no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.