The Embedding Tax

Every team building semantic caching or RAG retrieval asks what the embedding calls cost, and at 100 million tokens a month it is $1 to $13 in list price. Priced against OpenAI, Cohere, and Voyage AI's current API rates, that is almost nothing, and a 2026 paper on Redis-backed semantic caching (Regmi and Pun, arXiv:2411.05276) found the technique cuts LLM API calls by up to 68.8%, with hit rates of 61.6 to 68.8% and positive-hit accuracy above 97%. The bill that actually grows with usage is the one nobody scopes upfront: vector storage. The same 100 million embeddings that cost $13 to generate can occupy anywhere from 19 GB to 1.23 TB on disk depending on dimension count and quantization, a 65x spread driven entirely by a model choice made once and rarely revisited. A second 2026 paper from Redis and Virginia Tech (Gill et al., arXiv:2504.02268) shows a smaller, domain-fine-tuned embedding model can match or beat a general-purpose one on cache-matching precision and recall, at a fraction of the footprint. Here's where the embedding tax actually lands, the arithmetic behind it, and a four-step framework for sizing this layer instead of defaulting to the biggest model on the pricing page.

Published 2026-08-18, updated 2026-09-22, by Dor Amir on the Nadir blog.

Filed under Caching.

Every team that builds semantic caching or RAG retrieval eventually asks the same question: how much are the embedding calls costing us. Priced against OpenAI, Cohere, and Voyage AI's current API rates, the honest answer is almost nothing. At 100 million tokens a month, embedding generation runs $1 to $13 in list price depending on the model, and a 2026 paper on Redis-backed semantic caching found the technique cuts LLM API calls by up to 68.8%, with positive-hit accuracy above 97%. That is the number most teams stop at when they scope the cost of adding a vector layer to their stack.

It is also the wrong number to stop at. The bill that actually grows with usage is not the embedding call, it is what that call produces: a vector that has to live somewhere, get searched, and get replicated for uptime. The same 100 million embeddings that cost $13 to generate can occupy anywhere from 19 GB to 1.23 TB on disk, a 65x spread driven entirely by a dimension and precision choice made once, usually by whoever copied the first code sample they found, and rarely revisited once it is in production.

Where the tax actually lands

Run the arithmetic on the two models most teams reach for first. OpenAI's text-embedding-3-small costs $0.02 per million tokens and returns a 1536-dimension vector. text-embedding-3-large costs $0.13 per million tokens, 6.5 times more, and returns 3072 dimensions by default. Neither number moves the needle on a monthly bill: at 100 million tokens, that is $2 versus $13.

Storage is a different shape of cost entirely, because it does not reset every billing cycle, it accumulates. A float32 vector costs 4 bytes per dimension. At 100 million vectors, a 3072-dimension store works out to roughly 1.23 terabytes; truncate the same model to 1536 dimensions and it is 614 GB; drop to a 1024-dimension model like Cohere Embed v3 and it is 410 GB; quantize a 1536-dimension store to 1 bit per dimension and it is about 19 GB. That last number is not a rounding error, it is a 32x reduction from the float32 version at the same dimension count, before a managed vector database's own per-gigabyte markup gets applied on top of whichever figure a team landed on by default.

The embedding call is the cheap part. At 100M tokens/month, list-price embedding costs run $1 to $13 across four models. Storing the resulting 100M vectors runs from 19 GB (1536d, binary-quantized) to 1.23 TB (3072d, float32), a 65x spread from dimension and precision choice alone.
The embedding call is the cheap part. At 100M tokens/month, list-price embedding costs run $1 to $13 across four models. Storing the resulting 100M vectors runs from 19 GB (1536d, binary-quantized) to 1.23 TB (3072d, float32), a 65x spread from dimension and precision choice alone.
ModelPriceDimensionsStorage per 100M vectors (float32)
OpenAI text-embedding-3-small$0.02/M tokens1536, truncatable to 256~614 GB at 1536, ~100 GB at 256
OpenAI text-embedding-3-large$0.13/M tokens3072, truncatable to 1536~1.23 TB at 3072, ~614 GB at 1536
Cohere Embed v3~$0.10/M tokens1024, binary-compressible~410 GB float32, ~13 GB binary-quantized

None of this shows up on the line item labeled "embeddings" in a cloud bill. It shows up on the vector database's storage and compute invoice, weeks after the embedding model got picked without much thought, because the API call itself was cheap enough that nobody thought picking the biggest model available carried a cost.

What the research says about right-sizing this layer

The GPT Semantic Cache paper (Regmi and Pun, arXiv:2411.05276) is a useful reference point for what a well-tuned cache actually buys: on a Redis-backed semantic cache matching new queries against previously answered ones, it cut LLM API calls by up to 68.8% across query categories, with cache hit rates ranging from 61.6% to 68.8% and positive-hit accuracy exceeding 97%, meaning the cache rarely served a stale or wrong answer when it decided to serve one at all. None of that result depended on using the largest embedding model available, it depended on the similarity threshold and matching logic being tuned for the traffic pattern.

A second 2026 paper, from a team at Redis and Virginia Tech (Gill et al., arXiv:2504.02268), tests the assumption directly: does a smaller embedding model hurt cache accuracy. Their answer is that a smaller, domain-specific model, fine-tuned for as little as one epoch on a synthetically generated dataset built for the task, can match or surpass a larger general-purpose embedding model's precision and recall for semantic cache matching. The implication for the cost stack above is direct: the model doing the expensive part of this job, the general-purpose flagship embedding model, is frequently not the right tool for a task as narrow as "is this new prompt close enough to one we already answered."

A four-step framework for sizing the embedding layer

  1. Split cache-key embeddings from retrieval embeddings. A semantic cache is answering a narrower question than RAG retrieval, is this prompt close enough to one already answered, versus which of these thousand documents is most relevant. Use a small or truncated model for the first job and reserve a larger model for the second, where recall quality has a real cost when it is wrong.
  1. Truncate before you reach for a bigger model. OpenAI's Matryoshka-trained v3 models accept a dimensions parameter that returns a shorter vector from the same model, at a proportional cut in storage and search cost, without a separate model to manage or a fine-tuning job to run.
  1. Quantize what you store, not just what you compute. Binary or int8 quantization on the stored vector, independent of what dimension the embedding call itself returned, is where the 32x storage reduction in the table above actually comes from. Most vector databases support this as a configuration flag, not a data migration.
  1. Batch what does not need to happen in real time. Backfilling an existing corpus, or re-embedding after a model swap, is exactly what the batch embedding APIs most providers offer, typically at a 50% discount, exist for. Reserve synchronous calls for embeddings that gate a live request.
def get_cache_key_embedding(text: str) -> list[float]:
    # Narrow question, cheap model, truncated dimensions.
    return openai.embeddings.create(
        model="text-embedding-3-small",
        input=text,
        dimensions=256,
    ).data[0].embedding

def get_retrieval_embedding(chunk: str) -> list[float]:
    # Broader question, keep the recall the larger model buys.
    return openai.embeddings.create(
        model="text-embedding-3-large",
        input=chunk,
    ).data[0].embedding

def backfill_corpus(chunks: list[str]) -> str:
    # Offline job, batch API, half the price, no latency budget to protect.
    return openai.batches.create(
        endpoint="/v1/embeddings",
        input_file_id=upload_batch_file(chunks, model="text-embedding-3-large"),
        completion_window="24h",
    )

Where Nadir fits

Nadir does not run a semantic response cache. A cache keyed on prompt similarity can hand one tenant's completion to another, so if you want the savings this post describes, run the cache in your own application, scoped to your own users, where you control the threshold and the embedding model. Nadir does not decide how you store or embed a RAG corpus either; that stays a separate layer. What Nadir does keep intact is provider prompt caching. Requests that reach a model are still routed against the configured quality floor, and a configured, priced benchmark adds nadir_metadata.benchmark_comparison.savings_usd to complete non-streaming responses.

Related reading


Sources: Regmi, S. and Pun, C.P., ["GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching"](https://arxiv.org/abs/2411.05276), arXiv:2411.05276, 2024. Gill, W. et al., ["Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data"](https://arxiv.org/abs/2504.02268), Redis and Virginia Tech, arXiv:2504.02268, 2025. OpenAI, Cohere, and Voyage AI API pricing pages, August 2026. Storage figures are illustrative arithmetic (dimensions x 4 bytes x vector count for float32, dimensions / 8 for binary quantization), not a vendor-published benchmark.

What Nadir is

Nadir is an LLM router. Nadir sizes every prompt and routes it to the cheapest model that still clears your quality bar. A trained pre-classifier scores each prompt in under 10 ms, with no LLM call in the routing step.

Nadir runs two ways. The decision API returns a model, reasoning-effort, cache, context, and policy recommendation without calling a model provider, beside the gateway you already run. That is how a shadow-mode evaluation works, and its projected savings stay advisory. The OpenAI compatible managed proxy executes the route, and migration is a two-line change: point the base URL at api.getnadir.com and set model to auto. On that path an optional verifier can score a complete non-streaming answer and escalate to a stronger model when it misses the configured bar. Streaming bypasses post-generation verification. BYOK is supported on every tier.

For coding agents, Nadir connects to Codex, Claude Code, or Cursor and recommends a model tier for delegated work. The agent decides whether to hand the task off and uses its own configured models, so Nadir needs no proxy and no provider keys on that path.

What the numbers are, and what they are not

Nadir publishes each evaluation with its scope. On checkable code, run-check-escalate solved 392 of 395 common HumanEval and MBPP problems (99.2%), graded by running the canonical tests; that applies only to tasks with runnable deterministic tests. Nadir-Tumbler posts an arena_score of 72.3 on RouterArena's public scorer, 5th of 23 routers, which measures the routing decision on RouterArena's own model pool. A reference-assisted RouterBench evaluation over 11,420 held-out triples produced a 60% lower projected cost than always-Opus with about 98% retained quality and a 1.7% catastrophic-route rate. That experiment gave the verifier the expensive-model reference answer, which production does not have, so it is a research ceiling and not the deployed path.

None of these is a production guarantee, a universal savings rate, or a forecast for any particular workload. Customer savings are reported from measured execution against a declared baseline, and customer quality only from outcome-labelled traffic. Projected savings and realized savings are separate artifacts and are never blended.

Design-partner program

Three rungs, picked by risk appetite. Rung 0 Shadow runs advisory decision calls alongside live traffic and returns a projected receipt, with nothing in the request path changed. Rung 1 Hosted is the two-line swap on a production slice and returns a realized receipt. Rung 2 On-prem is a supervised six-week proof of concept inside the partner's VPC, where no prompt, response, or usage reaches Nadir. The commitments are the same at every rung. Apply for a rung directly: Rung 0 Shadow, Rung 1 Hosted, or Rung 2 On-prem. Not sure which fits? Start at getnadir.com/contact/?reason=design-partner.

Licensing

NadirClaw is the self-hosted core, source-available under the PolyForm Noncommercial License. Source-available is the correct label; NadirClaw is not open source. Nadir Route's hosted plan has no base fee and charges a variable fee only on measured savings from requests Nadir executed.

Pages on this site

Machine-readable summaries of this site: llms.txt and llms-full.txt.