Every team that builds semantic caching or RAG retrieval eventually asks the same question: how much are the embedding calls costing us. Priced against OpenAI, Cohere, and Voyage AI's current API rates, the honest answer is almost nothing. At 100 million tokens a month, embedding generation runs $1 to $13 in list price depending on the model, and a 2026 paper on Redis-backed semantic caching found the technique cuts LLM API calls by up to 68.8%, with positive-hit accuracy above 97%. That is the number most teams stop at when they scope the cost of adding a vector layer to their stack.
It is also the wrong number to stop at. The bill that actually grows with usage is not the embedding call, it is what that call produces: a vector that has to live somewhere, get searched, and get replicated for uptime. The same 100 million embeddings that cost $13 to generate can occupy anywhere from 19 GB to 1.23 TB on disk, a 65x spread driven entirely by a dimension and precision choice made once, usually by whoever copied the first code sample they found, and rarely revisited once it is in production.
Where the tax actually lands
Run the arithmetic on the two models most teams reach for first. OpenAI's text-embedding-3-small costs $0.02 per million tokens and returns a 1536-dimension vector. text-embedding-3-large costs $0.13 per million tokens, 6.5 times more, and returns 3072 dimensions by default. Neither number moves the needle on a monthly bill: at 100 million tokens, that is $2 versus $13.
Storage is a different shape of cost entirely, because it does not reset every billing cycle, it accumulates. A float32 vector costs 4 bytes per dimension. At 100 million vectors, a 3072-dimension store works out to roughly 1.23 terabytes; truncate the same model to 1536 dimensions and it is 614 GB; drop to a 1024-dimension model like Cohere Embed v3 and it is 410 GB; quantize a 1536-dimension store to 1 bit per dimension and it is about 19 GB. That last number is not a rounding error, it is a 32x reduction from the float32 version at the same dimension count, before a managed vector database's own per-gigabyte markup gets applied on top of whichever figure a team landed on by default.
| Model | Price | Dimensions | Storage per 100M vectors (float32) |
|---|---|---|---|
| OpenAI text-embedding-3-small | $0.02/M tokens | 1536, truncatable to 256 | ~614 GB at 1536, ~100 GB at 256 |
| OpenAI text-embedding-3-large | $0.13/M tokens | 3072, truncatable to 1536 | ~1.23 TB at 3072, ~614 GB at 1536 |
| Cohere Embed v3 | ~$0.10/M tokens | 1024, binary-compressible | ~410 GB float32, ~13 GB binary-quantized |
None of this shows up on the line item labeled "embeddings" in a cloud bill. It shows up on the vector database's storage and compute invoice, weeks after the embedding model got picked without much thought, because the API call itself was cheap enough that nobody thought picking the biggest model available carried a cost.
What the research says about right-sizing this layer
The GPT Semantic Cache paper (Regmi and Pun, arXiv:2411.05276) is a useful reference point for what a well-tuned cache actually buys: on a Redis-backed semantic cache matching new queries against previously answered ones, it cut LLM API calls by up to 68.8% across query categories, with cache hit rates ranging from 61.6% to 68.8% and positive-hit accuracy exceeding 97%, meaning the cache rarely served a stale or wrong answer when it decided to serve one at all. None of that result depended on using the largest embedding model available, it depended on the similarity threshold and matching logic being tuned for the traffic pattern.
A second 2026 paper, from a team at Redis and Virginia Tech (Gill et al., arXiv:2504.02268), tests the assumption directly: does a smaller embedding model hurt cache accuracy. Their answer is that a smaller, domain-specific model, fine-tuned for as little as one epoch on a synthetically generated dataset built for the task, can match or surpass a larger general-purpose embedding model's precision and recall for semantic cache matching. The implication for the cost stack above is direct: the model doing the expensive part of this job, the general-purpose flagship embedding model, is frequently not the right tool for a task as narrow as "is this new prompt close enough to one we already answered."
A four-step framework for sizing the embedding layer
- Split cache-key embeddings from retrieval embeddings. A semantic cache is answering a narrower question than RAG retrieval, is this prompt close enough to one already answered, versus which of these thousand documents is most relevant. Use a small or truncated model for the first job and reserve a larger model for the second, where recall quality has a real cost when it is wrong.
- Truncate before you reach for a bigger model. OpenAI's Matryoshka-trained v3 models accept a
dimensionsparameter that returns a shorter vector from the same model, at a proportional cut in storage and search cost, without a separate model to manage or a fine-tuning job to run.
- Quantize what you store, not just what you compute. Binary or int8 quantization on the stored vector, independent of what dimension the embedding call itself returned, is where the 32x storage reduction in the table above actually comes from. Most vector databases support this as a configuration flag, not a data migration.
- Batch what does not need to happen in real time. Backfilling an existing corpus, or re-embedding after a model swap, is exactly what the batch embedding APIs most providers offer, typically at a 50% discount, exist for. Reserve synchronous calls for embeddings that gate a live request.
def get_cache_key_embedding(text: str) -> list[float]:
# Narrow question, cheap model, truncated dimensions.
return openai.embeddings.create(
model="text-embedding-3-small",
input=text,
dimensions=256,
).data[0].embedding
def get_retrieval_embedding(chunk: str) -> list[float]:
# Broader question, keep the recall the larger model buys.
return openai.embeddings.create(
model="text-embedding-3-large",
input=chunk,
).data[0].embedding
def backfill_corpus(chunks: list[str]) -> str:
# Offline job, batch API, half the price, no latency budget to protect.
return openai.batches.create(
endpoint="/v1/embeddings",
input_file_id=upload_batch_file(chunks, model="text-embedding-3-large"),
completion_window="24h",
)
Where Nadir fits
Nadir does not run a semantic response cache. A cache keyed on prompt similarity can hand one tenant's completion to another, so if you want the savings this post describes, run the cache in your own application, scoped to your own users, where you control the threshold and the embedding model. Nadir does not decide how you store or embed a RAG corpus either; that stays a separate layer. What Nadir does keep intact is provider prompt caching. Requests that reach a model are still routed against the configured quality floor, and a configured, priced benchmark adds nadir_metadata.benchmark_comparison.savings_usd to complete non-streaming responses.
Related reading
- Semantic caching eliminates 30 to 50% of your LLM API calls. Most teams have never implemented it.
- Your RAG pipeline fetches 20 chunks per query. The model reads 3. The other 17 are billed in full.
- Gemini 2.5 Pro has a 1M token context window. Loading a 500-page document costs $1.00 per query. Your RAG pipeline costs $0.007.
- Prompt caching's write tax
Sources: Regmi, S. and Pun, C.P., ["GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching"](https://arxiv.org/abs/2411.05276), arXiv:2411.05276, 2024. Gill, W. et al., ["Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data"](https://arxiv.org/abs/2504.02268), Redis and Virginia Tech, arXiv:2504.02268, 2025. OpenAI, Cohere, and Voyage AI API pricing pages, August 2026. Storage figures are illustrative arithmetic (dimensions x 4 bytes x vector count for float32, dimensions / 8 for binary quantization), not a vendor-published benchmark.