Works with OpenAI, Claude, Gemini, Grok, and more

Cut your LLM API costs by up to 80%

SemaCache is an intelligent caching proxy for LLM APIs. It returns cached responses when a semantically similar query has been seen before — saving you money on every repeated question.

Drop-in replacement — change one line
# Before — calling OpenAI directly
client = OpenAI(api_key="sk-...")

# After — just change the base URL
client = OpenAI(
    api_key="sc-your-key",
    base_url="https://www.semacache.io/api/v1"
)
Free tier included
No code changes required
Supports OpenAI, Claude, Gemini & Grok
~5ms
Exact match latency
~20ms
Semantic match latency
80%
Avg cost reduction
99.9%
Uptime SLA
Interactive Demo

Try it yourself

Explore the dashboard and see the caching pipeline in action \u2014 no signup required. All data is mocked.

Dashboard
Operational

Overview

Monitor cache performance and usage.

898.3K tokens saved ($42.86)

1.3M total \u00b7 70% from cache

70%
tokens saved
Total Requests
12,847

all time

Cache Hits
8,993

5,621 exact · 3,372 semantic

Cache Misses
3,854

LLM pass-throughs

Cost Saved
$42.86

898.3K tokens saved

Cache Hit Rate

70.0%
hit rate

8,993 / 12,847

Cache is 195\u00d7 faster

Latency

Exact5ms
Semantic22ms
Pass-through2.3s
5ms
Exact
22ms
Semantic
2.3s
Native

Daily Requests

7d
Mon
Tue
Wed
Thu
Fri
Sat
Sun
Exact
Semantic
Miss
Avg Latency
486ms

all requests

Total Tokens
1.3M

642K prompt \u00b7 642K completion

Total Spend
$18.42

without cache: $61.28

Recent Queries

Latest cache interactions

View all
QueryModelTypeLatencyTime
What is the capital of France?gpt-4o-miniEXACT4ms2:34 PM
Explain machine learning basicsgpt-4oSEMANTIC19ms2:33 PM
How to sort a list in Python?gpt-4o-miniEXACT3ms2:31 PM
Benefits of cloud computing?claude-sonnet-4NATIVE2.2s2:29 PM
Translate hello to Spanishgpt-4o-miniSEMANTIC21ms2:27 PM
SQL query to find duplicatesgpt-4oEXACT5ms2:25 PM

\ud83d\udcca Preview of the SemaCache dashboard with mock data. Sign up to see your real metrics.

Stop overpaying for LLM API calls

You're burning $438/mo on repeat queries

Up to 40% of your API calls return the same or similar answers. SemaCache intercepts them before they hit your LLM — so you only pay once.

Configure your usage

~$0.02/request at average token usage

25K
1K10K100K500K
40%
ConservativeMost teams: 30–60%Aggressive

Your monthly savings

$175

That's $2.1k/year back in your pocket

40%

cost reduction

10K

free cache hits

2d

payback period

$438/mo$272/mo

Stop leaving $2.1k/year on the table.

Pro pays for itself in 2 days. Cancel anytime. No risk.

Start Saving Now

Trusted by developers building with OpenAI, Claude, Gemini, Grok, and custom models. One line of code. Instant savings.

Features

Three tiers of intelligent caching

Every request flows through a fast pipeline: exact hash → semantic similarity → LLM passthrough. Each tier is cheaper and faster than calling the LLM directly.

Exact Match Cache

MD5 hash lookup in Redis. Identical queries return cached responses in under 5ms.

Semantic Match Cache

Gemini-powered embeddings with pgvector similarity search. Catches paraphrased queries automatically.

Multi-Provider Routing

One endpoint for OpenAI, Anthropic Claude, Google Gemini, and xAI Grok. SemaCache auto-detects the provider from the model name and routes accordingly.

Encrypted Key Storage

Store your LLM API keys securely in the dashboard. AES-256 encrypted at rest — keys never leave our servers.

Real-Time Analytics

Dashboard with cache hit rates, latency metrics, cost savings, and daily request volume per API key.

OpenAI-Compatible API

Drop-in replacement for any OpenAI SDK client. Works with Python, JavaScript, Go, and every other language.

Open Source

Official SDKs on GitHub

MIT-licensed client libraries for every major language. Zero OpenAI dependency — each SDK talks directly to the SemaCache API and parses cache metadata from response headers.

How it works

From request to response in milliseconds

01

Your app sends a request

Point your OpenAI client at SemaCache. Your app sends requests as usual — no code changes needed beyond changing the base URL.

02

Exact match check

We hash the query and check Redis. If the identical query was asked before, the cached response is returned in ~5ms.

03

Semantic similarity search

If no exact match, we embed the query with Gemini and search our pgvector index. Paraphrased queries like "What's France's capital?" match "Capital of France?" with high confidence.

04

LLM passthrough & cache

On full miss, we route to the correct provider (OpenAI, Anthropic, Gemini, or Grok based on model name), return the response, and cache it for future hits.

Live production benchmarks

Text, images, and video — all cached.

Every API call goes through the same three-tier pipeline. The first request generates and caches. Every repeat returns instantly — whether it’s a chat reply, a 4K image, or a generated video.

133×
Chat speedup
16×
Image speedup
76×
Video speedup
<1s
All cache hits

Measured end-to-end on production (Google Cloud Run), including full network round-trip. Chat: OpenAI GPT-4o Mini & Gemini 2.0 Flash. Image: OpenAI GPT Image 1 & Google Imagen 4.0. Video: Google Veo 2 & Veo 3. Same caching applies to xAI Grok and all other supported models.

Supported Models

Works with every major LLM provider

Built-in support for OpenAI, Anthropic Claude, Google Gemini, xAI Grok, Imagen, and Veo. Plus register any OpenAI-compatible endpoint as a custom model.

O

OpenAI

Chat Completions

gpt-5.6-solgpt-5.6-terragpt-5.6-lunagpt-5.5gpt-5.5-progpt-5.4gpt-5.4-minigpt-5.4-nanogpt-4.1gpt-4ogpt-4o-minio3o4-mini
A

Anthropic Claude

Chat Completions

claude-fable-5claude-opus-4-8claude-opus-4-7claude-sonnet-5claude-sonnet-4-6claude-haiku-4-5
G

Google Gemini

Chat Completions

gemini-3.5-flashgemini-3.1-pro-previewgemini-3-flash-previewgemini-3.1-flash-litegemini-2.5-progemini-2.5-flashgemini-2.5-flash-lite
X

xAI Grok

Chat Completions

grok-4.20grok-4grok-4-fastgrok-3grok-3-minigrok-3-fast
I

Image Generation

OpenAI, Google, xAI

gpt-image-1.5gpt-image-1gpt-image-1-miniimagen-4.0-generate-001imagen-4.0-ultra-generate-001imagen-4.0-fast-generate-001gemini-3-pro-image-previewgemini-3.1-flash-image-previewgrok-imagine-imagegrok-imagine-image-pro
V

Video Generation

Google Veo, xAI

veo-3.1-generate-previewveo-3.1-fast-generate-previewveo-3.1-lite-generate-previewveo-3.0-generate-001veo-3.0-fast-generate-001veo-2.0-generate-001grok-imagine-video
Pro & Enterprise

Bring your own model

Register any OpenAI-compatible endpoint — vLLM, Ollama, Together AI, Groq, Fireworks, or your own self-hosted model. SemaCache caches responses from custom models the same way it caches OpenAI, Claude, and Gemini.

  • Register via dashboard or API — set base URL, model name, and auth
  • Full three-tier caching: exact → semantic → passthrough
  • Works with any provider that speaks OpenAI-compatible format
# Register "my-llama" in Dashboard → Custom Models
# Then use it like any built-in model

from openai import OpenAI

client = OpenAI(
  api_key="sc-your-key",
  base_url="https://www.semacache.io/api/v1"
)

response = client.chat.completions.create(
  model="my-llama",
  messages=[
    {"role": "user", "content": "Hello!"}
  ]
)

Use Cases

Built for every LLM-powered workflow

If your app calls an LLM API, SemaCache can save you money. Here are the most common patterns.

💬

Customer Support Chatbots

Users ask the same questions hundreds of times: "How do I reset my password?", "What are your hours?", "How do I cancel?" SemaCache recognizes all variations and returns cached answers in milliseconds — cutting your API bill by 60–80%.

72% cache hit rate typical
📚

RAG & Knowledge Base Q&A

Retrieval-augmented generation apps often get the same questions about the same documents. SemaCache caches the LLM's synthesized answers so repeat queries skip the entire RAG pipeline — embedding lookup, context assembly, and LLM call.

~20ms vs ~3s response time
🛒

E-Commerce Product Descriptions

Generating product descriptions, summaries, or recommendations with LLMs? The same products get described over and over. Cache the output once and serve it instantly on every page load.

Save $200–$2,000/mo
🖼️

Image & Video Generation

Marketing teams regenerate the same hero images and product videos repeatedly. SemaCache caches generated media — a cached image returns in <1s vs 10–15s from the provider. Cached video returns in <1s vs 30–60s.

16–76× faster on cache hit
🔄

Internal Tools & Dashboards

Internal apps that summarize reports, translate content, or classify tickets often process the same inputs. SemaCache eliminates redundant LLM calls across your team without any code changes.

Zero code changes needed
🧪

Development & Testing

Running the same prompts during development and CI/CD burns through API credits. SemaCache returns cached responses instantly so your test suite runs in seconds, not minutes — and costs nothing on repeat runs.

95%+ hit rate in CI/CD

Pricing

Start free, scale with confidence

Every plan includes multi-provider support and encrypted key storage.

Free

For experimentation and side projects

$0forever
  • 1,000 requests / month
  • 1 API key
  • Text + image caching
  • 7-day audit logs
  • Community support
Get Started
Most Popular

Pro

For developers shipping to production

$9/mo
  • 50,000 requests / month
  • 5 API keys
  • Text + image + video caching
  • Custom model registry
  • 30-day audit logs
  • Email support
Get Started

Enterprise

For teams at scale

$39/mo
  • 500,000 requests / month
  • Unlimited API keys
  • Text + image + video caching
  • Custom model registry
  • 90-day audit logs
  • Priority support
Contact Sales

Ready to cut your LLM costs?

Get started in under a minute. No credit card required. Change one line of code and start saving.