Tutorials

Short walks through the first time you use SemaCache. The docs are the API reference. Start here if you have not made a request yet.

1. Make your first request

You need two keys. The SemaCache key (sc-) tells us who you are. The provider key (sk- for OpenAI) is the one we use when the answer is not cached yet.

  1. 1

    Create an account

    Open Sign in and choose Sign up. Confirm your email, then sign in. You land on the dashboard.

  2. 2

    Create a SemaCache API key

    Go to API Keys and create a key. Copy it once. It starts with sc-. The dashboard will not show the full key again.

  3. 3

    Save your OpenAI key

    Go to Settings and add an OpenAI provider key. SemaCache stores it encrypted and uses it only when a request misses the cache. You can also send it on each request as x-upstream-api-key instead of saving it.

  4. 4

    Install the OpenAI library and send one prompt

    pip install openai
    from openai import OpenAI
    
    client = OpenAI(
        api_key="sc-your-key",
        base_url="https://www.semacache.io/api/v1",
    )
    
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "What is the capital of France?"}],
    )
    print(response.choices[0].message.content)

    Leave out x-upstream-api-key if you saved the OpenAI key in Settings. This first call goes to OpenAI, then SemaCache stores the answer.

2. See a cache hit

Run the same prompt again and read the response header. The first call is NATIVE (a miss). The second identical call is EXACT (served from cache, no new OpenAI bill for that completion).

raw = client.chat.completions.with_raw_response.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
)
print(raw.headers["x-semcache-match-type"])
print(raw.parse().choices[0].message.content)

Run that block twice. You should see NATIVE, then EXACT. A paraphrase such as "Which city is the capital of France?" can come back as SEMANTIC while you still have embedding tokens. Header details are in the docs.

3. Call SemaCache from LiteLLM

LiteLLM is a Python library that calls many model providers through one function. You do not install anything inside SemaCache. You tell LiteLLM to send the HTTP request to SemaCache instead of to OpenAI.

pip install litellm
import litellm

response = litellm.completion(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
    api_key="sc-your-key",
    api_base="https://www.semacache.io/api/v1",
    extra_headers={"x-upstream-api-key": "sk-your-openai-key"},
)
print(response.choices[0].message.content)

Keep the openai/ prefix. LiteLLM then uses the OpenAI request format, and the name after the slash is the model SemaCache routes. Use openai/claude-haiku-4-5 or openai/gemini-2.5-flash-lite the same way. A prefix like anthropic/claude-haiku-4-5 makes LiteLLM call Anthropic itself and skip SemaCache. Omit extra_headers if that provider key is already in Settings.

4. Call SemaCache from LangChain

LangChain is a library for chains and agents. You do not install it inside SemaCache. You tell ChatOpenAI to send the HTTP request to SemaCache instead of to OpenAI.

pip install langchain-openai
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="gpt-4o-mini",
    api_key="sc-your-key",
    base_url="https://www.semacache.io/api/v1",
    default_headers={"x-upstream-api-key": "sk-your-openai-key"},
)

response = llm.invoke("What is the capital of France?")
print(response.content)

Keep using ChatOpenAI for Claude and Gemini too. Set model to claude-haiku-4-5 (and pass max_tokens=64) or gemini-2.5-flash-lite.ChatAnthropic and ChatGoogleGenerativeAI call those providers themselves and skip SemaCache. If you use init_chat_model, pass model_provider="openai". Omit default_headers if that provider key is already in Settings. Do not call set_llm_cache. That cache answers inside your process, so SemaCache never sees the second call. The answer is on response.content. TypeScript is in the docs.

5. Use Claude or Gemini

The client stays the OpenAI library. You change the provider key in Settings and the model string. SemaCache reads the model name and calls that provider on a miss.

  1. 1

    Save the provider key

    In Settings, add an Anthropic key for Claude or a Gemini key for Google. Use a standard Anthropic key (sk-ant-api03-), not an identity-linked key. One SemaCache sc- key can have several provider keys saved.

  2. 2

    Send a Claude prompt

    response = client.chat.completions.create(
        model="claude-haiku-4-5",
        max_tokens=64,
        messages=[{"role": "user", "content": "What is the capital of France?"}],
    )
    print(response.choices[0].message.content)
  3. 3

    Or send a Gemini prompt

    response = client.chat.completions.create(
        model="gemini-2.5-flash-lite",
        messages=[{"role": "user", "content": "What is the capital of France?"}],
    )
    print(response.choices[0].message.content)

    Run either call twice. The second identical prompt is EXACT, the same as OpenAI. If you did not save the key, pass it on the request with extra_headers={"x-upstream-api-key": "..."}.

6. Turn on paraphrase matching

Exact match only returns a stored answer when the request is identical. Paraphrase matching is semantic cache. Free includes 10,000 embedding tokens and 500 stored vectors a month. New accounts start on Managed.

  1. 1

    Confirm Settings is on Managed

    Open Settings. Under semantic matching, Managed means SemaCache embeds the prompt for you. Off is exact match only. Bring your own means you send a 768-number embedding on the JSON body and we do not call an embedding model.

    After the 10,000 tokens are used, a new prompt returns HTTP 402 and asks you to upgrade to Pro (2,000,000 tokens) or switch to Off. The same request, sent again unchanged, still hits the exact cache.

  2. 2

    Store one answer

    client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "What is the capital of France?"}],
    )

    That response header is NATIVE.

  3. 3

    Ask the same thing in different words

    raw = client.chat.completions.with_raw_response.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "Which city is France's capital?"}],
    )
    print(raw.headers["x-semcache-match-type"])
    print(raw.headers["x-semcache-confidence"])

    A paraphrase hit is SEMANTIC. x-semcache-confidence is the similarity score, from 0 to 1. The text is the stored answer, not a new completion. If it comes back NATIVE, the wording was too far from the original for your threshold. The default is 0.95. Lower it for that request with the header x-similarity-threshold: 0.80.

7. Stream the response

Pass stream=True the same way you would with OpenAI. SemaCache still stores the full answer. A later identical call, streamed or not, hits that same cache entry. A hit is replayed as server-sent events. A miss waits for the provider, stores the completion, then sends it as those same events.

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Name three European capitals."}],
    stream=True,
)
for chunk in stream:
    if not chunk.choices:
        continue
    text = chunk.choices[0].delta.content or ""
    print(text, end="")

Run it twice with the same messages. The second run is served from cache and still arrives as chunks your existing stream loop can print. Leave stream off when you want one JSON object back instead.

Next: the API reference for headers, cache control, and every model id.