Bifrost

Provider prompt caching

Cache controls, token accounting, provider support, and diagnosis.

Provider prompt caching reuses computation at the model provider. It is separate from Bifrost's Redis response cache. A streamed request can use provider caching even though it is ineligible for the Redis response cache. cacheHit and the response-cache headers describe Redis; use token counters to measure provider caching.

Supported paths

The public endpoint and upstream transport are independent. For example, a client using /v1/chat/completions can reach Azure through /responses. Bifrost normalizes provider usage before rendering JSON, SSE, or Responses WebSocket events. Public model aliases are replaced with the selected deployment's upstream model name.

AdapterImplemented text transportCache handling
openaiResponses, Chat Completions; native Responses WebSocketKey, policy, retention and content breakpoints; read/write counters
azureopenaiResponses, Chat CompletionsSame controls; deployment determines model and feature support
azurefoundryChat CompletionsOpenAI-shaped cache fields; support depends on the deployed model
openaicompatibleChat CompletionsStandard fields and documented usage aliases; backend support varies
anthropicMessagesAutomatic and block-level cache_control; disjoint usage and write lifetimes
googleaistudiogenerateContentcachedContentTokenCount; external extra_body.cachedContent references
deepseekChat Completions, ResponsesNested cache counters; Chat also accepts prompt_cache_hit_tokens
moonshotChat CompletionsTop-level usage.cached_tokens mapped to cache reads
minimaxChat CompletionsNested cached-token usage
zaiChat CompletionsNested cached-token usage
vercelResponses, Chat CompletionsUsage normalization and forwarding of providerOptions, including caching options
openrouterNo text handler in this gatewayRegistered for reranking; no prompt-cache text path

This lists gateway support, not a guarantee that every model exposes every feature. Text caching controls are not attached to embeddings, reranking, standalone images, transcription, or video jobs. The Google Interactions implementation in this repository is for video, not a second text route.

OpenAI and Azure

Bifrost preserves prompt_cache_key, prompt_cache_options, prompt_cache_retention, and eligible content prompt_cache_breakpoint markers. Marked instruction blocks remain ordered blocks when translating Chat to Responses. File materialization preserves their markers too.

For GPT-5.6, mode: "explicit" with no breakpoints disables caching. Older models use the legacy retention field and can reject the new policy. Azure PTU-M does not support explicit breakpoints. The documented minimum eligible prefix is 1,024 tokens. See Azure prompt caching and OpenAI prompt caching.

With the standard AI SDK OpenAI provider, keep the openai options namespace even when changing baseURL and using a public alias:

import { createOpenAI } from "@ai-sdk/openai";
import { streamText } from "ai";

const gateway = createOpenAI({
  baseURL: process.env.GATEWAY_BASE_URL,
  apiKey: process.env.GATEWAY_API_KEY,
});

const result = streamText({
  model: gateway("my-public-model"),
  providerOptions: {
    openai: {
      promptCacheKey: "tenant:example:manual-v1",
      promptCacheOptions: { mode: "explicit", ttl: "30m" },
    },
  },
  messages: [{
    role: "user",
    content: [{
      type: "text",
      text: stableManual,
      providerOptions: {
        openai: { promptCacheBreakpoint: { mode: "explicit" } },
      },
    }, { type: "text", text: question }],
  }],
});

const usage = await result.usage;
console.log(usage.inputTokenDetails.cacheReadTokens);
console.log(usage.inputTokenDetails.cacheWriteTokens);

Use a stable prefix long enough for the selected model. SDK versions differ: verify that the version in the application emits these options and exposes these counters. generateText uses the same options without streaming. With multi-step tool loops, inspect total usage as well as each step; a single step's zero does not describe the whole interaction.

The earlier extra_body form of the two OpenAI policy fields remains accepted. Supplying the same control both directly and inside extra_body is rejected. A native policy that cannot be represented by Anthropic or Google is rejected rather than silently removed. Audio breakpoints require Chat Completions; Responses does not support that combination.

Other providers

For Anthropic, /v1/messages accepts top-level automatic cache_control and explicit markers on system/content/tool blocks, including tool_use and tool_result. Native markers require a Messages upstream. OpenAI-shaped clients targeting Anthropic can request automatic caching through extra_body: { cache_control: { type: "ephemeral" } }. TTLs of 5m and 1h have different write prices. Bifrost merges cumulative initial/final stream usage, adds the three disjoint input buckets, and retains native cache_creation detail. See Anthropic prompt caching.

For Vercel, top-level REST providerOptions: { gateway: { caching: "auto" } } and its extra_body equivalent are forwarded. Bifrost does not silently enable a paid write policy. The adapter implements Chat and Responses, not Vercel's Messages endpoint. See Vercel automatic caching.

Google resource references must already exist in the selected provider project. Creating, extending, and deleting those resources is outside the gateway's inference API. See generateContent. DeepSeek, Moonshot, MiniMax, and Z.AI retain their provider-managed caching behavior; see DeepSeek, Moonshot usage, MiniMax, and Z.AI.

Metrics and cost

Public protocolRead tokensWrite tokens
Chat Completionsusage.prompt_tokens_details.cached_tokensusage.prompt_tokens_details.cache_write_tokens
Responses JSON/SSE/WebSocketusage.input_tokens_details.cached_tokensusage.input_tokens_details.cache_write_tokens
Messages JSON/SSEusage.cache_read_input_tokensusage.cache_creation_input_tokens

Chat clients must request stream_options.include_usage: true to receive stream usage. Bifrost always requests usage from Chat upstreams for its own accounting. Responses stream accounting uses the canonical chunks directly, so rendering cannot erase counters or turn an absent provider metric into an observed zero. Messages may carry the final input/cache totals in message_delta; consumers must apply that cumulative snapshot.

Internally, reads and writes are disjoint subsets of total input. Ordinary input equals input minus reads minus writes. Anthropic's input_tokens instead excludes both cache buckets; the adapter adds them when normalizing, and the Messages renderer subtracts them again.

cacheWriteTokensByTtl records write subsets by lifetime in seconds. Messages exposes the native cache_creation object; OpenAI-shaped responses use the optional Bifrost extension cache_write_tokens_by_ttl within input details. Operation metadata retains these buckets. Configure cacheWriteCentsPerMTokensByTtl in catalog or deployment pricing for distinct lifetime rates, e.g. {"300": 625, "3600": 1000}. Context tiers can override these rates. Unknown lifetime prices fall back to the configured aggregate write rate. Deployment pricing overrides replace the entire catalog pricing object, so include all needed rates. The Anthropic catalog includes both lifetimes.

Investigating zero cache reads

  1. Capture one logical request's public protocol, selected provider/deployment, upstream usage, public usage, and SDK usage. Do not record credentials or full private prompts.
  2. Repeat a long stable prefix on the same deployment with the same key and a short changing suffix. Wait for the first request to finish before checking reuse. Test JSON and streaming independently.
  3. Distinguish an explicit provider zero from missing usage or an SDK field that defaults to zero. Responses and Messages may emit protocol-compatible zero defaults; internal optional counters preserve whether the provider reported a value.
  4. If upstream reports a hit but the public response does not, investigate translation. If both carry a hit but the application shows zero, investigate the SDK version and the usage field being read.
  5. If upstream reports zero, check prefix identity, eligibility, lifetime, write policy, provider project/region, deployment changes, fallback and load. More context alone does not guarantee reuse.

Bifrost's HTTP router does not pin a prompt cache key to a deployment. Different cache scopes or fallbacks can therefore reduce reuse. Use one compatible deployment while diagnosing; do not disable health fallback merely to increase a cache metric. A stable key is not a request for Bifrost to store provider KV state, and neither transport preservation nor a test suite guarantees a production hit percentage.

Dashboard usage

Overview shows recorded total tokens (input plus output) and defaults its activity chart to tokens. Usage tables and CSV exports include cache reads, writes, uncached input and unclassified input for models and actors. These input details are subsets, never additions to the total.

  • Cached input: input tokens read from the provider cache.
  • Uncached input: input minus cache reads, calculated per record only when both counts exist. This includes cache writes; it is not the ordinary-input billing bucket described above.
  • Unclassified input: known input whose cache read count was not reported.
  • Cache writes: separately reported tokens written to cache. Missing writes remain unknown.

The reuse percentage uses classified input only. Coverage counts show how many requests or attempts reported each counter. A dash means unreported; a reported zero remains zero. Missing usage is never estimated, and older gateways without these additional fields display unavailable cache details.

Metrics shows provider cache usage across deployment attempts, including retries and fallbacks, with separate time-series controls and per-deployment reuse and reporting coverage. Request totals count each request once. Response-cache hits refer to Bifrost's own response cache, not provider prompt-cache hits. These surfaces use the retained operation and attempt records; changes do not backfill missing historical provider usage.

On this page