Observability
Durable operation lifecycles, upstream attempts, structured logs, and OpenTelemetry.
Every public execution has two correlation keys: x-request-id, which may come from the client, and
an internal operationId. Neither is used as a metric label. Application logs are structured JSON on
stdout/stderr and include the request id, method, path, status, and duration.
Durable lifecycle
The current observability model uses four tables:
| Table | Purpose | Retention |
|---|---|---|
gateway_operations | One row from request start through a verified terminal outcome. | 30 days operational target. |
upstream_attempts | One row per deployment interaction, inserted before the upstream call. | 30 days operational target. |
payload_samples | Redacted, AEAD-encrypted forensic samples, one per finished operation. | 7 days by default. |
payload_access_audit | One immutable audit row for every forensic payload lookup, including the refused and the empty ones. | Operator audit policy. |
An operation is either in_progress or finished. Its outcome is success, incomplete, blocked,
error, cancelled, abandoned, or unknown. degraded is orthogonal: a recovered retry or fallback
is still successful but remains visible. A database constraint rejects new success, incomplete, or
blocked rows unless terminalVerified=true.
The stream supervisor records headers/first-event/first-reasoning/first-output timing, last progress, maximum inter-event gap, upstream/downstream bytes, frame classes, the physical transport terminator, and the normalized semantic terminal. EOF without a terminal, malformed JSON/SSE, duplicate terminals, and empty invalid responses are protocol errors. Only a verified terminal rewards deployment health.
GET /admin/logs lists operation summaries. GET /admin/logs/{operationId} returns the operation and
its complete attempt timeline. GET /admin/observability/summary returns outcome, stall, retry,
active/abandoned, and latency aggregates — for a trailing window=5m|1h|24h, or for an explicit
start/end range of at most 31 days, which is what the dashboard's own range control sends.
Dashboard metrics
The dashboard's Metrics section filters by public model, deployment, operation and date range. Today and calendar dates use UTC; custom end dates are inclusive. Charts switch between traffic, errors, latency and deployment token usage. Tables show public model consumption, deployment attempts, and failures grouped by deployment, kind, phase and provider status.
GET /admin/observability/metrics requires usage:read and accepts start, end, bucket=hour|day,
publicModel, deploymentId and operation. The range is [start, end), limited to 31 days, and
uses request start time for both request and attempt aggregates. Buckets are anchored at start.
Each request counts once; upstream tokens and errors count each attempt, including recovered retries.
With a deployment filter, request metrics describe requests that used that deployment, including
their fallbacks, while deployment metrics include only its own attempts. Request costs are not
attributed to individual attempts. Error rates divide errors by finished requests or attempts.
Only retained records contribute. Deleted deployments remain visible through their recorded attempt metadata. Missing token totals show as unavailable; reported totals include only known usage. Reasoning and cache counts are subsets of usage, not additional tokens to add to the total.
Rerank request summaries record only the public model, document count and byte totals, top_n, and
provider-option names. Response summaries contain only result count, indexes, usage, and cost. Query
and document contents are never stored in normal rows. searchUnits is persisted on operations and
attempts and included in /admin/usage aggregates.
Payload privacy
Ordinary rows contain summaries and HMAC fingerprints, never request or response bodies. IP addresses
and user agents are fingerprinted rather than stored verbatim. Every finished operation keeps one
sample under the versioned encryption keyring — the successful ones too, because a request that
returned 200 and the wrong answer is the case you cannot reconstruct from metadata. Encryption and
fingerprinting use separate, purpose-bound derived subkeys; request, response, error, and attempt
components are each capped at 32 KiB after recursive secret-field redaction. What bounds the
exposure is the retention window (OBSERVABILITY_PAYLOAD_RETENTION_DAYS), not chance.
Reading one
GET /admin/logs/{operationId}/payload needs payloads:read (master key, owner, or admin). Every
lookup writes a durable audit row and emits a structured audit log carrying its outcome —
revealed, missing, sealed, or unreadable — and a successful one also updates the sample's
access timestamp. The gateway never falls back to plaintext.
GET /admin/logs/{operationId} already reports whether a sample stands behind the operation and
whether it can be read, so a caller (and the dashboard) can tell "swept by retention" from "sealed"
without spending an audit entry to find out.
Sealing payloads
OBSERVABILITY_PAYLOAD_ACCESS=sealed keeps capture exactly as it is and closes reading for
everyone: the route refuses before it touches the envelope, so the plaintext never exists in the
process, and the attempt is audited like any other. Reopening is a deployment change — the variable
plus a restart — which is the point: no role, session, or stolen master key can lift it from the
outside. The key that sealed the sample still lives in the gateway, so this is an access policy, not
a cryptographic guarantee against the gateway itself; for that, shorten retention or run sealed and
export nothing.
OpenTelemetry
OpenTelemetry is off by default. Enable it only after an OTLP collector is available:
OTEL_ENABLED=true
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318The gateway exports the ingress-to-close operation span with routing, upstream-attempt, stream,
render, and persistence children, plus low-cardinality counters, histograms, active gauges, queue
depth, and persistence-loss metrics. Payload bodies are never emitted to OpenTelemetry. PostgreSQL,
/health/ready, the observability summary, and structured JSON logs remain sufficient for the initial
operating model.
Rerank usage also increments bifrost_search_units_total; it remains independent from
bifrost_tokens_total.
Initial alerts
- Any successful/incomplete/blocked operation without a verified terminal, abandoned operation, or persistence loss is critical.
- Stalls above 2% over five minutes (minimum 20 requests), protocol errors above 1%, or p95 first output above 25 seconds require investigation.
- Degraded/retry above 5% and client cancellations above 5% are warnings.
Use the request id supplied by the client to find the operation, then inspect its attempts and terminal timeline. See Admin API and Routing.