API overview
Auth, error shape, and how to import the spec — shared across every /v1/* endpoint.
Use this reference to choose an endpoint and understand authentication, errors, and request correlation. The OpenAPI specification contains the exact schemas and can be imported into Bruno, Postman, or Insomnia.
Examples use GATEWAY for the origin (for example, http://localhost:4000) and API_KEY for a
virtual key. The model field always names a public model configured in Bifrost.
Authentication
- Inference (
/v1/*): master key or virtual key, viaAuthorization: Bearer <key>or thex-api-keyheader. Query-string credentials are rejected because URLs leak through logs, history, and referrers./v1/modelsand/v1/models/{model}are the only unauthenticated exceptions — see Model discovery. - Admin (
/admin/*): use the master key for API automation. When dashboard authentication is enabled, operator sessions can also access routes permitted by their role.
Full model: Security.
Error shape
Most errors use the OpenAI-compatible envelope. /v1/messages uses Anthropic's shape and
/v1/rerank uses OpenRouter's { "error": { "code": <status>, "message": "..." } } shape:
{ "error": { "message": "...", "type": "invalid_request_error", "param": "model", "code": "model_not_found" } }Validation failures produced by the gateway include the field path and validation detail. Structured
upstream request rejections also preserve their actionable message, param, and code, after
bounding the fields and redacting the configured upstream model and credentials. The raw provider
body is never copied into the response. Operational failures (authentication, permissions, missing
deployments, and rate limits) and 5xx responses keep stable gateway messages so routing internals do
not leak. See Troubleshooting for the full class/status/code table.
Request correlation
Every response carries x-request-id (echoed if you send one). Always include it when reporting an
issue — see Headers for the full header reference and
Observability for how to look one up in the logs.
Importing into Bruno
Import Collection → OpenAPI V3, select apps/gateway/openapi.yaml, then set the baseUrl server
variable (e.g. http://localhost:4000) and the bearer token in Bruno's auth UI.
Typical flow to serve a model
POST /admin/deployments— create a deployment:publicModel+adapterKey+upstreamModel+ inline credentials.POST /admin/keys— issue a virtual key withallowedModels.- Call an inference endpoint with the public model name in
model.
POST /admin/deployments/resolve validates the profile/operations/transports without saving — see
First deployment.
The endpoints
| Endpoint | Contract | Page |
|---|---|---|
POST /v1/chat/completions | OpenAI Chat Completions | Chat completions |
POST /v1/responses, WS /v1/responses, GET/DELETE /v1/responses/:id, GET /v1/responses/:id/input_items | OpenAI Responses | Responses |
POST /v1/messages | Anthropic Messages | Messages |
POST /v1/images/generations, POST /v1/images/edits | OpenAI Images | Images |
POST/GET/DELETE /v1/videos, GET /v1/videos/:id/content | OpenAI-shaped Videos | Videos |
POST /v1/embeddings | OpenAI Embeddings | Embeddings |
POST /v1/rerank | OpenRouter Rerank | Reranking |
POST /v1/audio/transcriptions | OpenAI Audio | Audio transcriptions |
GET /v1/models, GET /v1/models/{model}, GET /v1/models/{model}/deployments | Model discovery | Model discovery |
Next steps
- Streaming — SSE shape, shared across every streamable endpoint.
- Headers — every request/response header in one table.
- Troubleshooting — error classes, status codes, and common symptoms.