# API Reference

OpenAI-compatible `/v1*` surface exposed by the gateway. Every route requires `Authorization: Bearer sk-blade-...`. This page is generated from `contracts/openapi.yaml` and resolves `$ref` one level deep; the bundled `openapi.yaml` (the full **/v1** spec) is the escape hatch for anything deeper.

### POST /v1/audio/speech

Text-to-speech (TTS, OpenAI-compatible).

JSON `{model, input, voice?}` → audio bytes (`audio/wav`). Served by
the Orpheus wrapper (voices: tara, leah, jess, leo, dan, mia, zac, zoe;
inline emotion tags `<laugh>`, `<sigh>`…). `input` ≤ 4096 characters.

**Request body** (`application/json`):

- `input` (string, required)
- `model` (string, required)
- `voice` (string, optional)

**Responses:**

- `200` — WAV audio (mono 24 kHz).
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `429` — Rate limit reached. Codes: - `rate_limit_exceeded`: the key's or the user's requests per minute (RPM) or tokens per minute (TPM) are used up; retry after `Retry-After`. - `insufficient_quota`: the user's daily token quota is used up; it resets at midnight UTC. Limits resolve per key first, then per user, then to the platform defaults. The OpenAI SDKs honour `Retry-After` natively (automatic backoff). The video route also returns its own 429, `video_jobs_quota_exceeded`, when too many jobs are still running; that cap is separate from rate limiting. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### POST /v1/audio/transcriptions

Audio transcription (ASR, OpenAI-compatible).

Multipart `file=<audio>` (wav, mp3, m4a, flac, ogg — 25 MB max) +
`model=<model_id tier=audio>`. The gateway proxies the multipart as-is
to the vLLM replica (Whisper/Voxtral), which returns `{text, usage:{type:
"duration", seconds}}`. v1 metering: estimated from the transcribed text
(metering="estimated") — the "audio seconds" unit is coming in M2.

**Request body** (`multipart/form-data`):

- `file` (string, required)
- `model` (string, required)

**Responses:**

- `200` — Transcription result.
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `413` — Request body beyond the limit (size guardrail). → `ErrorEnvelope`
- `429` — Rate limit reached. Codes: - `rate_limit_exceeded`: the key's or the user's requests per minute (RPM) or tokens per minute (TPM) are used up; retry after `Retry-After`. - `insufficient_quota`: the user's daily token quota is used up; it resets at midnight UTC. Limits resolve per key first, then per user, then to the platform defaults. The OpenAI SDKs honour `Retry-After` natively (automatic backoff). The video route also returns its own 429, `video_jobs_quota_exceeded`, when too many jobs are still running; that cap is separate from rate limiting. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### POST /v1/chat/completions

Chat completions (core route, stream + non-stream).

OpenAI-compatible. Supported fields: `model`, `messages`, `stream`,
`stream_options`, `max_tokens`, `temperature`, `top_p`, `stop`, `seed`,
`frequency_penalty`, `presence_penalty`, `n`, `web_search_options`,
`tools`, `tool_choice`. Out-of-scope fields (`response_format`…) are
relayed as-is to vLLM if it supports them, otherwise a clear error is
returned (SPEC §5.1). On a Harmony-format model whose `auto` tool calls need it,
when `tool_choice` is `auto` or absent and no function tool sets `strict`, the
gateway sends every function tool with `strict: true` (here and on
`/v1/responses`), which constrains the call to the Harmony format; set `strict`
on any tool yourself to opt out.

**The pass-through is deliberate, and two things keep it honest.** An allowlist
would refuse the engine extras callers legitimately use (`chat_template_kwargs`,
`top_k`, `repetition_penalty`, the guided-decoding family), so an unknown field is
relayed rather than rejected. What a caller gets instead is the answer BEFORE and
AFTER: `supported_parameters` on `GET /v1/models` lists, per model, what this
gateway stands behind, and any parameter it had to drop is named on the response in
`x-shadow-ignored-params`. A field that is neither listed nor announced was relayed
to the engine — the one case left where "accepted" does not prove "applied".

Guardrails (400 otherwise): `model` known, body size ≤ limit,
`max_tokens` ≤ the model's `max_output_tokens`,
`prompt_tokens + max_tokens ≤ max_model_len`, every `messages[].role` in the enum.

**Two parameters are answered per MODEL**, and `supported_parameters` on
`GET /v1/models` says which beforehand:

- `reasoning_effort` — HONOURED where the model exposes a reasoning control:
  the gateway translates it onto that control (`chat_template_kwargs.
  enable_thinking`) instead of relaying a field the engine validates and then
  ignores. `off` is a synonym of `none`. On a model whose chat template reads
  the grade itself (the Harmony models, e.g. `gpt-oss-20b`) it is relayed:
  `minimal` becomes `low`, and `none` is dropped because no grade turns that
  deliberation off. On a model with no such control, or `none` on a Harmony
  model, it is dropped and named in the `x-shadow-ignored-params` response
  header, never a silent no-op. An unknown value is a 400.
- `stop` — REFUSED (400 `unsupported_parameter`) on a model whose deliberation
  comes back in its own channel (`reasoning.separate_channel: true`). Stop
  sequences are matched against that channel too, so a sequence the model
  happens to think about ends the generation before the answer exists, with the
  `finish_reason: "stop"` of a normal completion — a failure the client cannot
  detect. Reproduced outside this platform (OpenRouter, qwen3-30b-a3b) and open
  upstream as vllm-project/vllm#38499.

`tools` carries TWO vocabularies. A `{"type": "function"}` entry is a
CLIENT function: relayed as-is, and the model's call comes back for the
caller to answer. An entry naming a SERVER tool
(`{"type": "web_search"}`, `{"type": "web_fetch", "max_uses": 3}`,
`{"type": "code_interpreter"}` — see docs/SERVER-TOOLS.md) is executed by
the gateway inside its own loop and never reaches vLLM, which understands
function declarations only. Any other `type` is a 400 rather than a
pass-through: a declaration nothing implements would otherwise be dropped
in silence while the caller believed a tool had run. Each server tool
needs its own account flag (`web_search_enabled`, `web_fetch_enabled`,
`code_interpreter_enabled`) and a model with native tool calling: a model
whose tool calls the serving stack cannot extract is refused up front
(400) rather than answered with raw JSON — see the per-model table in
docs/SERVER-TOOLS.md.

`web_search_options` (PAAS-5, opt-in, see docs/SERVER-TOOLS.md): the
legacy entry point for the same web_search tool. When present, the
gateway runs a tool loop against the Staan WebSearch4AI vendor and
returns a cited answer (`[n]` inline markers +
`message.annotations`). Requires the account flag `web_search_enabled`
and a model with native tool calling.

STREAMING with any server tool declared: the gateway is no longer a
relay, and the loop publishes its activity as extension chunks with
`"choices": []` and a `"web_search"` object (full ordering in
STREAMING.md §2.2 bis):

- `{"status": "searching", "query": "..."}` — a web_search call about to run;
- `{"status": "fetching", "url": "..."}` — a web_fetch call about to run;
- `{"status": "running_code", "language": "python", "code_chars": <n>}` —
a code_interpreter run about to start. The PROGRAM is not echoed here: the
client already has it in the tool call, and an activity chunk is not the
place to re-broadcast the user's data.
The key stays `web_search` for EVERY tool: it is the name the first tool
shipped with, and a rename would break deployed clients. A replacement
naming the tool and the call's outcome will be emitted ALONGSIDE it for
one release before this one is retired.
- `{"status": "done", "searches": <n>, "price_per_search": <float>}` —
once per request, even when no call ran. `searches` counts web_search
ONLY (it is a billed unit, not a total of tool calls), so a fetch-only or
compute-only turn legitimately ends on `searches: 0`.

Per-call OUTCOMES are not published on this surface: a failed call is
indistinguishable from a successful one here, and the request still
returns 200 (a tool failure is never a failed request). A client that
needs to know WHICH tool ran and whether it succeeded should use
POST /v1/responses, whose output items name the tool and carry
`status: "failed"`. A final chunk carries `delta.annotations` with the
`url_citation` list. In the non-stream response, the same information is
exposed as a top-level `web_search` object (`{"searches": <n>,
"price_per_search": <float>}`) alongside
`choices[].message.annotations`.

If `stream=false` → JSON `ChatCompletion` response.
If `stream=true`  → `text/event-stream` SSE stream of `ChatCompletionChunk`
terminated by `data: [DONE]` (native relay of the vLLM stream — see
STREAMING.md). With `stream_options.include_usage=true`, a final chunk
carries `usage`.

**Request body** (`application/json`):

`ChatCompletionRequest`
- `frequency_penalty` (number, optional)
- `max_completion_tokens` (integer, optional) — OpenAI's current name for `max_tokens`, and the same knob here: same cap, same context check, and an error about it names this field. Sent together, this one wins (as upstream).
- `max_tokens` (integer, optional) — Capped by the gateway: ≤ the model's `max_output_tokens` (published on `GET /v1/models`; the global guard-rail unless the model raises it) AND prompt_tokens + max_tokens ≤ the model's max_model_len. OMITTED is the normal case and costs nothing: the engine then fills the window it has left, computed from the real prompt length. The gateway states a ceiling upstream only when its own cap is the binding one — it never derives a ceiling from its own prompt ESTIMATE, which is how a request with no `max_tokens` used to be refused outright on a model whose cap equals its window.
- `messages` (array of `ChatMessage`, required)
- `model` (string, required) — Public model_id from the catalog.
- `n` (integer, optional) — Number of completions. n>1 ⇒ num_requests=n for metering. Requires a non-zero `temperature`: below 1e-5 the engine samples greedily, every completion would be identical, and it refuses the request rather than return copies. Asking for both is a 400 naming `n`, raised before any GPU is woken.
- `presence_penalty` (number, optional)
- `seed` (integer, optional)
- `stop` (string or array of string, optional)
- `stream` (boolean, optional) — If true, response as an SSE stream (native relay — see STREAMING.md).
- `stream_options` (`StreamOptions`, optional)
- `temperature` (number, optional)
- `tools` (array of object, optional) — Function declarations, relayed to vLLM as-is, AND/OR the gateway's built-in server tools (`{"type": "web_search"}`, `{"type": "web_fetch"}`, `{"type": "code_interpreter"}`) — the same vocabulary /v1/responses accepts. A built-in entry is executed by the gateway inside its own tool loop and stripped from the upstream body; an unknown `type` is a 400 with `param: "tools"`, which narrows the older "out-of-scope fields are relayed as-is" contract on purpose (a `{"type": …}` vLLM does not understand is an upstream 400 at best and silently ignored at worst). `{"type": "web_search"}` and `web_search_options` are the same switch: agreeing duplicates are merged, disagreeing ones are a 400.
- `top_p` (number, optional)
- `web_search_options` (object, optional) — Opt-in web search (PAAS-5). When present, the gateway injects a `web_search` tool, executes model-emitted searches server-side via the Staan WebSearch4AI vendor (queries LEAVE Shadow infrastructure — see docs/SERVER-TOOLS.md), and returns a cited answer. Each executed search is billed at defaults.price_web_search. Requires the account flag `web_search_enabled` and a model with native tool calling. Mirrors OpenAI's field of the same name.

**Responses:**

- `200` — Completion. JSON if `stream=false`; SSE stream if `stream=true`. → `ChatCompletion`
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `403` — The model is served by an external inference provider and the caller's organization has not enabled external providers (PAAS-206 / PAAS-32 X1). Deliberately a 403 that NAMES the model, where an internal-only model answers 404: the model's existence is not a secret here — only its enablement is — and the error is the enablement funnel. Raised before the rate limiter and before any upstream call, so a refused request costs nothing and, in particular, spends no supplier money. Retrying will not help; an administrator enabling external providers for the organization will. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `413` — Request body beyond the limit (size guardrail). → `ErrorEnvelope`
- `429` — Rate limit reached. Codes: - `rate_limit_exceeded`: the key's or the user's requests per minute (RPM) or tokens per minute (TPM) are used up; retry after `Retry-After`. - `insufficient_quota`: the user's daily token quota is used up; it resets at midnight UTC. Limits resolve per key first, then per user, then to the platform defaults. The OpenAI SDKs honour `Retry-After` natively (automatic backoff). The video route also returns its own 429, `video_jobs_quota_exceeded`, when too many jobs are still running; that cap is separate from rate limiting. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### POST /v1/completions

Legacy completions (compat for older clients).

Legacy OpenAI-compatible. `prompt` string or array of strings. Same
guardrails, same stream/non-stream modes (native SSE relay) as chat.

**Request body** (`application/json`):

`CompletionRequest`
- `max_tokens` (integer, optional)
- `model` (string, required)
- `n` (integer, optional)
- `prompt` (string or array of string, required)
- `seed` (integer, optional)
- `stop` (string or array of string, optional)
- `stream` (boolean, optional)
- `stream_options` (`StreamOptions`, optional)
- `temperature` (number, optional)
- `top_p` (number, optional)

**Responses:**

- `200` — Completion (JSON or SSE depending on `stream`). → `Completion`
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `403` — The model is served by an external inference provider and the caller's organization has not enabled external providers (PAAS-206 / PAAS-32 X1). Deliberately a 403 that NAMES the model, where an internal-only model answers 404: the model's existence is not a secret here — only its enablement is — and the error is the enablement funnel. Raised before the rate limiter and before any upstream call, so a refused request costs nothing and, in particular, spends no supplier money. Retrying will not help; an administrator enabling external providers for the organization will. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `413` — Request body beyond the limit (size guardrail). → `ErrorEnvelope`
- `429` — Rate limit reached. Codes: - `rate_limit_exceeded`: the key's or the user's requests per minute (RPM) or tokens per minute (TPM) are used up; retry after `Retry-After`. - `insufficient_quota`: the user's daily token quota is used up; it resets at midnight UTC. Limits resolve per key first, then per user, then to the platform defaults. The OpenAI SDKs honour `Retry-After` natively (automatic backoff). The video route also returns its own 429, `video_jobs_quota_exceeded`, when too many jobs are still running; that cap is separate from rate limiting. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### POST /v1/detect

Object/region detection (document-CV utility heads).

Image in → structured regions out. JSON `{model, image, ...knobs}` where
`image` is base64 or a `data:image/…;base64,…` URI. Serves detection heads
(e.g. `tatr` — table structure). Synchronous, one JSON document, NO streaming
(`stream:true` → 400). Model addressing is exact (no default). Billed per
processed image.

**Request body** (`application/json`):

- `image` (string, required) — base64 or data-URI image
- `model` (string, required)

**Responses:**

- `200` — Detection result.
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### GET /v1/models

List the models in the catalog.

OpenAI format `{object:"list", data:[Model…]}`. The list comes from
`models.yaml` (the gateway catalog), not vLLM.

**Responses:**

- `200` — Catalog. → `ModelList`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`

### POST /v1/parse

Structured parsing (document-CV utility heads).

Image in → structured text/table out. JSON `{model, image, ...knobs}`.
Serves parsing heads (`trocr` — handwriting; `got-ocr-2` — general OCR;
`deplot` — chart to data table; `pix2text` — math formula to LaTeX;
`tableformer` — table structure recognition).
The response carries a polymorphic `text`/`table`/`tables`/`output` field per
model. Synchronous, one JSON document, NO streaming (`stream:true` → 400).
Billed per processed image.

**`tableformer` additionally requires `table_bboxes`** — it recognises the
structure inside a table it is GIVEN, it does not locate the table. Get the
boxes from `tatr` on `/v1/detect`, or pass the full-page box. A body without it
is rejected with 400 immediately (no replica is woken).

**Request body** (`application/json`):

- `image` (string, required) — base64 or data-URI image
- `max_new_tokens` (integer, optional) — Generation ceiling for the text-generating heads (`trocr`, `deplot`). Clamped to the model's own bound.
- `model` (string, required)
- `table_bboxes` (array of array, optional) — REQUIRED for `tableformer`: pixel boxes of the tables to recognise. Ignored by the other parse heads.

**Responses:**

- `200` — Parse result (shape varies by model).
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### POST /v1/responses

Responses API (server tools, item-based output) — stateless only.

OpenAI-compatible Responses surface, implemented as a TRANSLATION LAYER over the
same gateway-owned server-tool loop `/v1/chat/completions` uses: the tools, the
budgets, the citations and the billing are identical, only the request and event
vocabularies differ (see docs/SERVER-TOOLS.md).

Accepted: `model`, `input` (a string, or an array of `message` /
`function_call` / `function_call_output` items), `instructions`, `tools`,
`tool_choice`, `parallel_tool_calls`, `max_output_tokens`, `temperature`,
`top_p`, `seed`, `stream`, `store` (false), `metadata`, `text`
(`{"format": {"type": "text"}}` only), `user`, `reasoning`
(`{"effort": …}`, see the schema).

**Stateless by design.** `store: true` and `previous_response_id` return 400
`unsupported_parameter`: no conversation is persisted, so accepting either would
mean answering a different question than the client believes it asked. `store`
is echoed as `false` in every response, including when it was omitted (OpenAI
defaults it to true). Any parameter not in the list above is a 400
`unknown_parameter` rather than being silently ignored.

**Not available here:** `ded/<slug>` aliases (a pricing surface of
`/v1/chat/completions`, refused with an explicit 400 rather than served at a
different price), and `n > 1` (the surface returns one response). Pool pricing is
spot like everywhere else: the multiplier in force at submission is stamped on
the usage event; the legacy `:spot` suffix is accepted as a deprecated no-op.

`tools` accepts `function` declarations (flattened, the Responses shape) plus the
built-in server tools — `{"type": "web_search"}`, `{"type": "web_fetch"}`,
`{"type": "code_interpreter"}` — which require the matching account entitlement
and a model with native tool calling. `tool_choice` cannot be combined with a
built-in tool (400): the gateway drives the loop and sets `tool_choice` itself on
every round.

**Output** is an item list: one `web_search_call` / `web_fetch_call` /
`code_interpreter_call` per executed server-tool call (`action` says what it did,
`status` whether it worked — a tool failure is a failed ITEM, never a failed
request), then a `reasoning` item when the model deliberated in its own channel,
one `function_call` per client function the model asked for, and the assistant
`message` whose `output_text` content part carries `url_citation` annotations.

The `reasoning` item carries the model's RAW deliberation in
`content[].reasoning_text` (`summary` stays empty — nothing here summarises it),
and streams as `response.reasoning_text.delta` events that close on
`response.reasoning_text.done` before the answer's first token, so a client can
show the thinking live and collapse it when the answer starts. It appears only
for models whose serving stack splits that channel
(`reasoning.separate_channel` on `GET /v1/models`); on every other model the
deliberation is part of the message and nothing can separate it.

`usage.output_tokens_details.reasoning_tokens` is present only when the engine
reported the split. It used to be hard-coded to 0, which reads as "this model did
not reason" on a response whose tokens were nearly all deliberation. The same rule
holds on `/v1/chat/completions`, where `usage.completion_tokens_details` is relayed
verbatim when the engine sends it and omitted when it does not: splitting the count
here would need a tokenizer, and an estimate presented as accounting is worse than
an absent field.

**Citations** carry BOTH shapes: `index` is the `[n]` marker number the model
writes inline (authoritative, the same value the chat surface returns) and
`start_index`/`end_index` are the best-effort character span of that marker in the
answer. A source the model never referenced has an EMPTY span at the end of the
text — models do cite numbers they never fetched, and there is no honest offset
for those. See docs/SERVER-TOOLS.md.

If `stream=false` → one JSON `Response`. If `stream=true` → a NAMED-event SSE
stream (`event:` + `data:` lines) with `sequence_number` on every event:
`response.created`, `response.output_item.added`/`.done`,
`response.<tool>_call.in_progress`/`.completed`, `response.content_part.added`/
`.done`, `response.output_text.delta`/`.done`,
`response.output_text.annotation.added`, and exactly one terminal event —
`response.completed`, `response.incomplete` (truncated answer, or the model
stream breaking after partial text) or `response.failed` (no answer produced).
There is no `[DONE]` sentinel: the terminal event is the terminator.

Metering: `usage_events.endpoint = 'responses'`, one token event per request
(streamed or not), plus one row per executed server-tool call under that tool's
own endpoint — the same rows the chat surface writes.

**Request body** (`application/json`):

`ResponsesRequest`
- `input` (string or array of object, required) — A plain string (one user turn), or the item array. Item types translated: `message` (roles user/assistant/system/developer — `developer` maps to a system turn), `function_call` and `function_call_output`, which is how a client answers its own function tool without a stored session. Content parts: `input_text`, `output_text`, `input_image`.
- `instructions` (string, optional) — Prepended as a system turn. When server tools are declared, the gateway's tool guidance is merged INTO it rather than added as a second system message (several chat templates reject a system turn that is not first).
- `max_output_tokens` (integer, optional) — Same cap as chat `max_tokens` (≤ the model's `max_output_tokens`, prompt + output ≤ max_model_len).
- `metadata` (object, optional) — Echoed back unchanged; not stored.
- `model` (string, required) — Public model_id from the catalog, `kind: chat`. `ded/<slug>` aliases are refused here (400) — use /v1/chat/completions. The legacy `:spot` suffix is a deprecated no-op (every pool price is spot).
- `parallel_tool_calls` (boolean, optional)
- `previous_response_id` (string, optional) — Not supported — a non-null value returns 400 `unsupported_parameter`. Replay the conversation in `input` instead.
- `reasoning` (object, optional) — `{"effort": "none"|"minimal"|"low"|"medium"|"high"}`. Translated onto the model's OWN control — `chat_template_kwargs.enable_thinking`, the one that acts on this fleet — so `none` turns the deliberation off and anything else turns it on. Only the extremes are actionable, because the control is an on/off switch and grading it would advertise a granularity nothing implements. A model with no such control (`reasoning.controllable: false` on `GET /v1/models`) returns 400 `unsupported_parameter` rather than accepting the parameter and ignoring it. `summary` is not supported: this surface streams the model's raw deliberation and nothing summarises it.
- `seed` (integer, optional)
- `store` (boolean, optional) — Must be false (or absent). `true` returns 400 `unsupported_parameter`: nothing is persisted, so a stored response could never be retrieved or continued.
- `stream` (boolean, optional)
- `temperature` (number, optional)
- `text` (object, optional) — Only `{"format": {"type": "text"}}` is accepted; a `json_schema` format returns 400 (structured decoding is not translated on this surface).
- `tool_choice` (string or object, optional) — Relayed to the model. Cannot be combined with a built-in tool (400 `unsupported_parameter`): the gateway owns the loop and sets `tool_choice` per round itself, so a caller's value would be overwritten.
- `tools` (array of object, optional) — Function declarations in the Responses (flattened) shape — `{"type": "function", "name", "parameters"}` — and/or built-in server tools.
- `top_p` (number, optional)
- `user` (string, optional)

**Responses:**

- `200` — The response. JSON if `stream=false`; a named-event SSE stream if `stream=true`. → `Response`
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `403` — The account is not entitled to a declared server tool, the model is disabled by an administrator, or the model is served by an external inference provider the organization has not enabled (`external_provider_not_enabled`, PAAS-206 — see the response of the same name). → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `413` — Request body beyond the limit (size guardrail). → `ErrorEnvelope`
- `429` — Rate limit reached. Codes: - `rate_limit_exceeded`: the key's or the user's requests per minute (RPM) or tokens per minute (TPM) are used up; retry after `Retry-After`. - `insufficient_quota`: the user's daily token quota is used up; it resets at midnight UTC. Limits resolve per key first, then per user, then to the platform defaults. The OpenAI SDKs honour `Retry-After` natively (automatic backoff). The video route also returns its own 429, `video_jobs_quota_exceeded`, when too many jobs are still running; that cap is separate from rate limiting. → `ErrorEnvelope`
- `502` — Unrecoverable upstream error (vLLM/Forge). → `ErrorEnvelope`
- `503` — Model unavailable (failed cold start, no replica). The client may retry. → `ErrorEnvelope`
- `504` — Upstream timeout exceeded (cold start too long, slow generation). → `ErrorEnvelope`

### GET /v1/video/generations

List the requester's 20 most recent video jobs.

**Responses:**

- `200` — List.

### POST /v1/video/generations

Submit a text-to-video job (asynchronous, 202).

Submit→poll: generation (1-3 min) runs as an async Forge job — the
202 returns immediately with a job_id. Quota: 2 non-terminal jobs
per user (429 code video_jobs_quota_exceeded). Limits:
704×480, 161 frames, 50 steps, prompt ≤ 2000 chars.

**Request body** (`application/json`):

- `fps` (integer, optional)
- `height` (integer, optional)
- `model` (string, required)
- `negative_prompt` (string, optional)
- `num_frames` (integer, optional) — rounded to 8k+1
- `num_inference_steps` (integer, optional)
- `prompt` (string, required)
- `seed` (integer, optional)
- `width` (integer, optional)

**Responses:**

- `202` — Job created. → `VideoGenerationJob`
- `400` — Invalid request: a guardrail was violated (max_tokens > cap, context exceeded) or the body is malformed. → `ErrorEnvelope`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
- `429` — Rate limit reached. Codes: - `rate_limit_exceeded`: the key's or the user's requests per minute (RPM) or tokens per minute (TPM) are used up; retry after `Retry-After`. - `insufficient_quota`: the user's daily token quota is used up; it resets at midnight UTC. Limits resolve per key first, then per user, then to the platform defaults. The OpenAI SDKs honour `Retry-After` natively (automatic backoff). The video route also returns its own 429, `video_jobs_quota_exceeded`, when too many jobs are still running; that cap is separate from rate limiting. → `ErrorEnvelope`

### GET /v1/video/generations/{job_id}

State of a video job (+ presigned mp4 URL when completed).

Reconciles state with the control plane on EVERY read (the presigned
S3 URL is re-resolved — the one returned previously expires).
Uniform 404 if the job does not exist or belongs to someone else.

**Parameters:**

- `job_id` (`path`, required, string)

**Responses:**

- `200` — Job. → `VideoGenerationJob`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`

### POST /v1/video/generations/{job_id}/cancel

Cancel a video job (best-effort).

**Parameters:**

- `job_id` (`path`, required, string)

**Responses:**

- `200` — Job (status cancelled if the cancellation took effect). → `VideoGenerationJob`
- `401` — Missing, malformed, unknown, or revoked key. → `ErrorEnvelope`
- `404` — Model unknown to the catalog. → `ErrorEnvelope`
