OpenAI-compatible API
Call Tenzro Network 1 with any OpenAI SDK: chat, responses, embeddings, images, audio and video, with streaming, API keys or pay-per-request.
Every Tenzro node serves an HTTP API in the OpenAI wire format. Point an existing OpenAI SDK at a Tenzro endpoint and it works unchanged: chat, responses, embeddings, images, audio and video. Where OpenAI publishes a path for a modality, Tenzro serves it at that path in that shape. Modalities with no vendor path, such as forecasting and detection, live under /v1/tenzro/.
Base URL: https://rpc.tenzro.xyz/v1, or http://<your-node>:8545/v1 on your own node.
Authentication and payment
Generation routes need one of two things:
- An API key, sent as the
X-Tenzro-Api-Keyheader. Usage is metered per call and billed to the account that owns the key. See API keys. - Payment per request. Send the request with no key and the node answers
402 Payment Requiredwith a challenge describing how to pay (x402 or MPP), in TNZO or stablecoins. Retry with the payment credential and the request proceeds. This is how agents and machines pay without an account relationship.
Model listing, generation lookups and video downloads are open.
Routes
| Method and path | What it does | Access |
|---|---|---|
POST /v1/chat/completions | Chat completions, streaming or not | key or 402 |
POST /v1/responses | The Responses shape over the same handler | key or 402 |
POST /v1/embeddings | Text embeddings | key or 402 |
POST /v1/images/generations | Text to image | key or 402 |
POST /v1/images/edits | Image to image (multipart) | key or 402 |
POST /v1/audio/transcriptions | Speech to text (multipart) | key or 402 |
POST /v1/audio/speech | Text to speech; returns raw audio | key or 402 |
POST /v1/videos | Text or image to video; returns a job | key or 402 |
GET /v1/videos/{id} | Video job status | open |
GET /v1/videos/{id}/content | The finished clip | open |
GET /v1/models | Every model this gateway can serve, with pricing | open |
GET /v1/models/{id} | One model | open |
GET /v1/generation?id=... | Recorded usage and cost for one completion | open |
Tenzro extensions, same access rules as the generation routes:
| Path | Modality |
|---|---|
POST /v1/tenzro/forecasts | Time-series forecasting |
POST /v1/tenzro/detections | Object detection |
POST /v1/tenzro/segmentations | Promptable and open-vocabulary segmentation |
POST /v1/tenzro/video/embeddings | Clip embedding |
Any JSON-RPC method is also reachable over REST at POST /v1/rpc/{method}. Request and response shapes for each modality are in Multimodal.
Client setup
from openai import OpenAI
client = OpenAI(
base_url="https://rpc.tenzro.xyz/v1",
api_key="unused",
default_headers={"X-Tenzro-Api-Key": "tnz_..."},
)
stream = client.chat.completions.create(
model="qwen3-8b",
messages=[{"role": "user", "content": "Explain metered settlement in one paragraph."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://rpc.tenzro.xyz/v1",
apiKey: "unused",
defaultHeaders: { "X-Tenzro-Api-Key": process.env.TENZRO_API_KEY! },
});
const res = await client.chat.completions.create({
model: "qwen3-8b",
messages: [{ role: "user", content: "hello" }],
});curl https://rpc.tenzro.xyz/v1/chat/completions \
-H 'content-type: application/json' \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-d '{"model":"qwen3-8b","messages":[{"role":"user","content":"hello"}]}'Chat completions
The standard fields are honoured: messages, max_tokens, temperature, top_p, stop, seed, frequency_penalty, presence_penalty, stream, stream_options and user, plus top_k, min_p and repetition_penalty. n must be 1: the network bills per completion.
Message content is a string or an array of parts: text, image_url (inline the bytes as a data: URI) for vision-language models, and input_audio and file for models that accept them. A part the model cannot take is refused by name rather than silently ignored.
Tenzro extensions ride in the same body. OpenAI SDKs pass unknown fields through.
models fallback model ids, tried in order
provider {"only": [...], "ignore": [...]} to narrow the provider pool
jurisdiction "DE,EU": route only to providers whose locality claim matches
jurisdiction_receipt "required": fail unless a signed locality receipt comes back
draft_n opt in to speculative decoding on models with a drafterThe response's model field is the candidate that actually served. Responses also carry cost_wei, generation_time_ms and tokens_per_second, and, where the provider signs, a tenzro_contentProvenance manifest and a tenzro_jurisdiction receipt. See Inference for the routing behind each field.
Responses
POST /v1/responses translates the Responses shape into a chat request, runs it through the same handler and translates the result back. Routing, pinning, settlement, provenance and streaming behave identically. input becomes messages, instructions a leading system message and max_output_tokens the token limit; usage uses the Responses names input_tokens and output_tokens. Prior turns are replayed in input: previous_response_id is refused because the node retains no responses.
Streaming
Set stream: true and the response is Server-Sent Events.
- Chat completions stream
chat.completion.chunkevents ending withdata: [DONE]. The final chunk carriesusageand the billing fields, so you learn what you were billed without opting in.stream_options: {"include_usage": true}also appends the standard empty-choicesusage chunk. - Responses stream typed events, from
response.createdthroughresponse.output_text.deltatoresponse.completed,response.incompleteorresponse.failed. The terminal event ends the stream; there is no[DONE]. - Resume. Every event id is
<completion_id>:<seq>. Reconnect withLast-Event-IDand the node replays the chunks you missed from a short-lived in-memory buffer. - Failover. If the serving provider drops mid-generation, the gateway continues the stream on another provider of the same model, replaying the text already emitted as an assistant prefix with the same seed. Pin
seedfor byte-identical continuation.
{"id":"chatcmpl-...","object":"chat.completion.chunk","model":"qwen3-8b",
"choices":[{"index":0,"delta":{},"finish_reason":"stop","native_finish_reason":"eos"}],
"usage":{"prompt_tokens":12,"completion_tokens":48,"total_tokens":60},
"cost_wei":"...","generation_time_ms":...,"tokens_per_second":...}
data: [DONE]finish_reason uses the OpenAI vocabulary; native_finish_reason carries the exact cause the engine reported. A client that closes the connection before the final chunk can read the same numbers later from GET /v1/generation?id=<completion_id>.
Over JSON-RPC, tenzro_chatStream streams the rich shape with reasoning and tool-call events, and POST /chat-stream serves the same rich shape as SSE.
Embeddings, transcriptions, speech and image generation return a single body. Video is a job: poll it, then download.
Model listing
Each entry in GET /v1/models carries the OpenAI fields plus the serving contract: context length, maximum output tokens, pricing in TNZO wei per token (and in USD where the operator declares a listing rate), the pricing model (PerToken, PerRequest, PerComputeTime or Dynamic), supported parameters, feature flags and the provider's declared jurisdiction. The list also carries a data_policy: prompts and completions are not written to disk, the stream resume buffer is in memory and expires, and usage accounting records metered units, cost and latency only, never content.
Errors
Errors use the OpenAI envelope, {"error": {"message", "type", "code"}}.
| Status | Meaning |
|---|---|
| 400 | Invalid request, or a field or content part the model cannot honour, named in the message |
| 401 | Missing or invalid API key on a key-gated model |
| 402 | Payment required; the body carries the payment challenge |
| 404 | No provider serves that model, or a provider pin admitted none |
| 409 | Video content requested before the job completed |
| 412 | No provider satisfies the jurisdiction pin, or a required receipt is unavailable |
| 429 | Rate limited |
| 502 | The provider was unreachable or returned an error |
| 504 | An image render outlived the wait; the message names the job to poll |
Next
- Modality shapes and examples: Multimodal.
- Paying per request from an agent: x402, Stablecoin payments.
- Lower latency and price on repeated prompts: Prefix and state reuse.