Skip to content
Tenzro
Documentation menu
Inference

OpenAI-compatible API

Call Tenzro Network 1 with any OpenAI SDK: chat, responses, embeddings, images, audio and video, with streaming, API keys or pay-per-request.

Every Tenzro node serves an HTTP API in the OpenAI wire format. Point an existing OpenAI SDK at a Tenzro endpoint and it works unchanged: chat, responses, embeddings, images, audio and video. Where OpenAI publishes a path for a modality, Tenzro serves it at that path in that shape. Modalities with no vendor path, such as forecasting and detection, live under /v1/tenzro/.

Base URL: https://rpc.tenzro.xyz/v1, or http://<your-node>:8545/v1 on your own node.

Authentication and payment

Generation routes need one of two things:

  • An API key, sent as the X-Tenzro-Api-Key header. Usage is metered per call and billed to the account that owns the key. See API keys.
  • Payment per request. Send the request with no key and the node answers 402 Payment Required with a challenge describing how to pay (x402 or MPP), in TNZO or stablecoins. Retry with the payment credential and the request proceeds. This is how agents and machines pay without an account relationship.

Model listing, generation lookups and video downloads are open.

Routes

Method and pathWhat it doesAccess
POST /v1/chat/completionsChat completions, streaming or notkey or 402
POST /v1/responsesThe Responses shape over the same handlerkey or 402
POST /v1/embeddingsText embeddingskey or 402
POST /v1/images/generationsText to imagekey or 402
POST /v1/images/editsImage to image (multipart)key or 402
POST /v1/audio/transcriptionsSpeech to text (multipart)key or 402
POST /v1/audio/speechText to speech; returns raw audiokey or 402
POST /v1/videosText or image to video; returns a jobkey or 402
GET /v1/videos/{id}Video job statusopen
GET /v1/videos/{id}/contentThe finished clipopen
GET /v1/modelsEvery model this gateway can serve, with pricingopen
GET /v1/models/{id}One modelopen
GET /v1/generation?id=...Recorded usage and cost for one completionopen

Tenzro extensions, same access rules as the generation routes:

PathModality
POST /v1/tenzro/forecastsTime-series forecasting
POST /v1/tenzro/detectionsObject detection
POST /v1/tenzro/segmentationsPromptable and open-vocabulary segmentation
POST /v1/tenzro/video/embeddingsClip embedding

Any JSON-RPC method is also reachable over REST at POST /v1/rpc/{method}. Request and response shapes for each modality are in Multimodal.

Client setup

python
from openai import OpenAI

client = OpenAI(
    base_url="https://rpc.tenzro.xyz/v1",
    api_key="unused",
    default_headers={"X-Tenzro-Api-Key": "tnz_..."},
)

stream = client.chat.completions.create(
    model="qwen3-8b",
    messages=[{"role": "user", "content": "Explain metered settlement in one paragraph."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://rpc.tenzro.xyz/v1",
  apiKey: "unused",
  defaultHeaders: { "X-Tenzro-Api-Key": process.env.TENZRO_API_KEY! },
});

const res = await client.chat.completions.create({
  model: "qwen3-8b",
  messages: [{ role: "user", content: "hello" }],
});
bash
curl https://rpc.tenzro.xyz/v1/chat/completions \
  -H 'content-type: application/json' \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -d '{"model":"qwen3-8b","messages":[{"role":"user","content":"hello"}]}'

Chat completions

The standard fields are honoured: messages, max_tokens, temperature, top_p, stop, seed, frequency_penalty, presence_penalty, stream, stream_options and user, plus top_k, min_p and repetition_penalty. n must be 1: the network bills per completion.

Message content is a string or an array of parts: text, image_url (inline the bytes as a data: URI) for vision-language models, and input_audio and file for models that accept them. A part the model cannot take is refused by name rather than silently ignored.

Tenzro extensions ride in the same body. OpenAI SDKs pass unknown fields through.

models                fallback model ids, tried in order
provider              {"only": [...], "ignore": [...]} to narrow the provider pool
jurisdiction          "DE,EU": route only to providers whose locality claim matches
jurisdiction_receipt  "required": fail unless a signed locality receipt comes back
draft_n               opt in to speculative decoding on models with a drafter

The response's model field is the candidate that actually served. Responses also carry cost_wei, generation_time_ms and tokens_per_second, and, where the provider signs, a tenzro_contentProvenance manifest and a tenzro_jurisdiction receipt. See Inference for the routing behind each field.

Responses

POST /v1/responses translates the Responses shape into a chat request, runs it through the same handler and translates the result back. Routing, pinning, settlement, provenance and streaming behave identically. input becomes messages, instructions a leading system message and max_output_tokens the token limit; usage uses the Responses names input_tokens and output_tokens. Prior turns are replayed in input: previous_response_id is refused because the node retains no responses.

Streaming

Set stream: true and the response is Server-Sent Events.

  • Chat completions stream chat.completion.chunk events ending with data: [DONE]. The final chunk carries usage and the billing fields, so you learn what you were billed without opting in. stream_options: {"include_usage": true} also appends the standard empty-choices usage chunk.
  • Responses stream typed events, from response.created through response.output_text.delta to response.completed, response.incomplete or response.failed. The terminal event ends the stream; there is no [DONE].
  • Resume. Every event id is <completion_id>:<seq>. Reconnect with Last-Event-ID and the node replays the chunks you missed from a short-lived in-memory buffer.
  • Failover. If the serving provider drops mid-generation, the gateway continues the stream on another provider of the same model, replaying the text already emitted as an assistant prefix with the same seed. Pin seed for byte-identical continuation.
{"id":"chatcmpl-...","object":"chat.completion.chunk","model":"qwen3-8b",
 "choices":[{"index":0,"delta":{},"finish_reason":"stop","native_finish_reason":"eos"}],
 "usage":{"prompt_tokens":12,"completion_tokens":48,"total_tokens":60},
 "cost_wei":"...","generation_time_ms":...,"tokens_per_second":...}

data: [DONE]

finish_reason uses the OpenAI vocabulary; native_finish_reason carries the exact cause the engine reported. A client that closes the connection before the final chunk can read the same numbers later from GET /v1/generation?id=<completion_id>.

Over JSON-RPC, tenzro_chatStream streams the rich shape with reasoning and tool-call events, and POST /chat-stream serves the same rich shape as SSE.

Embeddings, transcriptions, speech and image generation return a single body. Video is a job: poll it, then download.

Model listing

Each entry in GET /v1/models carries the OpenAI fields plus the serving contract: context length, maximum output tokens, pricing in TNZO wei per token (and in USD where the operator declares a listing rate), the pricing model (PerToken, PerRequest, PerComputeTime or Dynamic), supported parameters, feature flags and the provider's declared jurisdiction. The list also carries a data_policy: prompts and completions are not written to disk, the stream resume buffer is in memory and expires, and usage accounting records metered units, cost and latency only, never content.

Errors

Errors use the OpenAI envelope, {"error": {"message", "type", "code"}}.

StatusMeaning
400Invalid request, or a field or content part the model cannot honour, named in the message
401Missing or invalid API key on a key-gated model
402Payment required; the body carries the payment challenge
404No provider serves that model, or a provider pin admitted none
409Video content requested before the job completed
412No provider satisfies the jurisdiction pin, or a required receipt is unavailable
429Rate limited
502The provider was unreachable or returned an error
504An image render outlived the wait; the message names the job to poll

Next