Inference
How Tenzro Network 1 routes an inference request: provider selection by price, latency, location and trust, confidential inference in TEEs and metered billing.
Inference on Tenzro Network 1 is a market. Many independent providers serve the same open models at their own prices, from their own hardware, under their own signed policies. When you send a request, the router picks the provider that best matches what you asked for, the provider serves it, and the network meters and settles the call. You call one endpoint; the network does the rest.
How a request flows
- Resolve the model. You name a model, or describe the job and let intent routing (below) choose one.
- Filter providers. The router keeps only providers that serve the model and satisfy your hard constraints: jurisdiction, confidential hardware, trusted certifications, a pinned provider, a price ceiling.
- Score and pick. Among the remaining offers it applies your strategy: price, latency, reputation or a weighted blend.
- Serve. The request runs on the chosen provider, over the peer-to-peer network or HTTP.
- Meter and settle. The call is billed on what it consumed, at the price the offer advertised, and settled on the chain.
Provider selection
| Strategy | Picks |
|---|---|
| price | The cheapest matching offer. |
| latency | The provider with the lowest observed latency. |
| reputation | The provider with the best record of paid, completed work. |
| weighted (default) | A blend of price, latency and reputation. |
Constraints apply before the strategy, so a cheap offer that fails a constraint is never considered.
- Location. Providers declare a jurisdiction: a country code plus optional regulatory blocs. On TEE hardware the claim is bound to the attestation; otherwise it is operator-asserted, and the receipt says which. Pin a request with
jurisdiction: "DE,EU"and the router only uses providers whose claim matches. Pinning fails closed: a provider with no claim never matches, and routing never falls back to an unpinned provider. Addjurisdiction_receipt: "required"to fail the call unless a signed receipt binding the request, the response and the claim comes back. - Trust. Every provider publishes a signed operator policy, and any issuer can certify or rate models and operators. You choose which issuers you trust and route only to providers and models that carry their credentials. See Trust and provenance.
- Confidential hardware. Require attested TEE execution for sensitive workloads (see below).
- Provider pinning. Name a provider, or give lists of providers to use or to avoid.
tenzro chat qwen3-8b --jurisdiction DE,EU --require-jurisdiction-receiptThe router also prefers a provider on your own local network when one serves the model, and it biases toward providers that already hold your prompt's prefix, which cuts time to first token and prefill cost. See Prefix and state reuse.
Intent routing
Intent routing selects the model for you. State a use case, a budget, a quality floor and where to sit between cheapest and strongest; the network picks the model, then hands off to provider selection.
use_case chat | code | reasoning | research | summarize | extract | embed
budget per-request cost cap, in wei
optimize 0.0 cheapest ... 1.0 strongest
quality_floor cheap | strong
payer_address the wallet whose balance is the hard ceilingtenzro_routeIntent returns the selection without running anything: the model, its tier, the estimated cost, a fallback chain, the reason, and the provider whose offer won. tenzro_chatByIntent resolves the same way and dispatches in one call, pinned to the offer it scored, so the price quoted is the price settled. When the node has an embedding model loaded, selection also accounts for how hard the prompt is, using each model's observed error rate on similar prompts.
tenzro inference route --use-case code --optimize 0.3
tenzro chat --use-case chat --optimize 0.6Confidential inference
For sensitive prompts, regulated data or licensed weights, route to providers that serve inside a hardware enclave: Intel TDX, AMD SEV-SNP, AWS Nitro or NVIDIA confidential GPUs. The provider's attestation evidence is fully verified per vendor, and TEE attestations can be verified on-chain. Inference verification on Tenzro is TEE-first: the attestation tells you what hardware and what software served your request.
Every response can also carry a signed provenance manifest over the output, the model and the provider. Pass require_signed: true and the call fails unless the manifest verifies against the provider's registered key. Retrieve a manifest later by content hash with tenzro_getContentProvenance.
Reliability
- Hedging. If the chosen provider has not answered by its own typical tail latency, the router sends the same request to the next-best provider and keeps whichever answers first. Only the winner bills you. Opt out per request if you need strict single dispatch.
- Failover. If both fail, the router retries on the next-best providers.
- Deadlines. Bound a request with a wall-clock deadline and the router abandons it cleanly instead of waiting on a straggler.
- Streaming failover. If a provider drops mid-stream, the gateway continues the stream on another provider serving the same model. See OpenAI-compatible API.
Operators can watch routing with tenzro_getRouterMetrics or tenzro inference router-metrics.
Metered billing
Every call is metered per use by default and settled on the chain. You pay the price the provider advertised, measured in the dimensions the call actually consumed: input and output tokens, cached prompt tokens, image tokens, audio and video seconds, or denoising work for media. The network's share is carved out of that price rather than added to it, so the quote is the bill. Pricing for a forwarded request comes from the offer the router scored, never from the provider's response, so a provider cannot re-price after serving.
You can pay in TNZO or in stablecoins, per request over HTTP 402 (x402 or MPP), through a micropayment channel, or with an API key billed to your account. Developers can add their own margin on top by passing a registered app_id. See Machine Economy and Payments.
Look up what any completion cost after the fact with GET /v1/generation?id=<completion_id> or tenzro_getGeneration.
Making a request
Most applications use the OpenAI-compatible API. The native JSON-RPC call is tenzro_chat, which is an owner method: it bills the caller, so it is signed by the paying account.
{
"jsonrpc": "2.0",
"id": 1,
"method": "tenzro_chat",
"params": {
"model": "qwen3-8b",
"message": "Summarise the attached policy in three bullet points.",
"max_tokens": 256,
"jurisdiction": "EU",
"require_signed": true
}
}From the CLI:
tenzro chat qwen3-8b
tenzro inference request qwen3-8b "hello" --require-teeDiscovery methods are open: tenzro_listModels, tenzro_listProviders, tenzro_listModelEndpoints and GET /v1/models. See RPC access.
Beyond chat
- Images, audio, video, embeddings, vision, detection, segmentation and forecasting: Multimodal.
- Very large mixture-of-experts models across many providers: Distributed MoE.
- Planning a goal across models, skills and tools:
tenzro_orchestrate, which checks the whole plan's estimated cost against the payer's balance before any step runs.