Skip to content
Tenzro
Documentation menu
Inference

Model serving

Run a model provider on Tenzro Network 1: join with one command, serve any open model, publish a signed policy and earn TNZO per use.

A model provider is a node that serves inference to the network and is paid per use. The machine can be a home PC with a GPU, a homelab, a rack in an independent data centre or a fleet in a neo-cloud. The network finds your hardware, routes requests to you, meters what you serve and settles it on the chain.

The quick path

With a tenzro-node running on the machine, one command does the whole onboarding:

bash
tenzro-node --roles ai --data-dir ./data
tenzro join --provider

tenzro join --provider detects your hardware (see Hardware discovery), posts your provider bond, registers you as a provider, sets default pricing, then downloads and serves the largest catalog model that fits the machine. Point it at a remote node with --rpc if the node runs elsewhere.

A node can take several roles under one identity. Add ai to an existing validator or storage node with --roles validator,ai. See Operators and roles.

The manual path

Choose a model from the registry, fetch it and serve it:

bash
tenzro model download qwen3-8b
tenzro model serve qwen3-8b

The download is peer-first and hash-verified: the node fetches the weights from any origin, checks them against the catalog's pinned hash and refuses to load anything that does not match.

Set your prices and post your bond:

bash
tenzro provider pricing set \
  --input-price-wei <wei-per-input-token> \
  --output-price-wei <wei-per-output-token>

tenzro provider bond post \
  --did did:tenzro:machine:<your-node> \
  --address <your-payout-address> \
  --amount <tnzo>

Pricing is yours to set. Beyond per-token rates, tenzro provider pricing set accepts a per-request floor, a per-millisecond compute rate, separate rates for cached-read and cached-write prompt tokens (see Prefix and state reuse), and rates for image tokens, audio seconds and video seconds. Check your configuration with tenzro provider pricing show and your status with tenzro provider status.

Deciding what a node serves is an operator action. Serving, stopping and deleting models, and changing prices or schedules, are admin methods: they need your node's admin token when you drive them over RPC.

Visibility

FlagWho can call the model
none (default)Anyone. The model is announced to the network, and any caller can use it by paying, with no prior relationship.
--gatedOnly callers holding an API key whose policy you agreed in advance. The model is not announced.
--privateOnly you, over a direct or LAN connection. No announcement and no presence in provider discovery.
bash
tenzro model serve qwen3-8b --gated
tenzro model serve my-finetune --private

Signed operator policy

Every provider publishes a signed policy: what it serves, where its hardware sits (declared jurisdiction), whether it runs inside a TEE, how it handles data, and the service levels it commits to. Callers and routers read the policy before they send work, and certification issuers can rate you against it.

You are slashed only for a provable breach of your own published policy, never for anything you did not commit to. Commit to what your hardware and operations actually deliver. See Operator policies and Slashing.

Confidential serving

If your machine has confidential-computing hardware (Intel TDX, AMD SEV-SNP, AWS Nitro or NVIDIA confidential GPUs), you can serve inside the enclave. The node produces attestation evidence that is fully verified per vendor, and callers who need confidentiality route only to attested providers. A TEE is evidence about where the computation ran, not a place where keys are held. See TEE.

Large models

A model too large for one machine can still be served:

  • Across your own machines. tenzro model serve forms a LAN cluster automatically when the model does not fit the biggest single machine on your segment. Use --cluster to force a split and --force-single to prevent one. Preview the plan with tenzro cluster preview <model-id>.
  • Across the network. Mixture-of-experts models can be sharded by expert across many providers. See Distributed MoE.

External engines

You can front an OpenAI-compatible engine you already run, such as vLLM, SGLang or llama-server, instead of loading weights in the node:

bash
tenzro model serve qwen3-8b --engine vllm --base-url http://127.0.0.1:8000

The node handles discovery, admission, metering and settlement; the engine handles batching and GPU memory.

Sealed distribution of private weights

Private weights move between nodes as encrypted shards, never in cleartext. The owner seals the artifact for named recipients, each identified by a DID and an X25519 public key, optionally pinned to an enclave measurement. The installing node verifies the manifest signature, every shard hash and the plaintext hash before anything reaches model storage, and a node pinned to an enclave measurement must prove that measurement first.

bash
# Recipient: publish the node's recipient key
tenzro model recipient-key

# Owner: seal and export the manifest (admin)
tenzro model seal my-finetune --path /models/my-finetune.gguf \
  --owner-did did:tenzro:machine:<owner> \
  --recipient "did:tenzro:machine:<recipient>=<x25519-pubkey-hex>"
tenzro model sealed-get my-finetune --out my-finetune.manifest.json

# Recipient: install and serve privately (admin)
tenzro model install-sealed --manifest-file my-finetune.manifest.json \
  --recipient-did did:tenzro:machine:<recipient>
tenzro model serve my-finetune --private

Discovery, health and schedule

Every served model appears in GET /v1/models with its price, context length, features and your declared jurisdiction. Routers probe providers continuously; successful, settled work raises your reputation and failures lower it. Reputation only rises through paid work, so it tracks real service.

Limit serving to certain hours if the machine has other jobs:

bash
tenzro schedule set --start 09:00 --end 21:00 --tz UTC
tenzro schedule show

Getting paid

Inference is metered per use by default: tokens, images, audio and video seconds, and denoising work for media. Each call settles in TNZO, or in stablecoins where the caller pays that way, and your share goes to the payout address bound to your provider identity. See SLA attestation and metering and Settlement.

Next