Skip to content
Tenzro
← All tutorials
Tutorial · Operate

Serve a model on Tenzro

Fetch hash-verified weights from any origin, serve the model on your node and publish it so the network can route paid inference to you.

Intermediate25 min

This tutorial takes one model from download to paid traffic by hand, without the one-command tenzro join --provider. You fetch the weights, check them against their manifest, serve them behind the node's OpenAI-compatible API and publish the model to the network.

Prerequisites

  • A running node with the ai role (tenzro-node --roles ai --data-dir ./data). See Join Network 1 as a provider.
  • A posted provider bond, or run tenzro join --provider once and then continue here.
  • Disk space for the model you choose.

1. Download the weights

bash
tenzro model download qwen3-4b --rpc http://127.0.0.1:8545

Weights on Network 1 are stored across several origins: other Tenzro nodes and storage providers, the model creator's release and public hubs. The node fetches from peers first and falls back to other origins. Whichever origin supplies the bytes, every file is checked against the model's content-addressed manifest before the node loads it, and a file that does not match is refused.

Pin one path when you need to:

bash
tenzro model download qwen3-4b --source network

Follow a large download with tenzro model progress qwen3-4b.

2. Check the hash record

Every model carries a hash record: the BLAKE3 root of its manifest, the SHA-256 of each file and the per-file manifest.

bash
tenzro model get-hash qwen3-4b

The model's identity on the network is that root, so certifications and ratings issued for the model refer to exactly these bytes. See Model provenance.

3. Serve it locally first

Serve the model privately to test it. A private model answers local callers only and is not announced.

bash
tenzro model serve qwen3-4b --private
tenzro chat qwen3-4b

The node exposes the model on its OpenAI-compatible routes:

bash
curl -s http://127.0.0.1:8545/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"qwen3-4b","messages":[{"role":"user","content":"Say hello."}]}'

If the model does not fit one machine, the node forms a pipeline cluster with other willing machines on your local network. Use --force-single to prevent that or --cluster to force it. See LAN clustering.

4. Publish it to the network

Stop the private instance and serve the model publicly:

bash
tenzro model stop qwen3-4b
tenzro model serve qwen3-4b

A public model is announced to the network. Any caller can reach it by paying per request, either with an API key you issue or with an HTTP 402 payment through x402 or MPP. To serve only callers you have agreed terms with, use --gated instead.

Set your prices, in wei per unit (1 TNZO is 10^18 wei):

bash
tenzro provider pricing set \
  --input-price-wei 100000000000000 \
  --output-price-wei 200000000000000
tenzro provider pricing show

5. Confirm the network sees you

bash
tenzro provider models
tenzro provider list

Then send a request through the network rather than to your own node:

bash
tenzro inference request qwen3-4b "What does a finality certificate prove?" \
  --rpc https://rpc.tenzro.xyz

6. Watch routing health

The router hedges against slow providers. If a provider has not answered by its observed tail latency, the same request goes to a backup and the first answer wins.

bash
tenzro inference router-metrics
# requests           total routed
# hedges_dispatched  primary still pending past the delay
# hedges_won         backup answered first
# deadline_exceeded  abandoned on the wall-clock deadline

A rising hedges_won or deadline_exceeded against your endpoint means you are the slow tail. Serve a smaller quantisation, add hardware or lower concurrency.

Next steps