Serve a model on Tenzro
Fetch hash-verified weights from any origin, serve the model on your node and publish it so the network can route paid inference to you.
This tutorial takes one model from download to paid traffic by hand, without the one-command tenzro join --provider. You fetch the weights, check them against their manifest, serve them behind the node's OpenAI-compatible API and publish the model to the network.
Prerequisites
- A running node with the
airole (tenzro-node --roles ai --data-dir ./data). See Join Network 1 as a provider. - A posted provider bond, or run
tenzro join --provideronce and then continue here. - Disk space for the model you choose.
1. Download the weights
tenzro model download qwen3-4b --rpc http://127.0.0.1:8545Weights on Network 1 are stored across several origins: other Tenzro nodes and storage providers, the model creator's release and public hubs. The node fetches from peers first and falls back to other origins. Whichever origin supplies the bytes, every file is checked against the model's content-addressed manifest before the node loads it, and a file that does not match is refused.
Pin one path when you need to:
tenzro model download qwen3-4b --source networkFollow a large download with tenzro model progress qwen3-4b.
2. Check the hash record
Every model carries a hash record: the BLAKE3 root of its manifest, the SHA-256 of each file and the per-file manifest.
tenzro model get-hash qwen3-4bThe model's identity on the network is that root, so certifications and ratings issued for the model refer to exactly these bytes. See Model provenance.
3. Serve it locally first
Serve the model privately to test it. A private model answers local callers only and is not announced.
tenzro model serve qwen3-4b --private
tenzro chat qwen3-4bThe node exposes the model on its OpenAI-compatible routes:
curl -s http://127.0.0.1:8545/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"qwen3-4b","messages":[{"role":"user","content":"Say hello."}]}'If the model does not fit one machine, the node forms a pipeline cluster with other willing machines on your local network. Use --force-single to prevent that or --cluster to force it. See LAN clustering.
4. Publish it to the network
Stop the private instance and serve the model publicly:
tenzro model stop qwen3-4b
tenzro model serve qwen3-4bA public model is announced to the network. Any caller can reach it by paying per request, either with an API key you issue or with an HTTP 402 payment through x402 or MPP. To serve only callers you have agreed terms with, use --gated instead.
Set your prices, in wei per unit (1 TNZO is 10^18 wei):
tenzro provider pricing set \
--input-price-wei 100000000000000 \
--output-price-wei 200000000000000
tenzro provider pricing show5. Confirm the network sees you
tenzro provider models
tenzro provider listThen send a request through the network rather than to your own node:
tenzro inference request qwen3-4b "What does a finality certificate prove?" \
--rpc https://rpc.tenzro.xyz6. Watch routing health
The router hedges against slow providers. If a provider has not answered by its observed tail latency, the same request goes to a backup and the first answer wins.
tenzro inference router-metrics
# requests total routed
# hedges_dispatched primary still pending past the delay
# hedges_won backup answered first
# deadline_exceeded abandoned on the wall-clock deadlineA rising hedges_won or deadline_exceeded against your endpoint means you are the slow tail. Serve a smaller quantisation, add hardware or lower concurrency.
Next steps
- Build a model-serving provider: several models, gated access, a signed policy and certification.
- Peer-first model fetch.
- Model serving and OpenAI-compatible API.