Serve a distributed MoE model
Hold a subset of a Mixture-of-Experts model's experts on your node and serve it together with other providers across the network.
A Mixture-of-Experts (MoE) model routes each token to a few experts out of many. On Tenzro, those experts do not have to sit on one machine. Each provider holds a subset, the network keeps a live map of who holds what, and a forward pass fans each token out to the holders of the experts it was routed to. A machine that could never hold the whole model can still earn on it.
The planner batches per holder: tokens whose top-k routing landed on the same expert and holder go out in one call. Calls travel over the holder's QUIC endpoint on the peer-to-peer data plane when it advertises one, and over HTTP otherwise.
Prerequisites
- A node with the
airole and a posted provider bond. See Join Network 1 as a provider. - Enough memory for the experts you intend to hold. A GPU is recommended.
- This tutorial uses
qwen3-30b-a3b; the same steps apply to any catalog MoE model with a per-expert checkpoint source.
1. See the model's expert topology
Start with the catalog-side shape: how many experts the model has, how many are active per token, and the layer layout.
tenzro moe catalog-shape --model-id qwen3-30b-a3b2. Prepare the experts you will hold
Extract a layer's experts from the catalog checkpoint, optionally block-quantise each projection, and publish each expert as a content-addressed blob. A preset (q4_k_m, q8_0, q4_k, q6_k) applies one format to all three FFN projections; --quant-json sets each projection independently. Omit both to publish dense f32.
tenzro moe prepare-experts \
--model-id qwen3-30b-a3b --layer 0 \
--experts 1,2,5 \
--quant q4_k_m
# returns a job id; poll it:
tenzro moe prepare-status --job-id <job-id>When the job completes, prepare-status lists a tenzro://blob/<hash> URI per expert and one for the layer's gating network. The block layout is GGUF-compatible, so a holder decodes a prepared expert with the same path it uses for a local GGUF file. Because each blob is addressed by its hash, anyone fetching it verifies every byte.
3. Load the experts and the gate
Load each expert's gate, up and down projections into the node's expert runtime, keyed by model, layer and expert. Use the blob URI from step 2 or a local safetensors file. Load the gating network for every layer you take part in.
tenzro moe load-expert \
--model-id qwen3-30b-a3b --layer 0 --expert 1 \
--uri tenzro://blob/<expert-hash>
tenzro moe load-gate \
--model-id qwen3-30b-a3b --layer 0 \
--uri tenzro://blob/<gate-hash>4. Confirm what is resident
tenzro moe statusEach expert and gate reports its residency tier (Warm in memory or Cold on disk), its size, the memory budget and whether a GPU backend is active. A holder can advertise more experts than fit in memory: the runtime keeps them in a size-bounded cache over a disk tier and promotes the experts a forward is about to use back to warm before it dispatches.
5. Read the network-wide shard map
tenzro moe shard-map --model-id qwen3-30b-a3b
tenzro moe replication-policyThe shard map aggregates every holder's declaration: expert coverage, replication per expert, which experts are hot and which are under-replicated. The replication policy, set by governance, says how many distinct providers each active expert should have. Holding an under-replicated expert is the fastest way to be useful to the network.
6. Inspect a dispatch plan
Before running a request you can see how it would fan out. Write a JSON file of top-k routing decisions:
[
{ "token_index": 0, "experts": [{ "layer": 0, "expert": 1 }, { "layer": 0, "expert": 5 }] },
{ "token_index": 1, "experts": [{ "layer": 0, "expert": 2 }, { "layer": 0, "expert": 5 }] }
]Then ask the planner for the per-holder batches it would send:
tenzro moe plan-dispatch \
--model-id qwen3-30b-a3b \
--routings-json ./routings.json7. Run a distributed layer forward
A distributed forward gates locally, hands the routing decisions to the planner, sends each per-holder batch (run locally when this node holds the expert, remotely otherwise) and recombines the outputs weighted by the gate probabilities.
tenzro moe forward \
--model-id qwen3-30b-a3b --layer 0 \
--d-model 2048 \
--hidden-file ./hidden.f32 \
--out-file ./combined.f32--hidden-file holds raw little-endian f32 hidden states, one row of d-model values per token. Add --allow-cold to permit dispatch to experts that are not warm.
Next steps
- Distributed MoE for the planner and the shard map in depth.
- LAN clustering for pipeline clusters on a local network.
- Prefix and state reuse for how MoE and other architectures reuse context.