Skip to content
Tenzro
Documentation menu
Inference

Distributed MoE

Serve very large mixture-of-experts models across many providers on Tenzro Network 1, with signed expert receipts and per-holder settlement.

Mixture-of-experts (MoE) models activate only a few expert networks for each token, so the compute per token follows the active path while the total weights can be far larger than any one machine holds. Tenzro Network 1 serves MoE models in two modes over the same provider population:

  • Full replica. A provider whose hardware holds the whole model serves it like any dense model.
  • Expert shards. For models too large for any single provider, many providers each hold a subset of the experts. A router runs the shared layers and the gating step, sends each batch of tokens to the providers holding the selected experts, and combines the results.

Expert sharding lets a network of ordinary machines serve a model that none of them could hold alone.

Roles

A provider declares one or more MoE roles:

RoleDoes
ReplicaHolds the full model and serves it on one machine. The default.
RouterRuns the gating step and the shared layers, fans batches out to expert holders and combines the outputs.
ExpertHolderHolds one or more experts and executes batches sent to them.

Holdings are part of the provider record and ride every provider heartbeat, so the network's view of who holds what converges without a separate announcement. A provider can be a pure expert holder that serves no full model of its own.

Shard map

The shard map is the live view of a model's experts across the network: which providers hold each expert, how many replicas exist, which experts are hot and which are under-replicated.

bash
tenzro moe shard-map --model-id qwen3.5-397b-a17b
tenzro moe replication-policy

Each holding records the expert's residency (warm in memory or cold on disk), the holder's measured serving throughput and the content-addressed tenzro://blob/<hash> the holder loaded the weights from. The planner ranks holders by measured throughput, not self-declared figures.

Dispatch

For each layer, the router gates the batch locally, then plans the dispatch: tokens whose routing selected the same expert on the same holder are grouped into one call. That is per-holder batching rather than per-token all-to-all traffic, which is what makes the scheme work over ordinary networks.

Batches travel over the holder's iroh QUIC endpoint when it has one, or over HTTP otherwise. Activations are compressed on the wire. Results stream back into a combine weighted by the gate probabilities, so the router does not wait on the slowest holder before starting.

Both router families in the catalog run through the same gating path: softmax top-k routing and sigmoid scoring with per-expert bias and shared experts. The gate weights describe themselves, so a holder loads either family with the same call.

Failure handling

A batch walks its planned holders warm-first. If every known holder of an expert fails, the affected tokens are replanned against a fresh shard map that excludes the failed providers, while results already gathered stay in the combine. Each failure costs the holder reputation.

By default a forward that cannot cover every token fails closed. With allow_partial, the router renormalises the gate weights over the experts that responded and reports what was missing, the distributed equivalent of serving with a reduced top-k during a partial outage.

Signed receipts and settlement

Every remote expert execution comes back signed. The holder commits to its output rows and signs the commitment together with a hash of the exact input it received, under its provider key. The router verifies each receipt before accepting the batch; a bad or missing receipt is treated as a holder failure.

A sample of verified batches is stored in full, chosen by the commitment hash itself so a holder cannot predict which ones are kept. Any node that holds the same expert can re-execute a stored batch and compare. An upheld dispute is a fraud proof against the holder's own signature and is penalised.

After each forward, the router settles the expert work per holder at that holder's advertised price, and credits the holder's reputation with the paid amount. Reputation therefore tracks paid, receipt-verified work.

Replication

The replication policy sets a minimum number of distinct holders for every active expert and a ceiling for hot experts. Governance tunes it.

Raising replication is pull-based. Each node periodically computes, by rendezvous hash over the providers already holding a model, which provider should cover each missing replica, and acts only on the rows naming itself: it fetches the expert from an existing holder's tenzro://blob/ URI and verifies it by hash. No node can make a peer allocate memory, and two nodes never race for the same replica slot.

Memory tiers and quantisation

A holder can advertise more experts than fit in memory. Experts live in a byte-bounded memory tier over a disk tier; before a forward dispatches, the experts its routing selected are read ahead into memory so they are warm when the batch arrives.

Experts can be prepared block-quantised, in GGUF-compatible Q8_0, Q4_K or Q6_K formats, per projection, so more experts stay resident in the same memory. Holders run the expert maths on CPU by default, with optional GPU backends; a holder with a GPU advertises it and the planner biases hot experts toward it.

Operating an expert holder

Loading and preparing experts decides what your hardware runs, so these are operator (admin) actions.

bash
# Slice a layer's experts out of a catalog checkpoint, quantise and publish them
tenzro moe prepare-experts --model-id qwen3.5-397b-a17b --layer 0 --quant q4_k_m
tenzro moe prepare-status --job-id <job-id>

# Load a gate and experts from content-addressed blobs
tenzro moe load-gate --model-id qwen3.5-397b-a17b --layer 0 --uri tenzro://blob/<hash>
tenzro moe load-expert --model-id qwen3.5-397b-a17b --layer 0 --expert 1 --uri tenzro://blob/<hash>

# What this node holds
tenzro moe status
JSON-RPCDoesAccess
tenzro_moeShardMapLive shard map for a modelopen
tenzro_moeReplicationPolicyCurrent replication policyopen
tenzro_moeCatalogShapeA model's expert topology from the catalogopen
tenzro_moePlanDispatchPlan how a batch of routings fans outopen
tenzro_moeExpertStatusResident experts and gates on this nodeopen
tenzro_moeListReceiptsStored execution receiptsopen
tenzro_moeExpertLoad, tenzro_moeGateLoad, unload variantsManage held expertsadmin
tenzro_moePrepareExperts, tenzro_moePrepareStatusPrepare and publish expert blobsadmin
tenzro_moeForwardRun a distributed forward for one layerowner
tenzro_moeDisputeReceiptRe-execute a stored receipt and dispute itowner

Catalog coverage

MoE families in the catalog include Qwen 3, 3.5 and 3.6 MoE models, Gemma 4 MoE, Kimi K2 and K3, DeepSeek V3 and V4, GLM 5.x, MiniMax, Nemotron Nano MoE and gpt-oss. Check a model's topology with tenzro_moeCatalogShape or tenzro moe catalog-shape --model-id <id>.