Tenzro Train
Decentralised training on Tenzro Network 1: low-communication rounds over independent hardware, trust tiers, robust aggregation, a chat objective for agent trajectories and rounds replayable from their seed.
Abstract
Tenzro Train is the decentralised training protocol of Tenzro Network 1. A sponsor posts a training task and funds it. Trainers on independent hardware run an inner training loop locally and exchange compressed outer gradients once per round. A committee of syncers aggregates the gradients under a rule suited to the run's trust tier and commits the result on the ledger. Every round is replayable from its seed, so anyone can recompute what the round should have produced and challenge a contribution that does not match. Alongside supervised and reinforcement-learning objectives, a chat fine-tuning objective trains models on agent trajectories, with the loss on the assistant's turns and tool calls. Contributions settle in TNZO against bonds.
1. Introduction
Training a model has usually meant renting a large, tightly connected cluster from a single provider. That limits who can train, concentrates cost and leaves no public record of what was trained on what.
Low-communication training methods change the constraint. If each participant trains for many local steps between synchronisations and sends compressed updates, training can run over ordinary internet links between machines that do not trust each other. Tenzro Train turns that into a protocol: tasks, enrolment, rounds, aggregation, verification and settlement on a public ledger.
The protocol is aimed at the Machine Economy. Its first job is to make models better at acting as agents, by training them on the trajectories agents actually produce.
2. Roles
| Role | Does | Bonds |
|---|---|---|
| Sponsor | posts a task spec, funds it, and for confidential runs seals the dataset | the task budget, held in escrow |
| Trainer | runs the inner loop on its hardware and submits outer gradients | a trainer bond |
| Syncer | aggregates gradients, applies the outer step and finalises rounds | a syncer bond |
| Validator | orders and finalises task, round and settlement records | a validator bond |
Every participant acts with a hardware-rooted DID. A trainer enrols as a did:tenzro:machine identity.
3. Tasks
A task spec states everything needed to run and verify training:
- the base model, by manifest root, and the modality (language, vision or time series);
- the objective (section 6);
- the dataset reference and how it is sharded;
- the trust tier and aggregation rule;
- the round schedule: inner steps per round, outer optimiser and learning rate;
- the communication settings: gradient quantisation, sparsification and streaming;
- the budget and reward terms.
Posting a task is an owner-class action: it must be signed by the sponsor's account, which funds the escrow.
4. Rounds
Each round runs the same cycle.
- Seed. The round's seed is derived from the task, the round number and entropy from finalised chain state. Nobody can choose it in advance.
- Assignment. From the seed, each trainer's data shards, sampling order and any randomised choices in training are fixed. The syncer committee for the round is selected from registered syncers with the same seed.
- Inner loop. Each trainer starts from the round's committed model state and trains locally for the configured number of steps.
- Outer gradient. Each trainer computes the difference between its local result and the round's starting state, compresses it, and submits it with a commitment. The commitment includes the payload hash and signed probes of the training trajectory, such as loss values and activation samples at checkpoints the seed selects.
- Aggregation. The syncer committee aggregates the submitted gradients with the task's aggregation rule and applies the outer optimiser step.
- Finalisation. The committee finalises the round by committing the new state root. Finalisation is idempotent: redundant submissions of the same result are accepted, and conflicting results are rejected.
- Settlement. Accepted contributions are credited, and the round's receipt is recorded.
If the committee cannot form a quorum within the round's grace window, it issues a no-endorsement certificate and the run advances, carrying the previous state forward. A missing syncer delays a round; it does not stall the run.
5. Replayable rounds
Every round is replayable from its seed. Everything that determines a round's outcome is either committed on the ledger or derived from the seed:
- the starting state root;
- the shard assignment and sampling order;
- the randomised choices in training and the positions of trajectory probes;
- the task's configuration;
- the gradients submitted, by hash;
- the aggregation rule and the resulting state root.
Replayability serves three purposes.
- Verification. Anyone can recompute a trainer's contribution for a round and compare the recomputed trajectory probes with the committed ones. A contribution that does not match can be challenged. A successful challenge excludes the contribution and slashes the trainer's bond.
- Provenance. A model's training history is a chain of committed rounds. A relying party can check which data, objective and contributions produced a given state.
- Recovery. A new syncer or a restarted trainer can reconstruct the run's state from the ledger and the stored payloads.
Payloads are content-addressed. Gradients and model states are stored and fetched by hash, and each chunk is verified as it arrives.
6. Objectives
| Objective | What it trains | Data |
|---|---|---|
| Supervised | next-token or task loss on the dataset | text, images or time series |
| Chat fine-tuning | a model on agent trajectories, with the loss on assistant turns only | chat and tool-call conversations |
| Reinforcement learning | a policy with group-relative advantages from a reward function | prompts and a reward reference |
6.1 Chat fine-tuning for agent trajectories
The chat fine-tuning objective trains a language model on conversations in the OpenAI chat shape: a list of messages and the tools available. Each assistant turn, including its tool calls, is one training example, rendered with the model's own chat template against the conversation so far. The context is masked; the loss falls only on what the assistant said and did.
This is the objective for teaching a model to act as an agent: take corrected runs of an agent, including its tool calls, and fine-tune the model so it behaves well with a minimal harness. The objective supports LoRA and QLoRA adapters and alternating-freeze aggregation for adapter runs.
A dataset shard is a JSONL file of conversations:
{"messages": [
{"role": "system", "content": "You are a procurement agent."},
{"role": "user", "content": "Rent one GPU for two hours at the index price."},
{"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "get_price_index", "arguments": "{\"class\": \"workstation\"}"}}]},
{"role": "tool", "content": "{\"price\": \"...\"}"},
{"role": "assistant", "content": "Booked at the index price."}
], "tools": [{"type": "function", "function": {"name": "get_price_index"}}]}6.2 Reinforcement learning
The reinforcement-learning objective samples a group of completions per prompt, scores them with the reward function the sponsor references, computes group-relative advantages, and takes an optimiser step on a clipped objective with a KL penalty. It needs no value model and leaves the outer-gradient exchange unchanged.
7. Trust tiers and aggregation
| Tier | Who may train | Aggregation rules |
|---|---|---|
| Open | anyone with a bond | mean, and alternating-freeze for adapter runs |
| Verified | bonded trainers meeting the sponsor's credential requirements | adds Byzantine-robust rules: trimmed mean, coordinate-wise median, Krum |
| Confidential | Verified, plus an attested enclave | as Verified, with sealed data |
Robust aggregation limits what a minority of dishonest trainers can do to a round. Replay-based challenges catch contributions that do not follow from the committed data. Bonds make both costly to attempt.
7.1 Confidential runs
For a confidential run, the sponsor seals each data shard to a trainer enclave's key using HPKE, and publishes a sealed dataset manifest that names the expected enclave measurement and key. A trainer enrols with an attestation report. The report is verified against pinned vendor roots, simulated reports are refused, and the report must commit to the enclave key and measurement from the sponsor's manifest, not values the trainer supplies. The shards are decrypted only inside the attested enclave; the host never sees the plaintext.
8. Communication efficiency
Each mechanism is declared in the task and enforced by the syncers.
- Quantisation. Outer gradients are quantised blockwise to Int8 or Int4. A submission whose quantisation differs from the task's is rejected.
- Sparsification. Top-k sparsification with error feedback sends only the largest components each round and carries the remainder forward.
- Streaming synchronisation. The model is partitioned into fragments; each round synchronises a subset, overlapping communication with computation.
- Delayed application. The aggregate from one round is applied at the next, so the inner loop never waits for synchronisation.
- Adaptive outer learning rate. The outer step scales with the agreement between submitted gradients.
- Pipeline groups. Trainers can enrol as stages of a pipeline, so a group jointly holds one replica and no single trainer needs to fit the whole model.
- Inner optimiser. The reference trainer supports Muon for matrix parameters, with AdamW for the rest, which reduces the number of outer synchronisations needed.
9. Economics
- The sponsor's budget is held in escrow and released per round against accepted contributions.
- Trainers are paid for accepted gradients; syncers for finalised rounds.
- Trainers and syncers post bonds. A successful challenge, a conflicting finalisation or a failure to deliver slashes the bond, and slashed amounts are burned.
- Payments settle in TNZO. See TNZO and Settlement for the Machine Economy.
10. Interfaces
Training is driven through the node's tenzro_training_* RPC methods and the tenzro train CLI:
| Action | CLI |
|---|---|
| Post a task (sponsor) | tenzro train post-task --spec task.json |
| Enrol a trainer | tenzro train enroll-trainer --task-id <id> --trainer-did <did> |
| Submit an outer gradient | tenzro train submit-gradient |
| Decide whether a round finalises, waits or advances | tenzro train decide-round --task-id <id> |
| Finalise a round (syncer) | tenzro train finalize-round |
| Challenge a contribution | tenzro train challenge-commitment |
| Inspect runs and receipts | tenzro train list-runs, get-run, get-receipt |
| Install a sealed manifest (confidential) | tenzro train install-sealed-manifest --manifest sealed.json |
State-changing methods require a signature from the acting account. Trainer software ships as the tenzro-trainer Python package; a node with training enabled provisions the trainer automatically.
11. Design choices
- Rust protocol, Python trainer. The protocol, aggregation, commitments and settlement are implemented in the node. The inner training loop uses the Python machine-learning ecosystem, where per-architecture model code lives.
- Low communication by default. Rounds exchange compressed outer gradients rather than per-step gradients, so training runs over ordinary links.
- Evidence over trust. Seeds, commitments and replay make a contribution checkable after the fact; bonds make cheating costly.
- One identity and one ledger. Trainers, syncers and sponsors use the same hardware-rooted identities and settle on the same chain as every other service.
12. Conclusion
Tenzro Train lets anyone with hardware contribute to training, lets anyone check how a model was trained, and gives agents a direct path to better models through the chat fine-tuning objective. For hands-on detail, see Tenzro Train and Training for agents.