Skip to content
Tenzro
← All tutorials
Tutorial · Inference and training

Embed text with Qwen3-Embedding

Create text embeddings on Tenzro Network 1 with Qwen3-Embedding and EmbeddingGemma, through the CLI, the OpenAI-compatible embeddings API or JSON-RPC.

Beginner10 min

Qwen3-Embedding 0.6B is a fast, permissively licensed text embedder and a good default for retrieval. EmbeddingGemma 300M adds Matryoshka truncation, so you can store shorter vectors with controlled quality loss. Both are served by the text-embedding runtime and exposed on the OpenAI-compatible POST /v1/embeddings route, which returns the familiar OpenAI response shape.

Prerequisites

  • The tenzro CLI installed. See Getting started.
  • To call the public endpoint: an API key (header X-Tenzro-Api-Key) or a wallet that pays per request over HTTP 402. See API keys.
  • To load a model yourself: a node you operate with the ai role.

1. Load a text embedder on your node

Pick a model from the catalog. Passing a catalog id makes the node fetch the ONNX graph and tokenizer into its models directory and register the encoder under that same id. The weights are hash-verified on arrival. Loading is an operator action, so run it against your own node.

bash
tenzro embed-text catalog

tenzro embed-text load --model qwen3-embedding-0.6b

2. Embed a string

--normalize returns an L2-normalised vector, so cosine similarity reduces to a dot product. Repeat --input to embed several strings in one call.

bash
tenzro embed-text run \
  --model qwen3-embedding-0.6b \
  --input "decentralised GPU capacity for inference" \
  --input "rent a machine by the hour" \
  --normalize

Expected output (abridged):

json
{
  "embeddings": [[0.0132, -0.0419, "..."], [0.0087, -0.0236, "..."]],
  "dim": 1024,
  "cost_wei": "…"
}

3. Embed queries and documents differently

Retrieval models prefix queries and documents with different instructions. Mark search queries with --role query and, optionally, describe the task; documents use the default document role.

bash
tenzro embed-text run \
  --model qwen3-embedding-0.6b \
  --role query \
  --task "Find providers that match a compute request" \
  --input "H100 capacity in a confidential enclave" \
  --normalize

4. Truncate with Matryoshka dimensions

EmbeddingGemma supports compact dimensions (768, 512, 256, 128) and renormalises after truncation, which helps when you store millions of vectors.

bash
tenzro embed-text load --model embeddinggemma-300m

tenzro embed-text run \
  --model embeddinggemma-300m \
  --input "compact vector for retrieval" \
  --requested-dim 256 \
  --normalize

EmbeddingGemma is under the Gemma terms, so the node must be started with --accept-license gemma first.

5. Use the OpenAI-compatible route

POST /v1/embeddings accepts a string or an array of strings in input. dimensions requests Matryoshka truncation. The Tenzro extensions normalize, input_type (query or document) and task map to the options above.

bash
curl -s https://rpc.tenzro.xyz/v1/embeddings \
  -H 'content-type: application/json' \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -d '{"model":"qwen3-embedding-0.6b","input":["hello","world"],"normalize":true}'

The response uses the OpenAI shape: data is a list of { "object": "embedding", "index", "embedding" }. Only the float encoding is served.

6. Call it over JSON-RPC

tenzro_textEmbed takes an inputs array, so one call embeds a whole batch.

bash
curl -s https://rpc.tenzro.xyz \
  -H 'content-type: application/json' \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tenzro_textEmbed","params":{"model_id":"qwen3-embedding-0.6b","inputs":["hello","world"],"normalize":true}}'

Next steps