Embed images with DINOv3
Create image embeddings on Tenzro Network 1 with DINOv3 for similarity, retrieval and clustering, through the CLI, /v1/embeddings or JSON-RPC.
The vision runtime serves the DINOv3, SigLIP2 and CLIP families. DINOv3 ViT-B/16 gives self-supervised image embeddings that work well for similarity search, retrieval and clustering. In this tutorial you load it on a node you run, embed and compare images from the CLI, and call a provider over the OpenAI-compatible embeddings route and JSON-RPC.
Prerequisites
- The
tenzroCLI andjqinstalled. - Two sample images, for example
a.pngandb.png. - To call the public endpoint: an API key (header
X-Tenzro-Api-Key) or a wallet that pays per request over HTTP 402. - To load a model yourself: a node you operate with the
airole. See Model serving.
1. Accept the licence and load DINOv3
DINOv3 is under Meta's custom terms, so start your node with the licence id before loading it:
tenzro-node --roles ai --accept-license dinov3Then load the model. The ONNX graph must already be on the node's disk; --catalog-id supplies the input size, embedding dimension and normalisation. Loading is an operator action, so run it against your own node.
tenzro embed-image catalog
tenzro embed-image load \
--model img \
--path /models/dinov3-vitb16.onnx \
--catalog-id dinov3-vitb162. Embed an image
PNG, JPEG and WebP are decoded, resized and normalised on the node. --normalize returns an L2-normalised vector, so cosine similarity becomes a dot product.
tenzro embed-image run --model img --image ./a.png --normalizeExpected output (abridged):
{
"embedding": [0.0213, -0.0078, "..."],
"dim": 768,
"generation_time_ms": 17,
"cost_wei": "…"
}3. Compare two images
similarity computes cosine similarity over two vectors of the same length. Extract each vector with jq, because the command reads a bare JSON array.
tenzro embed-image run --model img --image ./a.png --normalize | jq '.embedding' > a.json
tenzro embed-image run --model img --image ./b.png --normalize | jq '.embedding' > b.json
tenzro embed-image similarity \
--image-embedding a.json \
--text-embedding b.jsonThe flag names come from the cross-modal case, but any two vectors from the same embedding space work, and two DINOv3 image embeddings are exactly that. Vectors of different dimensions are refused.
4. Use the OpenAI-compatible route
POST /v1/embeddings accepts typed content parts. An image part carries base64 bytes and is embedded by the vision model named in model. One request is either all text or all images.
IMG=$(base64 < a.png | tr -d '\n')
curl -s https://rpc.tenzro.xyz/v1/embeddings \
-H 'content-type: application/json' \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-d "{\"model\":\"img\",\"input\":[{\"type\":\"image\",\"image_base64\":\"$IMG\"}],\"normalize\":true}"The response uses the OpenAI embeddings shape, with one data entry per image.
5. Call it over JSON-RPC
tenzro_imageEmbed takes model_id (the id you loaded under, not the catalog id) and base64 image bytes.
curl -s https://rpc.tenzro.xyz \
-H 'content-type: application/json' \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tenzro_imageEmbed\",\"params\":{\"model_id\":\"img\",\"image_base64\":\"$IMG\",\"normalize\":true}}"The JSON-RPC response adds a settlement object that records how the call was paid.
Next steps
- Embed text for the same index: Embed text with Qwen3-Embedding.
- Store and query vectors: Databases.
- All vision routes: Multimodal inference.