Skip to content
Tenzro
← All tutorials
Tutorial · Inference and training

Segment images with SAM 2

Run point- and box-prompted image segmentation on Tenzro Network 1 with SAM 2, from the CLI, the HTTP API or JSON-RPC.

Intermediate15 min

The segmentation runtime serves SAM 2 (base and large) alongside lighter SAM variants. You give it point or box prompts, the encoder computes and caches an embedding per image, and the decoder returns one mask per prompt with a quality score. In this tutorial you load SAM 2 on a node you run, prompt it from the CLI, and call a provider on the network over HTTP and JSON-RPC.

Prerequisites

  • The tenzro CLI installed. See Getting started.
  • A sample image, for example image.png.
  • To call the public endpoint: an API key (header X-Tenzro-Api-Key) or a wallet that pays per request over HTTP 402. See API keys.
  • To load a model yourself: a node you operate with the ai role. See Model serving.

1. Accept the licence on your node

SAM 2 ships under Meta's custom terms. A node refuses to load it until the operator has read and accepted those terms by starting the node with the licence id:

bash
tenzro-node --roles ai --accept-license meta-sam

2. Load SAM 2

SAM ships as an encoder and decoder pair. --catalog-id supplies the decoder layout and input resolution; --model is the id callers use. Run this against your own node, because loading is an operator action.

bash
tenzro segment catalog

tenzro segment load \
  --model seg \
  --encoder-path /models/sam2-base/encoder.onnx \
  --decoder-path /models/sam2-base/decoder.onnx \
  --catalog-id sam2-base

3. Segment with point prompts

Prompts come from a JSON file. Coordinates are original-image pixels, and is_foreground says whether the point is inside the object (true) or outside it (false).

bash
cat > prompts.json <<'EOF'
[
  {"type": "point", "x": 412, "y": 310, "is_foreground": true},
  {"type": "point", "x": 120, "y": 500, "is_foreground": false}
]
EOF

tenzro segment run \
  --model seg \
  --image image.png \
  --prompts prompts.json

Expected output (abridged):

json
{
  "masks": [
    { "width": 1024, "height": 768, "mask": [0, 0, 1, 1, "..."], "score": 0.93 }
  ],
  "generation_time_ms": 61,
  "cost_wei": "…",
  "settlement": { "status": "settled", "...": "..." }
}

4. Segment with a box prompt

A box gives an object-level mask. x0,y0 is the top-left corner and x1,y1 the bottom-right, in original-image pixels.

bash
cat > box.json <<'EOF'
[{"type": "box", "x0": 120, "y0": 80, "x1": 540, "y1": 420}]
EOF

tenzro segment run --model seg --image image.png --prompts box.json

A good pattern is to take the boxes from RF-DETR detection and pass each one as a box prompt.

5. Call a provider over HTTP

POST /v1/tenzro/segmentations takes the image as base64 and a prompts array. Masks come back base64-encoded to keep the response compact.

bash
IMG=$(base64 < image.png | tr -d '\n')

curl -s https://rpc.tenzro.xyz/v1/tenzro/segmentations \
  -H 'content-type: application/json' \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -d "{\"model\":\"seg\",\"image_base64\":\"$IMG\",\"prompts\":[{\"type\":\"point\",\"x\":412,\"y\":310,\"is_foreground\":true}]}"

The same route also serves open-vocabulary segmentation models: send text_prompt instead of prompts, naming a model that supports text.

6. Call it over JSON-RPC

tenzro_segment uses model_id, which is the id the model was loaded under, not the catalog id.

bash
curl -s https://rpc.tenzro.xyz \
  -H 'content-type: application/json' \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tenzro_segment\",\"params\":{\"model_id\":\"seg\",\"image_base64\":\"$IMG\",\"prompts\":[{\"type\":\"box\",\"x0\":120,\"y0\":80,\"x1\":540,\"y1\":420}]}}"

Each response carries cost_wei and a settlement object, so you can reconcile what you paid per call.

Next steps