Segment images with SAM 2
Run point- and box-prompted image segmentation on Tenzro Network 1 with SAM 2, from the CLI, the HTTP API or JSON-RPC.
The segmentation runtime serves SAM 2 (base and large) alongside lighter SAM variants. You give it point or box prompts, the encoder computes and caches an embedding per image, and the decoder returns one mask per prompt with a quality score. In this tutorial you load SAM 2 on a node you run, prompt it from the CLI, and call a provider on the network over HTTP and JSON-RPC.
Prerequisites
- The
tenzroCLI installed. See Getting started. - A sample image, for example
image.png. - To call the public endpoint: an API key (header
X-Tenzro-Api-Key) or a wallet that pays per request over HTTP 402. See API keys. - To load a model yourself: a node you operate with the
airole. See Model serving.
1. Accept the licence on your node
SAM 2 ships under Meta's custom terms. A node refuses to load it until the operator has read and accepted those terms by starting the node with the licence id:
tenzro-node --roles ai --accept-license meta-sam2. Load SAM 2
SAM ships as an encoder and decoder pair. --catalog-id supplies the decoder layout and input resolution; --model is the id callers use. Run this against your own node, because loading is an operator action.
tenzro segment catalog
tenzro segment load \
--model seg \
--encoder-path /models/sam2-base/encoder.onnx \
--decoder-path /models/sam2-base/decoder.onnx \
--catalog-id sam2-base3. Segment with point prompts
Prompts come from a JSON file. Coordinates are original-image pixels, and is_foreground says whether the point is inside the object (true) or outside it (false).
cat > prompts.json <<'EOF'
[
{"type": "point", "x": 412, "y": 310, "is_foreground": true},
{"type": "point", "x": 120, "y": 500, "is_foreground": false}
]
EOF
tenzro segment run \
--model seg \
--image image.png \
--prompts prompts.jsonExpected output (abridged):
{
"masks": [
{ "width": 1024, "height": 768, "mask": [0, 0, 1, 1, "..."], "score": 0.93 }
],
"generation_time_ms": 61,
"cost_wei": "…",
"settlement": { "status": "settled", "...": "..." }
}4. Segment with a box prompt
A box gives an object-level mask. x0,y0 is the top-left corner and x1,y1 the bottom-right, in original-image pixels.
cat > box.json <<'EOF'
[{"type": "box", "x0": 120, "y0": 80, "x1": 540, "y1": 420}]
EOF
tenzro segment run --model seg --image image.png --prompts box.jsonA good pattern is to take the boxes from RF-DETR detection and pass each one as a box prompt.
5. Call a provider over HTTP
POST /v1/tenzro/segmentations takes the image as base64 and a prompts array. Masks come back base64-encoded to keep the response compact.
IMG=$(base64 < image.png | tr -d '\n')
curl -s https://rpc.tenzro.xyz/v1/tenzro/segmentations \
-H 'content-type: application/json' \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-d "{\"model\":\"seg\",\"image_base64\":\"$IMG\",\"prompts\":[{\"type\":\"point\",\"x\":412,\"y\":310,\"is_foreground\":true}]}"The same route also serves open-vocabulary segmentation models: send text_prompt instead of prompts, naming a model that supports text.
6. Call it over JSON-RPC
tenzro_segment uses model_id, which is the id the model was loaded under, not the catalog id.
curl -s https://rpc.tenzro.xyz \
-H 'content-type: application/json' \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tenzro_segment\",\"params\":{\"model_id\":\"seg\",\"image_base64\":\"$IMG\",\"prompts\":[{\"type\":\"box\",\"x0\":120,\"y0\":80,\"x1\":540,\"y1\":420}]}}"Each response carries cost_wei and a settlement object, so you can reconcile what you paid per call.
Next steps
- Embed the segmented crops for search: Embed images with DINOv3.
- All vision, audio and timeseries routes: Multimodal inference.
- Run this privately on confidential hardware: TEE.