Skip to content
Tenzro
← All tutorials
Tutorial · Inference and training

Transcribe audio with Whisper Turbo

Transcribe speech on Tenzro Network 1 with Whisper large-v3-turbo, through the CLI, the OpenAI-compatible transcriptions API or JSON-RPC.

Beginner10 min

The audio catalog covers Moonshine, Distil-Whisper, Whisper large-v3-turbo, Parakeet and Canary. Whisper Turbo is a strong general-purpose default; Moonshine Tiny suits edge devices. In this tutorial you load Whisper Turbo on a node you run, transcribe a file from the CLI, and call a provider through the OpenAI-compatible POST /v1/audio/transcriptions route and JSON-RPC. Speech recognition is billed per second of audio transcribed.

Prerequisites

  • The tenzro CLI installed. See Getting started.
  • An audio file (WAV, MP3 or FLAC), for example meeting.wav.
  • To call the public endpoint: an API key (header X-Tenzro-Api-Key) or a wallet that pays per request over HTTP 402.
  • To load a model yourself: a node you operate with the ai role. See Model serving.

1. Load the transcription model

--catalog-id supplies the decoding pipeline, audio window and checkpoint shape, and applies the model's licence tier. The three paths point at files already on the node. Loading is an operator action, so run it against your own node.

bash
tenzro transcribe catalog

tenzro transcribe load \
  --model asr \
  --encoder-path /models/whisper-turbo/encoder_model.onnx \
  --decoder-path /models/whisper-turbo/decoder_model_merged.onnx \
  --tokenizer-path /models/whisper-turbo/tokenizer.json \
  --catalog-id whisper-large-v3-turbo

2. Transcribe a file

The runtime handles the audio preprocessing and detokenisation.

bash
tenzro transcribe run --model asr --audio meeting.wav

Expected output (abridged):

json
{
  "text": "Let's move the training run to the new cluster before Friday.",
  "segments": [],
  "language": "en",
  "cost_wei": "…"
}

3. Pin the language and ask for timestamps

If you know the language, pass it to skip language detection. --timestamps returns per-segment start and end times.

bash
tenzro transcribe run \
  --model asr \
  --audio interview.wav \
  --language en \
  --timestamps

Each entry in segments has text, start_seconds and end_seconds.

4. Use the OpenAI-compatible route

POST /v1/audio/transcriptions takes a multipart form with file and model, and optional language, response_format, temperature, timestamp_granularities and prompt fields.

bash
curl -s https://rpc.tenzro.xyz/v1/audio/transcriptions \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -F file=@meeting.wav \
  -F model=asr \
  -F language=en

Without an API key the route answers 402 Payment Required with a challenge you can pay over x402 or MPP. See Pay for inference in stablecoins.

5. Call it over JSON-RPC

tenzro_transcribe takes the audio as base64 plus optional language, timestamps and temperature.

bash
AUDIO=$(base64 < meeting.wav | tr -d '\n')

curl -s https://rpc.tenzro.xyz \
  -H 'content-type: application/json' \
  -H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
  -d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tenzro_transcribe\",\"params\":{\"model_id\":\"asr\",\"audio_base64\":\"$AUDIO\",\"language\":\"en\",\"timestamps\":true}}"

The response carries the transcript, the metered units and cost_wei, and a settlement object.

Next steps