Transcribe audio with Whisper Turbo
Transcribe speech on Tenzro Network 1 with Whisper large-v3-turbo, through the CLI, the OpenAI-compatible transcriptions API or JSON-RPC.
The audio catalog covers Moonshine, Distil-Whisper, Whisper large-v3-turbo, Parakeet and Canary. Whisper Turbo is a strong general-purpose default; Moonshine Tiny suits edge devices. In this tutorial you load Whisper Turbo on a node you run, transcribe a file from the CLI, and call a provider through the OpenAI-compatible POST /v1/audio/transcriptions route and JSON-RPC. Speech recognition is billed per second of audio transcribed.
Prerequisites
- The
tenzroCLI installed. See Getting started. - An audio file (WAV, MP3 or FLAC), for example
meeting.wav. - To call the public endpoint: an API key (header
X-Tenzro-Api-Key) or a wallet that pays per request over HTTP 402. - To load a model yourself: a node you operate with the
airole. See Model serving.
1. Load the transcription model
--catalog-id supplies the decoding pipeline, audio window and checkpoint shape, and applies the model's licence tier. The three paths point at files already on the node. Loading is an operator action, so run it against your own node.
tenzro transcribe catalog
tenzro transcribe load \
--model asr \
--encoder-path /models/whisper-turbo/encoder_model.onnx \
--decoder-path /models/whisper-turbo/decoder_model_merged.onnx \
--tokenizer-path /models/whisper-turbo/tokenizer.json \
--catalog-id whisper-large-v3-turbo2. Transcribe a file
The runtime handles the audio preprocessing and detokenisation.
tenzro transcribe run --model asr --audio meeting.wavExpected output (abridged):
{
"text": "Let's move the training run to the new cluster before Friday.",
"segments": [],
"language": "en",
"cost_wei": "…"
}3. Pin the language and ask for timestamps
If you know the language, pass it to skip language detection. --timestamps returns per-segment start and end times.
tenzro transcribe run \
--model asr \
--audio interview.wav \
--language en \
--timestampsEach entry in segments has text, start_seconds and end_seconds.
4. Use the OpenAI-compatible route
POST /v1/audio/transcriptions takes a multipart form with file and model, and optional language, response_format, temperature, timestamp_granularities and prompt fields.
curl -s https://rpc.tenzro.xyz/v1/audio/transcriptions \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-F file=@meeting.wav \
-F model=asr \
-F language=enWithout an API key the route answers 402 Payment Required with a challenge you can pay over x402 or MPP. See Pay for inference in stablecoins.
5. Call it over JSON-RPC
tenzro_transcribe takes the audio as base64 plus optional language, timestamps and temperature.
AUDIO=$(base64 < meeting.wav | tr -d '\n')
curl -s https://rpc.tenzro.xyz \
-H 'content-type: application/json' \
-H "X-Tenzro-Api-Key: $TENZRO_API_KEY" \
-d "{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tenzro_transcribe\",\"params\":{\"model_id\":\"asr\",\"audio_base64\":\"$AUDIO\",\"language\":\"en\",\"timestamps\":true}}"The response carries the transcript, the metered units and cost_wei, and a settlement object.
Next steps
- Embed transcripts for search: Embed text with Qwen3-Embedding.
- All audio, vision and video routes: Multimodal inference.
- Keep recordings private on confidential hardware: TEE.