Skip to content
  1. Developers
  2. Models
  3. Transcription
Built by VidmoatAvailable

Vidmoat transcription

Speech to text with a timestamp on every word, ready to feed straight into an addCaptions command. It runs on our own servers using faster-whisper, detects the language when you do not give one, and caches results, so transcribing the same file again is free.

id POST /v1/ai/transcriptions

What it does well

  • A start and end time on every word.
  • Language detected automatically, or pass a hint.
  • Already-transcribed sources come back from cache at no cost.

Limits

  • 30 minutes of media per request on the API.
  • Speaker labels are not included.

Where to call it

REST APIPOST /v1/ai/transcriptionsMCPtranscribe
Chat botsTelegram, WhatsApp and Discord
Vidmoat appCaptions in the editor

Try it

curl
curl -X POST https://api.vidmoat.com/v1/ai/transcriptions \
  -H "Authorization: Bearer $VIDMOAT_KEY" \
  -H "Content-Type: application/json" \
  -d '{"src":"https://example.com/interview.mp4"}'

A test key answers with a sample and spends nothing. Every parameter and error is in the API reference.

Read more