Convert audio or video into a clean, structured text transcript with confidence scores.
What this worker does
Transcribes spoken audio from any source into accurate, readable text across 100+ languages. Upload an audio file (MP3, WAV, M4A, OPUS, or FLAC) and Speech-To-Text automatically detects the spoken language, returns a full timestamped transcript, and provides a confidence score for each segment. Optionally, you can specify the source language to improve accuracy for niche terminology or accented speech. Output includes plain text, JSON with timestamps, and detected language metadata.