Product Text to Speech Voice Cloning Voice Design Audiobook Studio Speech to Text Voice Library Use cases Faceless YouTube Short-form & UGC Video ads Audiobooks eLearning All use cases Resources Guides How-to Compare FAQ Pricing Start free Log in

Speech to Text: Transcribe Audio and Video in 30 Languages

Upload a recording or a video, get a transcript with a timestamp on every word, and export it as TXT, SRT or VTT. Auto-detects the language among 30; files up to 200 MB. Billed at 1 credit per character, from the same credits as Text to Speech.

3,000 free credits ≈ three minutes of speech · No credit card required · One-click Google sign-up

00:00:00,000 → 00:00:02,640Welcome back to the channel. Today we're breaking down

00:00:02,640 → 00:00:05,120the three habits that quietly drain your focus,

00:00:05,120 → 00:00:07,300and the one fix that took me a week to notice.

Illustrative SRT output: word timings grouped into cues of up to 42 characters. Real files carry your own text.

How it works

  1. Upload a fileDrop in audio (WAV, MP3, M4A, FLAC, OGG, Opus, AAC, WebM) or video (MP4, MOV, MKV, AVI, WMV, FLV, MPEG, M4V) up to 200 MB. Long recordings are processed in segments.
  2. TranscribeLet the language auto-detect or pick one of 30. The transcript comes back with a timestamp on every word.
  3. Edit and exportFix any misheard word in the app, then export TXT, SRT or VTT. Subtitle cues are grouped at up to 42 characters per line.

What makes LeapFun speech to text different

30 languages, auto-detected

English, Chinese (Mandarin), Cantonese, Japanese, Korean, Spanish, French, German, Russian, Italian, Portuguese, Arabic, Indonesian, Thai, Vietnamese, Turkish, Hindi, Malay, Dutch, Swedish, Danish, Finnish, Polish, Czech, Filipino, Persian, Greek, Romanian, Hungarian and Macedonian. Leave the language on auto and the transcriber identifies it; set it explicitly when a recording mixes accents or is mostly names.

Speech to Text vs. Studio subtitles

Speech to TextAudiobook Studio subtitles
InputAudio or video you already haveA script you narrate in LeapFun
How the timing is madeRecognition, word by wordFrom the narration timeline — no recognition, exact text
Best forInterviews, recordings, videos from other tools, repurposing old contentCaptions for audio generated here
Billing1 credit per transcript characterIncluded in the Studio export on every plan (free tier adds an attribution line)
ExportTXT, SRT, VTTSRT, VTT alongside WAV/MP3

What's free, what's paid

PlanPriceCreditsWhat it unlocks
Free$03,000 credits on sign-up + 1,000 each active monthTranscripts for personal and evaluation use · SRT/VTT carry an attribution line
Starter$6 / mo40,000 credits / monthClean subtitle files · commercial rights
Creator$19 / mo200,000 credits / monthRoughly 20 hours of talks a month · unlimited Studio projects

Transcription is billed at 1 credit per character of the transcript on every plan. Speech runs about 900 characters a minute, so 1,000 credits ≈ 1 minute of talk — the same rule of thumb as narration. Full tiers, yearly pricing and Pro / Team plans on the pricing page.

Frequently asked questions

Is LeapFun speech to text free?

Yes to start. Your 3,000 sign-up credits cover about three minutes of speech, and every active month adds 1,000 more. On the free tier the SRT/VTT export ends with a short attribution line; paid plans export clean files and include commercial rights.

How is transcription billed?

One credit per character of the transcript, on every plan, from the same credit pool as Text to Speech — there is no per-file charge. A 10-minute talk is roughly 1,500 words, about 9,000 characters, so about 9,000 credits.

Which file types and sizes are supported?

Audio: WAV, MP3, M4A, FLAC, OGG, Opus, AAC and WebM. Video: MP4, MOV, MKV, AVI, WMV, FLV, MPEG and M4V. Up to 200 MB per file; long recordings are transcribed in segments, so a 10-minute file is routine.

Do I get word-level timestamps?

Yes. Every word carries a start time, which is what makes the SRT/VTT export usable as captions; cues are grouped at up to 42 characters and a 0.8-second gap.

What is the difference between Speech to Text and Studio subtitles?

Speech to Text transcribes audio you upload — it listens. Studio subtitles are generated from the narration timeline of a LeapFun project, so the text is exact and no recognition is involved. Use Speech to Text for recordings you did not generate here.

Can I edit the transcript before exporting?

Yes. The app shows the text, the word timeline and the cue lines; correct any word, then export TXT, SRT or VTT.

Ready to transcribe?

Sign up in one click, get 3,000 credits, and turn your first recording into a timed transcript in minutes.

Transcribe a file