Free tools

WebM to text

Upload a WebM audio or video file and get a clean transcript with speaker labels, timestamps, and an AI summary. Handles the Opus and Vorbis streams that browsers and web tools produce.

What this tool does

WebM is what the browser writes when a web app records audio through the MediaRecorder API. Discord voice messages, Loom-style screen recorders, in-browser demo recorders, and many web meeting tools all export WebM (Opus audio, VP9 video). Upload the .webm to RecordX and you get the full transcript, speaker turns, and a summary - no conversion step needed.

The free tier lets you upload files up to 5 minutes long with no signup required to try. Create a free account and you get 60 transcription minutes a month for longer files, exports, and history. Same engine that powers the paid plans.

How it works

  1. Upload the file. Sign up at recordx.io, open the web app, and drag your file into the upload area.
  2. Wait for the transcript. Short files come back in seconds; longer files a few minutes. Speaker separation and the AI summary run automatically.
  3. Download or ask the AI. Read, edit, and export the transcript, copy the summary and action items, or ask questions of the transcript in the built-in editor.

Formats supported

WebM with Opus or Vorbis audio, both audio-only .webm and video .webm with an audio track. WebM often lands with an inaccurate reported duration - RecordX handles that transparently.

What kinds of WebM files work

Beyond file uploads

The same RecordX account handles new conversations as they happen:

FAQ

Why is my WebM file's duration wrong?

WebM files from MediaRecorder sometimes have no duration header. RecordX detects and handles this - the transcript is still complete and correctly timestamped.

Should I convert WebM to MP3 first?

No. Upload WebM as-is. Converting first only adds a lossy step.

Can I upload a video-only WebM with no audio?

The audio track is what gets transcribed. A video WebM without an audio track has nothing to transcribe.

Does the transcript include the speaker who is on screen and the one who isn't?

Diarization runs on audio, so both on-screen and off-screen speakers are labelled as separate turns.

Start transcribing free