macOS
Universal DMG for Apple Silicon and Intel
Built by a team that has worked with global media, research, and Fortune 500 companies
Join 100k+ users across the web
Record straight from your mic or drop in a file. No account needed, save it to your account when you're ready.
Upload a file, record live, paste a link, or import from your cloud, then watch it transcribe.
Paste a link from YouTube, TikTok, Instagram, and more. We download the audio and transcribe it with speaker labels, subtitles, and translation. Nothing to upload.
From the first upload to publish-ready captions and translations, one workflow, nothing bolted on.
Explore all featuresA complete AI transcription and translation toolkit, accurate, fast, and ready for production.
Whisper VoiceKit speech models tuned for accents, jargon, and noisy rooms, so you spend less time fixing transcripts.
Automatically detects who said what and labels every speaker, even on overlapping conversations.
Transcribe in one language and translate to another in a single pass, with formatting preserved.
No setup, no manual cleanup. Bring a file, leave with publish-ready text and captions.
Drag in audio or video, MP3, WAV, M4A, MP4 and more. Or paste a link or call the API.
Our Whisper VoiceKit engine processes your file, separates speakers, and produces a clean, time-coded transcript.
Translate to 100+ languages, then export TXT, SRT, VTT, DOCX or PDF, or pull results via API.
However you work with audio, RealtimeVoiceKit fits in.
Repurpose episodes into show notes, blog posts, and captioned clips that rank and reach more people.
Thousands of podcasters, journalists, and developers use RealtimeVoiceKit every day.
“We cut our editing time in half. The speaker labels are scarily accurate, and the subtitle export drops straight into our workflow.”
“Turnaround that used to take a freelancer a full day now takes minutes. Transcripts are clean enough to publish with light edits.”
“The API was a two-line integration. Webhooks instead of polling mean I never wrote a single retry loop.”
Start free, upgrade when you need more minutes and translation.
Turn on Medical Mode for sharper accuracy on medications, procedures, conditions, and dosages, available in English, Spanish, German, and French.
Review transcripts before clinical use.
RealtimeVoiceKIT is built to protect your recordings, transcripts, and account, with encryption, granular consent, and full control over your data.
TLS in transit and encryption at rest, with bcrypt-hashed passwords and hashed API keys.
Export your data or delete your account anytime. We never sell your data or train models on your content.
Granular cookie consent, and we honor Global Privacy Control as an automatic opt-out.
We support your data rights under GDPR and CCPA, with a Data Processing Addendum and a published list of sub-processors.
Download RealtimeVoiceKIT for macOS, Windows, or Linux. Capture your microphone and supported system audio, transcribe in real time, and keep every session synced with the web.
Universal DMG for Apple Silicon and Intel
64 bit installer for Windows
x86_64 AppImage for Linux
Free to download. Sign in with the same RealtimeVoiceKIT account you use on the web.
RealtimeVoiceKIT is available on Google Play and the App Store. Record, transcribe, summarize, translate, and export subtitles from Android, iPhone, or iPad.
Download RealtimeVoiceKIT from Google Play or the App Store and sign in with the same account you use on the web.
Capture Google Meet calls, webinars, or any browser tab, record your microphone in one tap, and dictate in real time. Every session lands in your RealtimeVoiceKIT library with an AI summary.
Install it now from the Chrome Web Store. Works on Chrome and Edge.
RealtimeVoiceKIT is an AI transcription and translation platform. Upload audio or video and get an accurate, speaker-labeled transcript with subtitles, plus optional translation into 100+ languages, through the web app or a developer API. Under the hood, transcription runs on Whisper VoiceKit speech models.
Common formats including MP3, WAV, M4A, and MP4 work out of the box. You can also transcribe from a URL or send files programmatically through the API.
We transcribe and translate across 100+ languages, so you can caption and localize content for a global audience in a single workflow.
Yes. Every transcript can be exported as plain text, timestamped text, SRT, or WebVTT, ready for any video player or editing suite.
Yes. The Free plan includes 10 minutes of transcription every month with speaker labels and subtitle export, no credit card required.
Yes. Premium and Pro plans include a REST API with rtvk_ keys and webhooks, so you can add transcription, subtitles, and translation directly to your own product.
Get 10 free minutes every month. No credit card, no setup, just upload and go.
Welcome back to the show, today we're talking about real-time audio.
Thanks for having me. The accuracy on technical terms has come a long way.