The larger Whisper Small English model is used by default. It takes longer, but handles accents, punctuation, and imperfect audio better than tiny browser models.
On-device transcription
Choose an audio or video recording to create an editable transcript with timestamps.
Recognition modelWhisper Small • accuracy modeAbout 410 MB on first use. Cached by Safari when storage allows.
AI
Choose a recording to begin.Waiting for device check
0%YOUR RECORDING STAYS HERE.
Only the speech-model files are downloaded. Your selected audio or video is decoded and transcribed inside Safari on this device.
“
Your words will appear here.
Choose a recording, press Create Transcript, and keep this tab open while the iPad works.
On Safari 26 or newer, WebGPU gives the AI direct access to the Apple GPU. If unavailable, the tool falls back to WebAssembly instead of uploading your file.
Edit the transcript, inspect timestamps, copy the text, or download plain text and subtitle files. Names should always be checked before publication.