On-device audio models
Small models for speech, on the device: studio sound, filler words cut, the language named, who said what. Nothing uploaded.
Every audio model runs on the chip already in the phone, so a five-minute recording is done in seconds and the audio never leaves the device. No per-minute bill, no upload, no queue.
Available models
On-device speech enhancement that gives your recording the warm, close-miked sound of a podcast studio, processed 100% on-device.
On-device filler detection that marks every um, uh and hmm to within 20ms, an hour of audio in 12s.
Word timestamps on device, for any transcriber. Cut, caption, and highlight on the word, at half the error, from a 0.7MB model.
Clip selection that turns a long recording into ranked shorts and highlights from a transcript. On-device on iPhone or Mac.
On-device spoken language identification across 99 languages, from a short stretch of audio, before transcription starts.
Transcribe 10 minutes of audio in 2s on an iPhone, 4.7x faster than Whisper, accurate word timestamps.
Available soon
Pricing
Every model is free up to 100k monthly active devices per SDK. Unlimited inference per user.