Voz.
Transcribeer 10 minuten audio in 2s op een iPhone, 4,7x sneller dan Whisper, met precieze woordtijdstempels.
Transcribeer 10 minuten audio in 2s op een iPhone.
Transcribeer een audio- of videobestand met Voz en krijg precieze begin- en eindtijden voor elk woord. 10 minuten kost 2s op een iPhone, op het apparaat, dus er wordt niets geüpload en niets per minuut afgerekend.
Op Apple silicon draait het model op de Neural Engine, dus transcriberen is razendsnel en verbruikt minder stroom. Op een Mac gaat een backcatalogus met podcasts, interviews en colleges er in één pass doorheen, en er wordt niets geüpload.
De tijdstempels zijn precies genoeg om op te monteren. Woordbegin landt binnen 83 ms van een forced aligner en woordeinde binnen 95 ms, voor precieze montage op basis van het transcript. Voz is een voor de Apple Neural Engine geoptimaliseerde versie van NVIDIA's Parakeet TDT 0.6B v3, met een nauwkeurigheid die op de meeste audio vergelijkbaar is met Whisper large-v3-turbo, en met een downloadgrootte van slechts 467 MB.
4,7x sneller dan Whisper.
10 minuten voordracht in 2s en 30 minuten in 6s op een iPhone 17 Pro, 7s op een iPhone 15 Pro. Op een Mac 4,7x sneller dan whisper.cpp large-v3-turbo op dezelfde podcast-audio, met een woordfoutpercentage van 7,40% over zes publieke Engelse sets waar Whisper 7,00% scoort.
Woordfoutpercentage op de Open ASR Leaderboard-datasets, Voz met zijn eigen tekstnormalisator, de cijfers van Whisper uit het leaderboard
| Dataset | Voz | Whisper large-v3-turbo |
|---|---|---|
| LibriSpeech test-clean | 2,19% | 2,13% |
| LibriSpeech test-other | 3,86% | 3,70% |
| GigaSpeech | 9,70% | 8,47% |
| SPGISpeech | 3,86% | 2,79% |
| Earnings-22 | 12,97% | 11,07% |
| AMI | 11,84% | 13,87% |
| Average | 7,40% | 7,00% |
Lees de rij die bij jouw audio past. Schone, dichtbij opgenomen spraak komt rond 2 tot 4% uit (LibriSpeech, SPGISpeech), podcasts en webvideo rond 10% (GigaSpeech), en vergaderruimtes en telefoongesprekken 12 tot 13% (AMI, Earnings-22), en daar verslaat het model Whisper. Cijfers per taal, elk op 10 minuten voorgelezen spraak, staan op de model card, van 3,31% voor Italiaans tot 39,46% voor Grieks. De iPhone-timings en de vergelijking met whisper.cpp zijn onze eigen runs en staan nog niet op de kaart.
Transcribeer elk audio- of videobestand
Neem een podcast of een video op en het volledige transcript is klaar voordat de export klaar is. Doorzoekbare tekst met een tijdstempel op elk woord, in elk van 25 talen, zonder het apparaat te verlaten.
Werk een achterstand weg
Een archief met interviews, colleges of vergaderopnames wordt in één pass doorzoekbaar op een Mac. Honderd uur gaat er in 20 minuten doorheen, en er wordt niets geüpload.
Knippen op een woord
Elk woord draagt een begin en een eind, dus een geselecteerd bereik in het transcript is een knip in een video-editor. Combineer het met Clips voor de shorts en Title om ze een naam te geven.
Inspiratie
Ideeën om te bouwen met Voz. Kopieer een prompt naar je coding-agent en ga.
Build a podcast player that transcribes every episode and makes it searchable, on the device.
Build a podcast player. When an episode downloads, run Desert Ant's Voz on-device to transcribe it with word timestamps, index the text, and let the user search across a whole feed and tap a result to jump to that moment. No cloud transcription, no per-minute bill. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build a meeting recorder that hands you a timestamped transcript before you leave the room.
Build a meeting recorder. When recording stops, Desert Ant's Voz transcribes the file in seconds on iPhone or Mac, so the timestamped transcript is ready right away and nothing is uploaded. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Turn voice memos into searchable, editable text notes, offline.
Build a voice-notes app where each memo is transcribed on-device with Desert Ant's Voz into editable, searchable text. Works on a plane, and the audio never leaves the phone. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Generate SRT subtitles for any video file on the Mac.
Build a small Mac tool that takes a video file, transcribes the audio with Desert Ant's Voz (Apple silicon), and writes a correctly timed .srt file. No upload, no API key. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Add captions to a video player that run on the device.
Transcribe the video's audio on-device with Desert Ant's Voz (a fast batch pass on iPhone or Mac), then render a synced caption track. Works offline and for private content that can't go to a cloud service. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build an interview app for journalists where sources never leave the phone.
Build an interview-recording app that transcribes on-device with Desert Ant's Voz, so a source's audio and words stay on the reporter's phone. Add speaker-friendly formatting and export. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build a full podcast studio: transcribe, clip, and title on the device.
Build a podcast tool that runs Desert Ant's Voz to transcribe an episode, Clips to pick the best moments as shorts, and Title to name and describe each clip, all on-device with no per-minute or per-token bill. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Clips (Swift). Install and API: https://desertant.com/docs/clips/. Title (Swift). Install and API: https://desertant.com/docs/title/. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Record, clean, transcribe, and clip: a whole creator pipeline, offline.
Build a creator pipeline: Desert Ant's Clear cleans the recording, Voz transcribes it, and Clips pulls the shorts, all on the device so a full episode never touches a server. Build it with the Desert Ant SDK. Clear (Swift, Kotlin, JavaScript / TypeScript). Install and API: https://desertant.com/docs/clear/. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Clips (Swift). Install and API: https://desertant.com/docs/clips/. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Wat het model doet
- Woordbegin binnen 83 ms en woordeinde binnen 95 ms van een forced aligner, gemiddeld, bij een frameresolutie van 80 ms.
- 10 minuten audio in 2s en 30 minuten in 6s op een iPhone 17 Pro; 290x realtime op een M3 Ultra. Losse korte clips halen 50 tot 62x, omdat elke clip een volledig venster van 15s betaalt.
- 7,40% woordfoutpercentage gemiddeld over zes Open ASR Leaderboard-sets (Whisper large-v3-turbo: 7,00%), 2,83% op een half uur voordracht (Whisper: 2,72%), 11,84% op AMI-vergaderingen (Whisper: 13,87%).
- De hele graph draait op de Neural Engine zonder CPU- of GPU-fallback; het piekgeheugen groeit niet mee met de lengte van de opname.
- 25 talen: Bulgaars, Kroatisch, Tsjechisch, Deens, Nederlands, Engels, Ests, Fins, Frans, Duits, Grieks, Hongaars, Italiaans, Lets, Litouws, Maltees, Pools, Portugees, Roemeens, Russisch, Slowaaks, Sloveens, Spaans, Zweeds en Oekraïens.
- 467 MB op schijf, op verzoek gedownload en gecachet; de eerste keer laden na een download duurt eenmalig 20s terwijl Core ML specialiseert, daarna 0,2s. Download tijdens de onboarding.
- Mono-audio met elke samplerate; de SDK resamplet en mixt naar beneden. Langere audio wordt bij pauzes in vensters van 15s geknipt en samengevoegd op de woorden waarover aangrenzende vensters het eens zijn.
- Gebouwd op NVIDIA's Parakeet TDT 0.6B v3, uitgebracht onder CC BY 4.0. De gewichten zijn ongewijzigd; de conversie, de compressie en de Neural Engine-runtime zijn van ons.
Voorbeeld-app
Clipper
Gebouwd met Voz, Clips, Title
Generate short clips from a video podcast or a long recording, fully on device. A macOS app and a command-line tool over the same core.
Voz transcribes with a time on every word, Clips ranks the best spans, and Title writes a title and description for each.
macOS 26 or later, Apple Silicon
brew tap desert-ant-labs/tap
brew install --cask clipper
Aan de slag
Voeg spraakherkenning toe aan je iOS or macOS-app in een paar regels code. Voz docs.
// Swift Package Manager
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0")
// target dependency
.product(name: "Voz", package: "desert-ant-core")
import Voz
let voz = try await Voz()
let result = try await voz.transcribe(url)
result.text // the transcript
result.words.first?.start // 80ms resolution
result.realtimeFactor // seconds of audio per second of wall clock
Add Voz from Desert Ant Labs to this Swift project (iOS, macOS). What it does: on-device spraak-naar-tekst op de Neural Engine. SDK: Swift (iOS, macOS) Repo: https://github.com/Desert-Ant-Labs/desert-ant-core#readme // Swift Package Manager .package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0") // target dependency .product(name: "Voz", package: "desert-ant-core") Reference: - Model page: https://desertant.com/models/voz/ - Full catalog and other models: https://desertant.com/llms.txt Add the SDK, then follow its README for the exact API and current version. Do not invent API names or method signatures; confirm them against the README.
Specs
- Snelheid
- 10 minuten audio in 2s op een iPhone 17 Pro, 30 minuten in 6s (7s op een iPhone 15 Pro); 30 minuten in 5,6s op een M3 Ultra, 319x realtime
- Nauwkeurigheid
- 7,40% WER over zes Open ASR Leaderboard-sets; 2,83% long-form; woordbegin 83 ms, woordeinde 95 ms gemiddelde fout
- Grootte op het apparaat
- 467 MB gecompileerde Core ML, op verzoek gedownload
- Talen
- 25 Europese talen; nauwkeurigheid van 3,31% (Italiaans) tot 39,46% (Grieks) op voorgelezen spraak
- Model
- NVIDIA Parakeet TDT 0.6B v3 (CC BY 4.0), geconverteerd naar Core ML en gecomprimeerd door Desert Ant Labs: log-mel-frontend, conformer-encoder, transducer-decoder, allemaal op de Neural Engine
- Platforms
- iOS, iPadOS, macOS, tvOS, visionOS (Core ML); alleen Apple
Voz is alleen voor Apple: de runtime stuurt Core ML rechtstreeks aan om de graph op de Neural Engine te houden, en er is geen build voor Android, Linux of web. Voz weet niet welke van zijn 25 talen het hoort, en een taal die het niet dekt levert met stelligheid onzin op in plaats van een fout, dus combineer het met Ear.
De nauwkeurigheid verschilt sterk per taal, woordeinden zijn de lastigere helft van de tijdstempels, en 467 MB is een echte download om tijdens de onboarding op te halen.
FAQ
Wat is Voz?
Transcribeer 10 minuten audio in 2s op een iPhone, 4,7x sneller dan Whisper, met precieze woordtijdstempels.
Draait Voz op het apparaat?
Ja. Voz draait op het apparaat, zonder serveraanroep, dus de data blijft bij de gebruiker.
Welke platforms ondersteunt Voz?
Voz wordt geleverd als een native on-device SDK voor Swift.
Hoeveel kost Voz?
Elk model is gratis voor maximaal 100k maandelijks actieve apparaten per SDK. Onbeperkte inference per gebruiker. Neem contact op voor aangepaste licenties.
Hoe nauwkeurig of snel is Voz?
10 minuten voordracht in 2s en 30 minuten in 6s op een iPhone 17 Pro, 7s op een iPhone 15 Pro. Op een Mac 4,7x sneller dan whisper.cpp large-v3-turbo op dezelfde podcast-audio, met een woordfoutpercentage van 7,40% over zes publieke Engelse sets waar Whisper 7,00% scoort.
Draait Voz op Android of Windows?
Nog niet. Voz draait vandaag op Apple silicon, waar het hele model op de Neural Engine draait. We verwachten Android en Windows in de komende maanden toe te voegen. Vertel ons wat je bouwt, dan laten we het je weten zodra ze er zijn.