Voz.
10 Minuten Audio in 2 s auf einem iPhone transkribieren, 4,7x schneller als Whisper, mit genauen Wort-Zeitstempeln.
10 Minuten Audio in 2 s auf einem iPhone transkribieren.
Transkribieren Sie eine Audio- oder Videodatei mit Voz und erhalten Sie genaue Start- und Endzeiten für jedes Wort. 10 Minuten brauchen 2 s auf einem iPhone, auf dem Gerät, also wird nichts hochgeladen und nichts pro Minute abgerechnet.
Auf Apple silicon läuft das Modell auf der Neural Engine, sodass die Transkription blitzschnell ist und weniger Strom verbraucht. Auf einem Mac läuft ein Archiv aus Podcasts, Interviews und Vorlesungen in einem Durchlauf durch, und nichts wird hochgeladen.
Die Zeitstempel sind präzise genug zum Schneiden. Wortanfänge liegen innerhalb von 83 ms eines Forced Aligners und Wortenden innerhalb von 95 ms, für präzises transkriptbasiertes Schneiden. Voz ist eine für die Apple Neural Engine optimierte Version von NVIDIAs Parakeet TDT 0.6B v3, mit einer Genauigkeit, die bei den meisten Audios mit Whisper large-v3-turbo vergleichbar ist, bei einer Downloadgröße von nur 467 MB.
4,7x schneller als Whisper.
10 Minuten Erzählung in 2 s und 30 Minuten in 6 s auf einem iPhone 17 Pro, 7 s auf einem iPhone 15 Pro. Auf einem Mac 4,7x schneller als whisper.cpp large-v3-turbo bei demselben Podcast-Audio, mit einer Wortfehlerrate von 7,40 % über sechs öffentliche englische Sätze, in denen Whisper 7,00 % erreicht.
Wortfehlerrate auf den Open-ASR-Leaderboard-Datensätzen, Voz mit eigenem Text-Normalizer, Whispers Werte aus dem Leaderboard
| Datensatz | Voz | Whisper large-v3-turbo |
|---|---|---|
| LibriSpeech test-clean | 2,19 % | 2,13 % |
| LibriSpeech test-other | 3,86 % | 3,70 % |
| GigaSpeech | 9,70 % | 8,47 % |
| SPGISpeech | 3,86 % | 2,79 % |
| Earnings-22 | 12,97 % | 11,07 % |
| AMI | 11,84 % | 13,87 % |
| Average | 7,40 % | 7,00 % |
Lesen Sie die Zeile, die zu Ihrem Audio passt. Saubere, nah mikrofonierte Sprache liegt bei etwa 2–4 % (LibriSpeech, SPGISpeech), Podcasts und Web-Video bei etwa 10 % (GigaSpeech) und Besprechungsräume und Telefonanrufe bei 12–13 % (AMI, Earnings-22), und genau dort schlägt das Modell Whisper. Die Werte pro Sprache bei jeweils 10 Minuten gelesener Sprache stehen auf der Modellkarte, von 3,31 % für Italienisch bis 39,46 % für Griechisch. Die iPhone-Zeiten und der whisper.cpp-Vergleich sind unsere eigenen Läufe und stehen noch nicht auf der Karte.
Jede Audio- oder Videodatei transkribieren
Nehmen Sie einen Podcast oder ein Video auf, und das vollständige Transkript ist fertig, bevor der Export abgeschlossen ist. Durchsuchbarer Text mit einem Zeitstempel an jedem Wort, in 25 Sprachen, ohne das Gerät zu verlassen.
Einen Rückstand abarbeiten
Ein Archiv aus Interviews, Vorlesungen oder Meeting-Aufnahmen wird in einem Durchlauf auf einem Mac durchsuchbar. Hundert Stunden laufen in 20 Minuten durch, und nichts wird hochgeladen.
Auf ein Wort schneiden
Jedes Wort trägt einen Anfang und ein Ende, sodass ein ausgewählter Bereich im Transkript ein Schnitt in einem Video-Editor ist. Kombinieren Sie es mit Clips für die Shorts und Title, um sie zu benennen.
Inspiration
Ideen zum Bauen mit Voz. Kopieren Sie einen Prompt in Ihren Coding-Agenten und los.
Build a podcast player that transcribes every episode and makes it searchable, on the device.
Build a podcast player. When an episode downloads, run Desert Ant's Voz on-device to transcribe it with word timestamps, index the text, and let the user search across a whole feed and tap a result to jump to that moment. No cloud transcription, no per-minute bill. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build a meeting recorder that hands you a timestamped transcript before you leave the room.
Build a meeting recorder. When recording stops, Desert Ant's Voz transcribes the file in seconds on iPhone or Mac, so the timestamped transcript is ready right away and nothing is uploaded. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Turn voice memos into searchable, editable text notes, offline.
Build a voice-notes app where each memo is transcribed on-device with Desert Ant's Voz into editable, searchable text. Works on a plane, and the audio never leaves the phone. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Generate SRT subtitles for any video file on the Mac.
Build a small Mac tool that takes a video file, transcribes the audio with Desert Ant's Voz (Apple silicon), and writes a correctly timed .srt file. No upload, no API key. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Add captions to a video player that run on the device.
Transcribe the video's audio on-device with Desert Ant's Voz (a fast batch pass on iPhone or Mac), then render a synced caption track. Works offline and for private content that can't go to a cloud service. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build an interview app for journalists where sources never leave the phone.
Build an interview-recording app that transcribes on-device with Desert Ant's Voz, so a source's audio and words stay on the reporter's phone. Add speaker-friendly formatting and export. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build a full podcast studio: transcribe, clip, and title on the device.
Build a podcast tool that runs Desert Ant's Voz to transcribe an episode, Clips to pick the best moments as shorts, and Title to name and describe each clip, all on-device with no per-minute or per-token bill. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Clips (Swift). Install and API: https://desertant.com/docs/clips/. Title (Swift). Install and API: https://desertant.com/docs/title/. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Record, clean, transcribe, and clip: a whole creator pipeline, offline.
Build a creator pipeline: Desert Ant's Clear cleans the recording, Voz transcribes it, and Clips pulls the shorts, all on the device so a full episode never touches a server. Build it with the Desert Ant SDK. Clear (Swift, Kotlin, JavaScript / TypeScript). Install and API: https://desertant.com/docs/clear/. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Clips (Swift). Install and API: https://desertant.com/docs/clips/. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Was das Modell macht
- Wortanfänge innerhalb von 83 ms und Wortenden innerhalb von 95 ms eines Forced Aligners, im Mittel, bei 80 ms Frame-Auflösung.
- 10 Minuten Audio in 2 s und 30 Minuten in 6 s auf einem iPhone 17 Pro; 290x Echtzeit auf einem M3 Ultra. Einzelne kurze Clips laufen mit 50–62x, weil jeder Clip ein volles 15-s-Fenster bezahlt.
- 7,40 % Wortfehlerrate im Mittel über sechs Open-ASR-Leaderboard-Sätze (Whisper large-v3-turbo: 7,00 %), 2,83 % bei einer halben Stunde Erzählung (Whisper: 2,72 %), 11,84 % bei AMI-Meetings (Whisper: 13,87 %).
- Der gesamte Graph läuft auf der Neural Engine ohne CPU- oder GPU-Fallback; der Spitzenspeicher wächst nicht mit der Länge der Aufnahme.
- 25 Sprachen: Bulgarisch, Kroatisch, Tschechisch, Dänisch, Niederländisch, Englisch, Estnisch, Finnisch, Französisch, Deutsch, Griechisch, Ungarisch, Italienisch, Lettisch, Litauisch, Maltesisch, Polnisch, Portugiesisch, Rumänisch, Russisch, Slowakisch, Slowenisch, Spanisch, Schwedisch und Ukrainisch.
- 467 MB auf der Festplatte, bei Bedarf geladen und im Cache gehalten; der erste Ladevorgang nach einem Download dauert einmalig 20 s, während Core ML spezialisiert, danach 0,2 s. Laden Sie beim Onboarding herunter.
- Mono-Audio in beliebiger Abtastrate; das SDK rechnet die Abtastrate um und mischt herunter. Längeres Audio wird an Pausen in 15-s-Fenster geschnitten und an den Wörtern zusammengefügt, auf die sich benachbarte Fenster einigen.
- Basiert auf NVIDIAs Parakeet TDT 0.6B v3, veröffentlicht unter CC BY 4.0. Die Gewichte sind unverändert; die Konvertierung, die Kompression und die Neural-Engine-Laufzeit stammen von uns.
Beispiel-App
Clipper
Gebaut mit Voz, Clips, Title
Generate short clips from a video podcast or a long recording, fully on device. A macOS app and a command-line tool over the same core.
Voz transcribes with a time on every word, Clips ranks the best spans, and Title writes a title and description for each.
macOS 26 or later, Apple Silicon
brew tap desert-ant-labs/tap
brew install --cask clipper
Erste Schritte
Fügen Sie spracherkennung mit wenigen Zeilen Code zu Ihrer iOS or macOS-App hinzu. Voz Docs.
// Swift Package Manager
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0")
// target dependency
.product(name: "Voz", package: "desert-ant-core")
import Voz
let voz = try await Voz()
let result = try await voz.transcribe(url)
result.text // the transcript
result.words.first?.start // 80ms resolution
result.realtimeFactor // seconds of audio per second of wall clock
Add Voz from Desert Ant Labs to this Swift project (iOS, macOS). What it does: On-Device-Sprache-zu-Text auf der Neural Engine. SDK: Swift (iOS, macOS) Repo: https://github.com/Desert-Ant-Labs/desert-ant-core#readme // Swift Package Manager .package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0") // target dependency .product(name: "Voz", package: "desert-ant-core") Reference: - Model page: https://desertant.com/models/voz/ - Full catalog and other models: https://desertant.com/llms.txt Add the SDK, then follow its README for the exact API and current version. Do not invent API names or method signatures; confirm them against the README.
Technische Daten
- Geschwindigkeit
- 10 Minuten Audio in 2 s auf einem iPhone 17 Pro, 30 Minuten in 6 s (7 s auf einem iPhone 15 Pro); 30 Minuten in 5,6 s auf einem M3 Ultra, 319x Echtzeit
- Genauigkeit
- 7,40 % WER über sechs Open-ASR-Leaderboard-Sätze; 2,83 % Langform; Wortanfänge 83 ms, Enden 95 ms mittlerer Fehler
- Größe auf dem Gerät
- 467 MB kompiliertes Core ML, bei Bedarf geladen
- Sprachen
- 25 europäische Sprachen; Genauigkeit von 3,31 % (Italienisch) bis 39,46 % (Griechisch) bei gelesener Sprache
- Modell
- NVIDIA Parakeet TDT 0.6B v3 (CC BY 4.0), von Desert Ant Labs zu Core ML konvertiert und komprimiert: Log-Mel-Frontend, Conformer-Encoder, Transducer-Decoder, alles auf der Neural Engine
- Plattformen
- iOS, iPadOS, macOS, tvOS, visionOS (Core ML); nur Apple
Voz ist nur für Apple: Die Laufzeit steuert Core ML direkt an, um den Graphen auf der Neural Engine zu halten, und es gibt keinen Build für Android, Linux oder Web. Voz weiß nicht, welche seiner 25 Sprachen es gerade hört, und eine nicht abgedeckte Sprache erzeugt selbstbewussten Unsinn statt eines Fehlers, kombinieren Sie es also mit Ear.
Die Genauigkeit schwankt stark je nach Sprache, Wortenden sind die schwierigere Hälfte der Zeitstempel, und 467 MB sind ein echter Download, den man beim Onboarding holt.
FAQ
Was ist Voz?
10 Minuten Audio in 2 s auf einem iPhone transkribieren, 4,7x schneller als Whisper, mit genauen Wort-Zeitstempeln.
Läuft Voz auf dem Gerät?
Ja. Voz läuft auf dem Gerät, ohne Serveraufruf, sodass die Daten beim Nutzer bleiben.
Welche Plattformen unterstützt Voz?
Voz wird als natives On-Device-SDK für Swift ausgeliefert.
Was kostet Voz?
Jedes Modell ist kostenlos für bis zu 100k monatlich aktive Geräte pro SDK. Unbegrenzte Inferenz pro Nutzer. Kontaktieren Sie uns für individuelle Lizenzen.
Wie genau oder schnell ist Voz?
10 Minuten Erzählung in 2 s und 30 Minuten in 6 s auf einem iPhone 17 Pro, 7 s auf einem iPhone 15 Pro. Auf einem Mac 4,7x schneller als whisper.cpp large-v3-turbo bei demselben Podcast-Audio, mit einer Wortfehlerrate von 7,40 % über sechs öffentliche englische Sätze, in denen Whisper 7,00 % erreicht.
Läuft Voz auf Android oder Windows?
Noch nicht. Voz läuft heute auf Apple silicon, wo das ganze Modell auf der Neural Engine läuft. Wir rechnen damit, Android und Windows in den kommenden Monaten hinzuzufügen. Sagen Sie uns, was Sie bauen, und wir melden uns, sobald sie verfügbar sind.