Voz.
iPhone에서 10분 오디오를 2초에 전사, Whisper보다 4.7배 빠르며 정확한 단어 타임스탬프.
iPhone에서 10분 오디오를 2초에 전사합니다.
Voz로 오디오나 비디오 파일을 전사하면 모든 단어의 정확한 시작·종료 시각을 얻습니다. iPhone에서 10분이 2초 걸리며, 기기에서 처리되므로 업로드되는 것도, 분당 청구되는 것도 없습니다.
Apple silicon에서 모델은 Neural Engine에서 돌아가므로, 전사가 번개처럼 빠르고 전력을 덜 씁니다. Mac에서는 팟캐스트, 인터뷰, 강의로 이뤄진 지난 자료가 한 번에 처리되고, 업로드되는 것이 없습니다.
타임스탬프는 편집에 쓸 만큼 정밀합니다. 단어 시작은 강제 정렬 기준의 83밀리초 이내, 단어 끝은 95밀리초 이내에 들어가, 전사 기반의 정밀한 편집이 가능합니다. Voz는 NVIDIA의 Parakeet TDT 0.6B v3를 Apple Neural Engine에 최적화한 버전으로, 대부분의 오디오에서 Whisper large-v3-turbo에 필적하는 정확도를 내며, 다운로드 크기는 467MB에 불과합니다.
Whisper보다 4.7배 빠릅니다.
iPhone 17 Pro에서 10분짜리 내레이션을 2초에, 30분을 6초에 (iPhone 15 Pro에서 7초). Mac에서는 같은 팟캐스트 오디오로 whisper.cpp large-v3-turbo보다 4.7배 빠르며, Whisper가 7.00%를 내는 여섯 개 공개 영어 집합에서 단어 오류율 7.40%입니다.
Open ASR Leaderboard 데이터셋에서의 단어 오류율, Voz는 자체 텍스트 정규화기 사용, Whisper 수치는 리더보드에서.
| 데이터셋 | Voz | Whisper large-v3-turbo |
|---|---|---|
| LibriSpeech test-clean | 2.19% | 2.13% |
| LibriSpeech test-other | 3.86% | 3.70% |
| GigaSpeech | 9.70% | 8.47% |
| SPGISpeech | 3.86% | 2.79% |
| Earnings-22 | 12.97% | 11.07% |
| AMI | 11.84% | 13.87% |
| Average | 7.40% | 7.00% |
당신의 오디오에 맞는 행을 읽으세요. 깨끗하고 가까이 마이크로 잡은 음성은 2~4%(LibriSpeech, SPGISpeech), 팟캐스트와 웹 영상은 10% 안팎(GigaSpeech), 회의실과 전화 통화는 12~13%(AMI, Earnings-22)이고, 바로 이 구간에서 모델이 Whisper를 앞섭니다. 각 10분짜리 읽기 음성에 대한 언어별 수치는 이탈리아어 3.31%부터 그리스어 39.46%까지 모델 카드에 있습니다. iPhone 타이밍과 whisper.cpp 비교는 우리 자체 실행값이며 아직 카드에 없습니다.
어떤 오디오나 비디오 파일이든 전사
팟캐스트나 영상을 녹음하면 내보내기가 끝나기 전에 전체 전사가 준비됩니다. 25개 언어 중 어느 것으로든 모든 단어에 타임스탬프가 붙은 검색 가능한 텍스트를, 기기를 벗어나지 않고요.
밀린 자료 처리하기
인터뷰, 강의, 회의 녹음 보관함이 Mac에서 한 번에 검색 가능해집니다. 100시간이 20분 만에 처리되고, 업로드되는 것이 없습니다.
단어에 맞춰 자르기
모든 단어에 시작과 끝이 붙으므로, 전사에서 선택한 구간이 영상 편집기에서 하나의 컷이 됩니다. 쇼츠는 Clips와, 이름 붙이기는 Title과 짝지으세요.
영감
Voz(으)로 만들 아이디어. 프롬프트를 코딩 에이전트에 복사해 시작하세요.
Build a podcast player that transcribes every episode and makes it searchable, on the device.
Build a podcast player. When an episode downloads, run Desert Ant's Voz on-device to transcribe it with word timestamps, index the text, and let the user search across a whole feed and tap a result to jump to that moment. No cloud transcription, no per-minute bill. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build a meeting recorder that hands you a timestamped transcript before you leave the room.
Build a meeting recorder. When recording stops, Desert Ant's Voz transcribes the file in seconds on iPhone or Mac, so the timestamped transcript is ready right away and nothing is uploaded. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Turn voice memos into searchable, editable text notes, offline.
Build a voice-notes app where each memo is transcribed on-device with Desert Ant's Voz into editable, searchable text. Works on a plane, and the audio never leaves the phone. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Generate SRT subtitles for any video file on the Mac.
Build a small Mac tool that takes a video file, transcribes the audio with Desert Ant's Voz (Apple silicon), and writes a correctly timed .srt file. No upload, no API key. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Add captions to a video player that run on the device.
Transcribe the video's audio on-device with Desert Ant's Voz (a fast batch pass on iPhone or Mac), then render a synced caption track. Works offline and for private content that can't go to a cloud service. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build an interview app for journalists where sources never leave the phone.
Build an interview-recording app that transcribes on-device with Desert Ant's Voz, so a source's audio and words stay on the reporter's phone. Add speaker-friendly formatting and export. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Install: SwiftPM desert-ant-core. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Build a full podcast studio: transcribe, clip, and title on the device.
Build a podcast tool that runs Desert Ant's Voz to transcribe an episode, Clips to pick the best moments as shorts, and Title to name and describe each clip, all on-device with no per-minute or per-token bill. Build it with the Desert Ant SDK. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Clips (Swift). Install and API: https://desertant.com/docs/clips/. Title (Swift). Install and API: https://desertant.com/docs/title/. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
Record, clean, transcribe, and clip: a whole creator pipeline, offline.
Build a creator pipeline: Desert Ant's Clear cleans the recording, Voz transcribes it, and Clips pulls the shorts, all on the device so a full episode never touches a server. Build it with the Desert Ant SDK. Clear (Swift, Kotlin, JavaScript / TypeScript). Install and API: https://desertant.com/docs/clear/. Voz (Swift). Install and API: https://desertant.com/docs/voz/. Clips (Swift). Install and API: https://desertant.com/docs/clips/. SDK source: https://github.com/Desert-Ant-Labs/desert-ant-core. Machine-readable catalog of every model and SDK: https://desertant.com/llms.txt.
모델이 하는 일
- 강제 정렬기 대비 단어 시작은 평균 83밀리초 이내, 단어 끝은 95밀리초 이내로, 80밀리초 프레임 해상도에서.
- iPhone 17 Pro에서 10분 오디오를 2초에, 30분을 6초에; M3 Ultra에서 실시간의 290배. 짧은 단일 클립은 50~62배인데, 모든 클립이 15초짜리 창 하나를 온전히 치르기 때문입니다.
- 여섯 개 Open ASR Leaderboard 집합에 걸친 평균 단어 오류율 7.40%(Whisper large-v3-turbo: 7.00%), 30분짜리 내레이션에서 2.83%(Whisper: 2.72%), AMI 회의에서 11.84%(Whisper: 13.87%).
- 그래프 전체가 CPU나 GPU 폴백 없이 Neural Engine에서 돌아가고, 최대 메모리가 녹음 길이에 따라 늘지 않습니다.
- 25개 언어: 불가리아어, 크로아티아어, 체코어, 덴마크어, 네덜란드어, 영어, 에스토니아어, 핀란드어, 프랑스어, 독일어, 그리스어, 헝가리어, 이탈리아어, 라트비아어, 리투아니아어, 몰타어, 폴란드어, 포르투갈어, 루마니아어, 러시아어, 슬로바키아어, 슬로베니아어, 스페인어, 스웨덴어, 우크라이나어.
- 디스크에서 467MB, 필요할 때 내려받아 캐시됩니다; 다운로드 후 첫 로드는 Core ML이 특화하는 동안 한 번 20초가 걸리고, 이후에는 0.2초입니다. 온보딩 중에 다운로드하세요.
- 임의의 샘플레이트의 모노 오디오; SDK가 리샘플링하고 다운믹스합니다. 긴 오디오는 멈춤 지점에서 15초짜리 창으로 잘리고, 이웃한 창이 합의하는 단어에서 이어 붙입니다.
- CC BY 4.0으로 공개된 NVIDIA의 Parakeet TDT 0.6B v3를 기반으로 합니다. 가중치는 그대로이며, 변환, 압축, Neural Engine 런타임이 우리의 것입니다.
예제 앱
Clipper
Voz, Clips, Title로 제작
Generate short clips from a video podcast or a long recording, fully on device. A macOS app and a command-line tool over the same core.
Voz transcribes with a time on every word, Clips ranks the best spans, and Title writes a title and description for each.
macOS 26 or later, Apple Silicon
brew tap desert-ant-labs/tap
brew install --cask clipper
시작하기
음성 인식을(를) 몇 줄의 코드로 iOS or macOS 앱에 추가하세요. Voz 문서.
// Swift Package Manager
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0")
// target dependency
.product(name: "Voz", package: "desert-ant-core")
import Voz
let voz = try await Voz()
let result = try await voz.transcribe(url)
result.text // the transcript
result.words.first?.start // 80ms resolution
result.realtimeFactor // seconds of audio per second of wall clock
Add Voz from Desert Ant Labs to this Swift project (iOS, macOS). What it does: Neural Engine에서 도는 온디바이스 음성-텍스트 변환. SDK: Swift (iOS, macOS) Repo: https://github.com/Desert-Ant-Labs/desert-ant-core#readme // Swift Package Manager .package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.1.0") // target dependency .product(name: "Voz", package: "desert-ant-core") Reference: - Model page: https://desertant.com/models/voz/ - Full catalog and other models: https://desertant.com/llms.txt Add the SDK, then follow its README for the exact API and current version. Do not invent API names or method signatures; confirm them against the README.
사양
- 속도
- iPhone 17 Pro에서 10분 오디오를 2초에, 30분을 6초에 (iPhone 15 Pro에서 7초); M3 Ultra에서 30분을 5.6초에, 실시간의 319배
- 정확도
- 여섯 개 Open ASR Leaderboard 집합에서 WER 7.40%; 롱폼 2.83%; 단어 시작 평균 오차 83밀리초, 끝 95밀리초
- 온디바이스 크기
- 컴파일된 Core ML 467MB, 필요할 때 다운로드
- 언어
- 25개 유럽 언어; 읽기 음성에서 정확도는 3.31%(이탈리아어)부터 39.46%(그리스어)까지
- 모델
- NVIDIA Parakeet TDT 0.6B v3 (CC BY 4.0)를 Desert Ant Labs가 Core ML로 변환·압축: 로그-멜 프런트엔드, 콘포머 인코더, 트랜스듀서 디코더, 전부 Neural Engine에서
- 플랫폼
- iOS, iPadOS, macOS, tvOS, visionOS (Core ML); Apple 전용
Voz는 Apple 전용입니다: 런타임이 그래프를 Neural Engine에 두려고 Core ML을 직접 구동하며, Android, Linux, 웹 빌드가 없습니다. Voz는 자신이 듣는 것이 25개 언어 중 무엇인지 알지 못하고, 지원하지 않는 언어에는 오류가 아니라 자신 있는 헛소리를 내놓으므로 Ear와 함께 쓰세요.
정확도는 언어마다 크게 다르고, 단어의 끝은 타임스탬프에서 더 어려운 절반이며, 467MB는 온보딩 중에 받기에 만만치 않은 다운로드입니다.
자주 묻는 질문
Voz은(는) 무엇인가요?
iPhone에서 10분 오디오를 2초에 전사, Whisper보다 4.7배 빠르며 정확한 단어 타임스탬프.
Voz은(는) 온디바이스로 실행되나요?
네. Voz은(는) 서버 호출 없이 기기에서 실행되므로, 데이터는 사용자의 기기에 그대로 남습니다.
Voz은(는) 어떤 플랫폼을 지원하나요?
Voz은(는) Swift용 네이티브 온디바이스 SDK로 제공됩니다.
Voz의 비용은 얼마인가요?
모든 모델은 SDK당 월간 활성 기기 10만 대까지 무료입니다. 사용자당 추론은 무제한입니다. 맞춤 라이선스는 문의하기 바랍니다.
Voz은(는) 얼마나 정확하고 빠른가요?
iPhone 17 Pro에서 10분짜리 내레이션을 2초에, 30분을 6초에 (iPhone 15 Pro에서 7초). Mac에서는 같은 팟캐스트 오디오로 whisper.cpp large-v3-turbo보다 4.7배 빠르며, Whisper가 7.00%를 내는 여섯 개 공개 영어 집합에서 단어 오류율 7.40%입니다.
Voz가 Android나 Windows에서 실행되나요?
아직은 아닙니다. Voz는 오늘 Apple silicon에서 실행되며, 거기서는 모델 전체가 Neural Engine에서 돌아갑니다. 앞으로 몇 달 안에 Android와 Windows를 추가할 예정입니다. 무엇을 만들고 계신지 알려주시면 출시될 때 알려드리겠습니다.