Models
Small, specialized models that run on the device, one per task.
Available
Accurate word timestamps for any transcript.
Studio sound, no cloud bill.

Create video shorts and highlights.

Detect language based on 30 seconds of audio.
Suggest emoji faster than you can type.

Generate topics and tags for posts and articles.
Filter PII on-device.

Rough sketch. Perfect shape.

Suggest a title and description for any text.
Detect language based on 3 words.

Detect and remove filler words in seconds.

Transcribe 10 minutes of audio in 2s on an iPhone.

Beta
Tag and rank images based on aesthetics and content.
Find people in photos, private on the device.
Flag nudity before upload or showing.

Extract typed JSON from any text.
Catch hate speech before it posts.

Label who said what in audio and video.