Who.
On-device speaker labeling that turns a recording into per-person turns with face boxes and timestamps, ready to cut or caption.
Label who said what in audio and video.
Record a two-hour two-camera interview and Who splits it by speaker, so you jump to what one person said or cut on every speaker change.
Who listens and watches the mouths at the same time, so overlaps and short interjections do not throw the labels off. Each turn carries a speaker id, a face-track id, a normalized face bounding box and millisecond timestamps, in the same shape a Whisper transcript uses.
Follow-the-speaker editing
Cut or switch camera on every speaker change. Use the face box to frame the shot in talking-head and multi-cam video.
Captions with names on them
Attach speaker labels to Whisper-compatible transcript turns. A viewer reads who said each line, not one anonymous block of text.
Find what one person said
Split a two-hour recording into per-person turns on the device. Filter by speaker and jump straight to the turn instead of scrubbing the timeline.
What the model does
- Per-turn: speaker id, face track id, normalized face bbox, millisecond timestamps.
- Whisper-compatible turns for captioning pipelines.
- Uses the picture as well as the audio, not audio alone.
Specs
- Platforms
- Apple (Swift) first
FAQ
What is Who?
On-device speaker labeling that turns a recording into per-person turns with face boxes and timestamps, ready to cut or caption.
Does Who run on device?
Yes. Who runs on the device, with no server call, so the data stays with the user.
Is Who available yet?
Who is in closed beta. Request early access on this page.
How much does Who cost?
Every model is free up to 100k monthly active devices per SDK. Unlimited inference per user. Contact us for custom licenses.
Early access
Tell us what you are building and we will get you set up.