Desert Ant Labs

The cerebellum for every product.

Desert Ant Labs is an on-device AI lab. We build the fastest models for your task, designed to run on the billions of devices people already own. Lightning fast, private, and no per-call cost. Little brains in every product.

The intelligence layer for apps

The cerebellum is the little brain, located near our neck and connected to our brain stem. A tenth of the brain by size, four fifths of its neurons. The cerebellum does the fast actions and the control functions: balance, timing, coordination, the skills you learned once and never think about. Nothing in it reasons. Its value is speed and reliability, so the rest of the brain is free to think.

Software gets better with the same layer. The big reasoning models run in the cloud, and the constant work can run there too: cleaning up audio, reading a photo, pulling data out of a message. We start with fast, specialized local models: built for one job, they compete with cloud inference on price, speed, and performance. Combine them with the cloud and you build better products, private by default and free to run. Use the right model for each job, and send work to the cloud when needed. That layer is what we build.

There's more compute in people's pockets than in the data centers

The world ships more than a billion capable phones, tablets, and laptops a year, most with a chip built for exactly this work, paid for and idle most of the day. Run the model there and the economics flip: no per-call cost, no round-trip, and nothing leaves the device.

What we build

Not the biggest model, the fastest one for the job. Many small models, each tuned for one task and speed with no per-call cost. Each comes with a native SDK for Swift, Kotlin, and JavaScript that does the heavy lifting, so a model drops into your product in a few lines of code. You build something unique faster, and what used to be a whole product to buy becomes a building block.

Specialized
Models specialized for their task, optimized for phones and desktop compute.
Lightning fast
Fast enough for every interaction, every keystroke and every frame, with no per-call cost.
Clean data
Trained on licensed and openly available data, so every model can be used commercially, without exposure.

We start from the product

We lead with the product and design our models for the best end-user experience. A model is a thousand product decisions: what it learns from, the shape of its inputs and outputs, where it trades accuracy for speed. We make those calls the way we design anything people touch, from years of shipping consumer apps.

Over the next decade more of the interface gets built by more people, and the model stops being a detail behind the product. The models with product taste, ready to drop in, are the ones those products will run on.

Public, and free to start

Every model and its benchmarks are public on Hugging Face, the SDK on GitHub. Free up to 100k monthly active devices per SDK. Building something with them, or want to build them with us? Write to us.