What It Takes to Run Machine Learning on the Device Itself
A model that depends on a strong internet connection is of limited use in many of the environments we build for, so the requirement that it run on the device itself tends to come before almost every other design decision. That rules out a server round trip and any dependency on a signal that may not be available at the moment the tool is needed.
The constraint is harder to meet than it sounds. A model tuned for maximum accuracy is usually too large and too slow for a mid-range phone or a basic tablet to handle, and the techniques that make it smaller, quantisation and pruning among them, reduce the size significantly but can cost accuracy along the way. There is no version of this trade-off that comes free, and designing as though there were tends to produce unpleasant surprises in the field.
What the trade-off buys is worth the cost. A model running on the device returns an answer in a fraction of a second with nothing to wait on, which means a health worker examining a patient or a farmer examining a damaged crop gets a result immediately, whether they are standing somewhere with strong signal or nowhere at all.
Running on-device also means giving up the safety net a cloud model has, since there is no larger system sitting behind it to fall back on when something confusing arrives. The model has to recognise its own uncertainty and act on it, rather than producing an answer regardless of how shaky the ground underneath that answer is. A confidence threshold decides whether a result is presented as settled or flagged for a person to check, and that threshold is set against real test cases rather than estimated.
Testing carries more weight here for a specific reason. A cloud model can be observed in real time from a central server, while a model distributed across thousands of separate devices cannot be watched that closely once it has left. Most of the safety work therefore happens before release, run against a fixed set of difficult and rare cases every time a new version ships.
The most technically impressive part of an AI system is rarely what determines whether the tool works in practice. That comes down to something quieter, which is whether it still does the job for the person relying on it, in whatever conditions that person happens to be working in.