On-device inference is not automatically better than cloud inference. This research note looks at the product constraints that make one architecture more sensible than the other.
The decision is multi-dimensional
- Latency and offline requirements
- Privacy expectations
- Model size and device capability
- Update frequency
- Server and inference cost