Work out edge device compute sizing instantly with clear inputs, formula shown and shareable results.
Required compute is operations per inference times inference rate, divided by achievable utilisation — and that last term is the one people forget. Real accelerators sustain only 20-50% of peak on typical models because of memory bandwidth and layer shapes, so a 0.036 TOPS raw demand needs a 0.1 TOPS part.
Edge compute
raw TOPS = GFLOP per inference x inferences per second / 1000; required = raw / utilisation
Peak TOPS assumes perfectly shaped dense matrix operations. Depthwise convolutions, small batches and activation memory traffic all leave the multipliers idle.
Substantially. Int8 inference typically runs 2-4 times faster than float and uses a quarter of the memory bandwidth, often with negligible accuracy loss.