Benchmarks show M5 Max delivers strong on-device inference for optimized models
- Recent Apple Silicon benchmarks posted yesterday showed an M5 Max running optimized models yielding 27–83 t/s across varying context lengths on inference tasks. - An online post noted M5 Max throughput of 27 t/s at small context and 83 t/s at large context across optimized models. - The benchmark post was shared July 7–8 on X by users discussing M1/M5 suitability for ML inference. (x.com)