Stacked M3 Ultra Mac Studios run trillion-parameter model in benchmark test

- Developer Adea0x posted on X on July 7 that stacked Apple Silicon Mac Studios with M3 Ultra ran a trillion-parameter model locally. - The benchmark’s key figure was about 28.6 tokens per second, above Apple’s own claim that one M3 Ultra can run 600B-parameter models. - Apple’s WWDC26 MLX sessions and LM Studio’s Apple-focused tooling remain the clearest public places to track follow-up demonstrations.

A July 7 post on X said a cluster of Apple Mac Studios with M3 Ultra chips ran a trillion-parameter model locally at about 28.6 tokens per second. The claim circulated as a fresh data point in a niche but fast-growing corner of AI computing: using Apple’s unified-memory desktops, linked together, to run models that would normally be associated with server-class GPU systems. Apple has said a single Mac Studio with M3 Ultra can run large language models with more than 600 billion parameters directly on device, and the company has been promoting MLX and distributed inference tools for local AI workflows. The benchmark did not come from Apple. It appeared in a social post cited in the source briefings, and similar public write-ups have described four-machine Mac Studio clusters using M3 Ultra systems, Thunderbolt 5 links and Apple’s MLX software stack to reach trillion-parameter inference. LM Studio materials and third-party reports tied to WWDC26 have also described a four–Mac Studio setup running Moonshot AI’s Kimi K2.6, a 1-trillion-parameter mixture-of-experts model. (apple.com) ### How can a Mac Studio run a model that large? Apple said on March 5, 2025 that M3 Ultra supports more than half a terabyte of unified memory and delivers over 800 GB/s of memory bandwidth. Apple also said the chip was built for AI workloads and that a single M3 Ultra Mac Studio could run LLMs with more than 600 billion parameters on device. (news.aibase.com) A trillion-parameter model pushes beyond that single-box claim, which is why the reports center on multiple machines. Third-party descriptions of the setup say four Mac Studios were linked into one inference cluster, pooling about 1.5 TB of unified memory across systems. Creative Strategies said its own four-machine M3 Ultra cluster used two 256 GB systems and two 512 GB systems. (apple.com) ### Where does the 28.6 tokens-per-second figure fit? The 28.6 tokens-per-second number is notable because it is in the same range as other public claims for clustered Mac Studio inference on trillion-parameter-class models. A June 22 report describing the Apple-LM Studio demonstration said generation speed could reach about 28 tokens per second on a four–Mac Studio cluster. Max Weinbach of Creative Strategies wrote in December 2025 that his four-unit setup ran a 1T-parameter Kimi K2 Thinking model at about 25 tokens per second. (creativestrategies.com) Those figures are not directly comparable without identical models, quantization settings and prompt lengths. But they point in the same direction: the cluster is being judged less by peak training-style performance than by whether it can sustain usable local inference on very large models. ### What made clustered Macs faster than earlier attempts? (news.aibase.com) Thunderbolt 5 and RDMA are central to the newer cluster claims. Apple said M3 Ultra adds Thunderbolt 5, and exo, an open-source clustering project, says it supports RDMA over Thunderbolt 5 to cut latency between devices. Creative Strategies wrote that Apple enabled RDMA access over Thunderbolt 5 in macOS 26.2 and said that change lifted Kimi K2 throughput from about 5 tokens per second on a standard connection to about 25 tokens per second. (creativestrategies.com) Apple’s WWDC26 session on MLX also highlighted distributed inference and named LM Studio among the tools in the local AI stack on Mac. That does not verify any one benchmark result, but it does show Apple is publicly supporting the software path used in these demonstrations. ### Is this meant for ordinary buyers? The price point in the public examples says no. (apple.com) Creative Strategies said its four-machine configuration would cost about $39,596 before tax, and Weinbach wrote that buying such a setup solely for personal inference “does not make sense” when cloud inference is cheaper per token. He said the case was stronger for production houses or engineering teams that already use Mac Studios and want private, local inference. (developer.apple.com) LM Studio’s public materials make a similar pitch around local control. Its site says users can run models privately on their own computers, and Apple’s Mac Studio marketing has paired LLM throughput claims with privacy language around on-device AI. ### What should readers watch next? WWDC26 materials remain the most concrete public trail for this line of work. Apple’s June 8-12 developer conference schedule and MLX sessions point to distributed inference as an active area, while LM Studio’s Apple-focused tooling and LM Link features provide a public software layer for future demos. (creativestrategies.com) Any next benchmark that names the exact model, quantization level, number of Mac Studios and prompt settings will be easier to compare with the July 7 post. (lmstudio.ai) (developer.apple.com)

Get your own daily briefing

Scout delivers personalized news, insights, and conversations tailored to your role and industry.

Download on the App Store

Shared from Scout - Be the smartest in the room.