Details of Backburner software that increased AI prefill speeds by 44 percent using iPhone 17 Pro Max and M4 Pro MacBook Pro.
Running advanced AI models like the Qwen3.8-27B locally can be quite challenging depending on the amount of system memory and VRAM available. The 24GB of combined memory found in M4 Pro MacBook Pro devices can create a significant bottleneck on the prefill speeds of such models.
A Reddit user offered a different solution to this problem by connecting an iPhone 17 Pro Max to a MacBook Pro via the USB-C port. Thanks to a specially developed software, an improvement of up to 44% in the prefill performance of the AI model was achieved.
Workload Sharing Between iPhone and MacBook
The developer known as u/StayLameBro followed a formula that shares the workload between the two devices to overcome the MacBook’s 24GB combined memory limit. For the prefill acceleration process, the MacBook Pro processes the first 40 layers of each 256-token cluster and transfers this information to the iPhone.
I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. by
u/StayLameBro in
LocalLLaMA
The iPhone 17 Pro Max, on the other hand, uses its own A19 Pro chip’s GPU to process layers 41 through 64. This process, by leveraging the device’s GPU capacity, makes the processes approximately 2.4 times faster than working alone.
The device’s Neural Engine unit is also active in this process, compiling old context data and shortening the process time. Thanks to this, the time to write tokens with a length of 140,000 contexts can be reduced from 279 milliseconds to 176 milliseconds.
Performance Gains and System Limitations
In tests, significant speed increases were recorded in different context windows during the preloading of a 2,000-token document. In 8K context, the MacBook alone offered a speed of 132 ticks/second, while with the iPhone this value increased to 177 ticks/second, providing a 35% increase.
In 16K context, performance increased from 109 ticks/second to 157 ticks/second, showing a 44% improvement. In the 32K context window, the speed increased from 101 kT/second to 130 kT/second, representing a 29% improvement.
However, this method doesn’t always yield miraculous results and encounters some technical limitations. For example, the iPhone has no accelerating effect on text rendering speeds below 64K, and this process is entirely the responsibility of the MacBook.
It is predicted that future iPhone 18 Pro models with the A20 Pro chip will offer higher efficiency thanks to their dual 16-core Neural Engine architecture. This hardware can provide a higher throughput in reasonable AI tasks compared to the current 7-core GPU.
This software, shared as open source on GitHub under the name backburner, offers an experimental area for users who want to use hardware resources more efficiently. Such experimental studies reveal the potential of portable devices to improve AI performance.
Do you think smartphones could become standard auxiliary hardware in the future to enhance the AI performance of computers?