iPhone as a second GPU speeds up Qwen 3.8 27B prefill by 29–44%
10/03/2026 — 10/03, 03:11·1 sources·1 reports
Story overview
On October 3, 2026, a developer posted on Reddit's LocalLLaMA community about using an iPhone 17 Pro Max as a second GPU alongside a 24GB M4 Pro MacBook. The phone carries part of the context window for Qwen 3.8 27B at IQ4_XS quantization, and according to the developer, end-to-end prefill speed improved by 29–44%.
The backdrop, as described in the post, is memory pressure on the MacBook. Every file or tool result the developer's agent reads involves a wait, and even with the wired limit raised to 20480, only 64k of 8-bit context fits next to Qwen 3.8 27B (IQ4_XS).
The post carries a disclaimer: the prefill tokens per second shown on the phone is computed only for the layers that device holds, so it does not represent overall throughput. The accurate figure, the developer says, is the end-to-end prefill rate, and the on-device display is already being fixed to report it. The numbers given in the post, the developer stresses, are accurate for end-to-end prefill rate. As of this report, the work stands at the point where the developer has published the figures and flagged that the display metric still needs correcting.
AI-generated from 1 reports · updated 2 hours ago
Latest turnA developer turned an iPhone 17 Pro Max into a second GPU for a 24 GB M4 Pro MacBook, offloading part of the context window for Qwen 3.8 27B (IQ4_XS). End-to-end prefill runs 29-44% faster; the prefill t/s shown on the phone only counts the layers it holds and is being fixed to report the end-to-end rate.

- Heat index
- 89
- Sources
- 1
- Reports
- 1
- First seen
- 4 hours ago
Reports on this story headlines open the original
A developer turned an iPhone 17 Pro Max into a second GPU for a 24 GB M4 Pro MacBook, offloading part of the context window for Qwen 3.8 27B (IQ4_XS). End-to-end prefill runs 29-44% faster; the prefill t/s shown on the phone only counts the layers it holds and is being fixed to report the end-to-end rate.
Other stories people are talking about
- 593SurgeNVIDIA launches 64GB DGX Spark desktop AI computer at $4,9997 sources
- 466NewApple tightens macOS Full Disk Access over AI agent risks5 sources
- 372Google releases Gemini 4 Argon as next-gen frontier AI model10 sources
- 262SurgeOpenAI DevDay 2026 launches Dots, Decisions API, and more3 sources
- 248SurgeSuno launches voice generation feature for narration and music3 sources
- 212SurgeMicrosoft AI releases MAI-Transcribe-2-Streaming real-time transcription model3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
