HaiAI123

Curated Global AI Tools Directory

NewHot story
89
heat index
New

iPhone as a second GPU speeds up Qwen 3.8 27B prefill by 29–44%

10/03/2026 — 10/03, 03:11·1 sources·1 reports

Story overview

On October 3, 2026, a developer posted on Reddit's LocalLLaMA community about using an iPhone 17 Pro Max as a second GPU alongside a 24GB M4 Pro MacBook. The phone carries part of the context window for Qwen 3.8 27B at IQ4_XS quantization, and according to the developer, end-to-end prefill speed improved by 29–44%.

The backdrop, as described in the post, is memory pressure on the MacBook. Every file or tool result the developer's agent reads involves a wait, and even with the wired limit raised to 20480, only 64k of 8-bit context fits next to Qwen 3.8 27B (IQ4_XS).

The post carries a disclaimer: the prefill tokens per second shown on the phone is computed only for the layers that device holds, so it does not represent overall throughput. The accurate figure, the developer says, is the end-to-end prefill rate, and the on-device display is already being fixed to report it. The numbers given in the post, the developer stresses, are accurate for end-to-end prefill rate. As of this report, the work stands at the point where the developer has published the figures and flagged that the display metric still needs correcting.

AI-generated from 1 reports · updated 2 hours ago

Latest turnA developer turned an iPhone 17 Pro Max into a second GPU for a 24 GB M4 Pro MacBook, offloading part of the context window for Qwen 3.8 27B (IQ4_XS). End-to-end prefill runs 29-44% faster; the prefill t/s shown on the phone only counts the layers it holds and is being fixed to report the end-to-end rate.

Related tools
24-hour heatpeak 100 · 4h ago
24 hours agonow

Reports on this story headlines open the original

Today
  1. A developer turned an iPhone 17 Pro Max into a second GPU for a 24 GB M4 Pro MacBook, offloading part of the context window for Qwen 3.8 27B (IQ4_XS). End-to-end prefill runs 29-44% faster; the prefill t/s shown on the phone only counts the layers it holds and is being fixed to report the end-to-end rate.

    Reddit · LocalLLaMAAI score 72

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to the hot board →