HaiAI123

Curated Global AI Tools Directory

SurgeIndustry update
↑ 159%

Strata Runs 100B-Parameter Models on 8GB GPUs

10/09/2026 — 10/10, 18:20·2 sources (all)·3 reports (all)·In the 2026-10-10 briefing

Key takeaways

  • A hands-on test by Appinn shows Strata can run 100B-parameter models on ordinary PCs with only 8GB of VRAM, verified on the RTX 5070, RTX 30…

Background & analysis

On October 8, 2026, an author named Wu Jiahao published a technical article describing a scheme called Strata that runs a 125B-parameter MoE large model on a single RTX 5090 with 32GB of VRAM, tested with driver version 617. The article framed the approach as bringing server-class MoE inference to ordinary PCs. The post appeared on Juejin on October 9 at 22:51.

Early the next day, on October 10 at 01:12, the same author published a follow-up tutorial on Juejin covering the full workflow for running the 125B model on a single RTX 5090 with 32GB of VRAM. It walked through installation, verification, and benchmarking, and listed a more specific test environment that included driver version 617.14.

Later that day, at 16:44 on October 10, the site Xiaozhong Software reported that it had tested Strata itself and found the tool can run a trillion-parameter-class large model on an ordinary computer with only 8GB of VRAM. The report said the approach was verified on a range of GPUs, including the RTX 5070, RTX 3060, RTX 4060 Ti, RX 9070 XT, RX 9060 XT, RX 7800 XT, RTX 3090, and even older Quadro cards. According to that report, this significantly lowers the GPU barrier for running very large models locally. The public information currently stops there: the reports establish that Strata is feasible, but they do not name the developer behind the scheme, describe its implementation in detail, or mention any next steps.

AI-generated from 3 reports · updated 3 hours ago

Latest turnStrata, tested hands-on by the Chinese site Xiaozhuanruanjian, can run 100B-parameter LLMs on ordinary PCs with only 8GB of VRAM. The outlet says it verified the setup on GPUs including the RTX 5070, RTX 3060, RTX 4060 Ti, RTX 3090, RX 9070 XT and older Quadro cards, lowering the hardware bar for local large-model use.

24-hour heatHeat index 145 · peak 163 · 4h ago
24 hours agonow

Reports on this story headlines open the original

Today
  1. Strata, tested hands-on by the Chinese site Xiaozhuanruanjian, can run 100B-parameter LLMs on ordinary PCs with only 8GB of VRAM. The outlet says it verified the setup on GPUs including the RTX 5070, RTX 3060, RTX 4060 Ti, RTX 3090, RX 9070 XT and older Quadro cards, lowering the hardware bar for local large-model use.

    小众软件AI score 64
  2. A hands-on tutorial walks through running a 125B large model on a single RTX 5090 with 32GB of VRAM, covering installation, verification and benchmarking. The author lists the exact test environment, including driver version 617.14.

    掘金AI score 65
Yesterday
  1. A technical write-up describes how Strata runs a 125B-parameter MoE model on a single RTX 5090 with 32GB of VRAM, tested on driver version 617. The author, Wu Jiahao, frames it as bringing server-grade MoE inference to an ordinary PC.

    掘金AI score 61

Other stories people are talking about

How is heat calculated?About the method

Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.

This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source. If you believe a headline or summary infringes your rights, email the contact address in the footer with the page URL and basis of your claim, and we will remove or replace it after review.

Surge
Discussion rising fast
New
First report within 6 hours
Rising
Still gathering discussion

Back to AI industry updates →