Why fine-tuning a 7B model needs 112 GB of VRAM
10/03/2026 — 10/03, 12:12·1 sources·1 reports
Story overview
A post published on DEV Community on October 3, 2026, works through the memory math of fine-tuning a 7B model. The article says that when people are asked how much VRAM this takes, the instinct is to answer that the model itself is 14 GB in fp16, so a little more than that should be enough. The real figure it gives is about 112 GB. Within that total, the model weights account for just 14 GB, while roughly 98 GB goes to gradients, optimizer states and activations. The post attributes this accounting to the ZeRO paper. Its argument is that once you see where those other 98 GB go, LoRA and QLoRA stop looking like clever tricks and start looking like the obvious choice. The excerpt opens by describing the 112 GB as the number you face before storing a single activation, then lists activations among the components of the remaining 98 GB, a small tension in how the breakdown is framed. The piece presents itself as an accounting exercise rather than a benchmark: the headline number is not the size of the model. So far no other report has offered different figures or added anything to the picture.
AI-generated from 1 reports · updated 1 hour ago
Latest turnFine-tuning a 7B model in fp16 actually needs around 112 GB of memory, and the weights account for only 14 GB of it, according to the memory accounting from the ZeRO paper. The remaining ~98 GB goes to gradients, optimizer states and activations. Once you see where the memory goes, LoRA and QLoRA stop looking like tricks and start looking obvious.

Reports on this story headlines open the original
Fine-tuning a 7B model in fp16 actually needs around 112 GB of memory, and the weights account for only 14 GB of it, according to the memory accounting from the ZeRO paper. The remaining ~98 GB goes to gradients, optimizer states and activations. Once you see where the memory goes, LoRA and QLoRA stop looking like tricks and start looking obvious.
DEV Community · AIAI score 82
Other stories people are talking about
- 554NVIDIA launches 64GB DGX Spark desktop AI computer at $4,9998 sources
- 539RisingApple tightens macOS Full Disk Access over AI agent risks7 sources
- 400Meta open-sources Muse Gadgets firmware and SDK5 sources
- 234SurgeAmazon Weighs Moving $8 Billion of Nvidia Chips into a Financing Vehicle3 sources
- 227Anthropic launches Claude Frontier Academy with $100M3 sources
- 226SurgeHugging Face open-sources AstaBrief for fast report generation3 sources
How is heat calculated?About the methodHide
Heat counts how many independent sources covered a story in the last 48 hours: one source counts once no matter how many posts it published, decaying with a 24-hour half-life. What ranks first is what many people are talking about.
This page aggregates public feeds. Headlines and summaries are machine-organized and remain the property of the original authors; verify important facts at the source.
- Surge
- Discussion rising fast
- New
- First report within 6 hours
- Rising
- Still gathering discussion
