Running a 125B-parameter model on a single RTX 4090 sounds impossible, because the weights are far larger than the card's memory. The practical answer is a mix of quantization, offload and careful tuning of the inference setup.
An RTX 4090 has 24 GB of VRAM. A 125B model in FP16 needs roughly ten times that amount just for the weights, before you count the KV cache, activations and runtime overhead.
So "running a 125B model on a 4090" always means the same thing: the model is compressed and then split between fast GPU memory and slower system RAM. There is no configuration where all of it sits in VRAM at full precision.
If you are new to local inference, start with smaller models first. The workflow is identical, only the numbers change. A good entry point is this beginner's guide to running AI models locally on your PC.
Quantization stores weights in fewer bits per value. FP16 uses 16 bits; a 4-bit format uses four.
For a 125B model, 4-bit is the usual compromise. It brings the weights down to a size measured in tens of gigabytes instead of hundreds. That is still far more than 24 GB, which is exactly why offload exists.
Offload means running part of the model on the GPU and the rest on the CPU, with the remaining weights held in system RAM.
A typical setup keeps as many layers in VRAM as will fit, and sends the rest to the CPU. Every generated token then passes through both parts of the model, so the slowest part sets the pace.
The consequences are predictable:
Speed on this setup is not a fixed number. It depends on how much of the model stays on the GPU and how long your context is.
Four factors dominate: the quantization level, the number of layers kept in VRAM, the size of the KV cache, and the context length you ask for. Shorter context and fewer offloaded layers push throughput up; bigger models and longer prompts push it down.
Treat 100 tokens per second as a target you approach by keeping the GPU share high, not as a guarantee of any specific configuration.
Local deployment at this scale is only realistic with open-weight models, where you can download the weights and quantize them yourself. Closed models that you can only call through an API cannot be offloaded or compressed on your own machine.
It also helps to compare models by real usage before you commit hours to a download and a slow first run. Leaderboards and side-by-side arenas are a quick way to shortlist candidates — see how to read the LMArena leaderboard and pick a model. For a broader overview of what is available with open weights, read open-weight AI models and how to run them for free.
Does a 125B model fit on a single RTX 4090?
No. Not in full precision and not after 4-bit quantization either — 24 GB of VRAM is simply too small for weights of this size, so offload is required.
What is quantization in simple terms?
It is storing model weights in fewer bits, such as 8-bit or 4-bit instead of 16-bit. It reduces memory use several times over at the cost of a small quality loss.
What is offload and why do I need it?
Offload keeps part of the model on the GPU and the rest in system RAM. You need it because a 125B model cannot fit entirely into VRAM.
Can I really reach 100 tokens per second?
It depends on how much of the model stays on the GPU, your quantization level and your context length. Treat 100 tokens per second as a target, and expect slower output as the offloaded share grows.