# How to Run a 125B-Parameter LLM on an RTX 4090: Quantization, Offload and 100 Tokens/s

> AI Release · @ai_release1 · https://ai-release.net/guides/kak-zapustit-model-na-125b-parametrov-na-rtx-4090-kvant_en.html

Running a 125B-parameter model on a single RTX 4090 sounds impossible, because the weights are far larger than the card's memory. The practical answer is a mix of quantization, offload and careful tuning of the inference setup.

## TL;DR

- A 125B model never fits on one consumer GPU in full precision, so quantization is the first step.
- 4-bit quantization shrinks the weights roughly four times compared with FP16.
- Offload moves part of the model into system RAM and keeps the rest on the GPU.
- 100 tokens per second is a target, not a promise: speed depends on how much of the model stays in VRAM.
- The more layers you offload, the more memory you save and the slower generation becomes.

## Why a 125B model does not fit on an RTX 4090

An RTX 4090 has 24 GB of VRAM. A 125B model in FP16 needs roughly ten times that amount just for the weights, before you count the KV cache, activations and runtime overhead.

So "running a 125B model on a 4090" always means the same thing: the model is compressed and then split between fast GPU memory and slower system RAM. There is no configuration where all of it sits in VRAM at full precision.

If you are new to local inference, start with smaller models first. The workflow is identical, only the numbers change. A good entry point is [this beginner's guide to running AI models locally on your PC](https://ai-release.net/guides/kak-zapustit-ii-model-lokalno-na-svoem-kompyutere-gayd.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

## Quantization: the main lever

Quantization stores weights in fewer bits per value. FP16 uses 16 bits; a 4-bit format uses four.

- 8-bit: about half the memory of FP16, quality almost unchanged.
- 4-bit: about a quarter of FP16 memory, with a small quality loss on most tasks.
- Below 4-bit: saves more memory, but quality drops faster and speed does not always improve.

For a 125B model, 4-bit is the usual compromise. It brings the weights down to a size measured in tens of gigabytes instead of hundreds. That is still far more than 24 GB, which is exactly why offload exists.

## Offload: how it works in practice

Offload means running part of the model on the GPU and the rest on the CPU, with the remaining weights held in system RAM.

A typical setup keeps as many layers in VRAM as will fit, and sends the rest to the CPU. Every generated token then passes through both parts of the model, so the slowest part sets the pace.

The consequences are predictable:

- More offload means more system RAM is required.
- CPU speed starts to matter as much as GPU speed.
- Tokens per second drop as the offloaded share grows.

## How to get close to 100 tokens per second

Speed on this setup is not a fixed number. It depends on how much of the model stays on the GPU and how long your context is.

Four factors dominate: the quantization level, the number of layers kept in VRAM, the size of the KV cache, and the context length you ask for. Shorter context and fewer offloaded layers push throughput up; bigger models and longer prompts push it down.

Treat 100 tokens per second as a target you approach by keeping the GPU share high, not as a guarantee of any specific configuration.

## Checklist before you start

- [ ] Confirm your VRAM budget (24 GB on an RTX 4090).
- [ ] Measure free system RAM, since offloaded weights live there.
- [ ] Pick a quantized version of the model instead of the original weights.
- [ ] Set the context length as low as your task allows.
- [ ] Test with a short prompt before running long generations.
- [ ] Measure tokens per second, not just "did it load".

## Choosing a model you can actually run

Local deployment at this scale is only realistic with open-weight models, where you can download the weights and quantize them yourself. Closed models that you can only call through an API cannot be offloaded or compressed on your own machine.

It also helps to compare models by real usage before you commit hours to a download and a slow first run. Leaderboards and side-by-side arenas are a quick way to shortlist candidates — see [how to read the LMArena leaderboard and pick a model](https://ai-release.net/guides/kak-polzovatsya-lmarena-arena-ai-chitaem-reyting-i-vybi.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide). For a broader overview of what is available with open weights, read [open-weight AI models and how to run them for free](https://ai-release.net/guides/kakie-ii-modeli-imeyut-otkrytye-vesa-i-kak-ih-zapustit.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

## FAQ

**Does a 125B model fit on a single RTX 4090?**
No. Not in full precision and not after 4-bit quantization either — 24 GB of VRAM is simply too small for weights of this size, so offload is required.

**What is quantization in simple terms?**
It is storing model weights in fewer bits, such as 8-bit or 4-bit instead of 16-bit. It reduces memory use several times over at the cost of a small quality loss.

**What is offload and why do I need it?**
Offload keeps part of the model on the GPU and the rest in system RAM. You need it because a 125B model cannot fit entirely into VRAM.

**Can I really reach 100 tokens per second?**
It depends on how much of the model stays on the GPU, your quantization level and your context length. Treat 100 tokens per second as a target, and expect slower output as the offloaded share grows.
