# LoRA fine-tuning on a single GPU: a step-by-step guide

> AI Release · @ai_release1 · https://ai-release.net/guides/lora-fayntyun-na-odnoy-videokarte_en.html

LoRA fine-tuning on a single GPU lets you adapt a small language model (SLM — a model from a few million up to 10 billion parameters that runs on consumer hardware) to your own task without retraining the whole model. This guide is for beginners with one consumer-grade GPU (8–24 GB VRAM) and a Linux or Windows machine with Python installed. It works best with compact models like Qwen3.5-0.8B, the Gemma family (2–9B), Gemini Nano-1 (1.8B) and Nano-2 (3.25B), or Phi-4 14B from Microsoft. If you haven't run a model locally yet, start with [How to run AI models locally on your PC: a beginner's guide](https://ai-release.net/guides/kak-zapustit-ii-model-lokalno-na-svoem-kompyutere-gayd.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

## Требования и подготовка
![Illustration: Требования и подготовка](https://ai-release.net/guides/img/lora-fayntyun-na-odnoy-videokarte-en-1.jpg)

- One consumer GPU with at least 8 GB VRAM (LoRA trains only small adapter weights, so small models fit easily; Phi-4 14B needs more memory).
- Python 3.10+ installed, plus the `transformers`, `peft` (LoRA implementation library) and `datasets` packages.
- A base SLM: Qwen3.5-0.8B, Gemma 2–9B, Gemini Nano-1 (1.8B) / Nano-2 (3.25B), or Phi-4 14B.
- A small dataset in JSONL format (question/answer pairs) — a few hundred examples are enough for a narrow task.
- Admin rights to install CUDA drivers, and ~20 GB free disk space for checkpoints.

**Пример:** install the core packages:

```bash
pip install transformers peft datasets
```

This works because `peft` adds LoRA adapters on top of a frozen base model, so only a tiny fraction of weights is trained.

## Пошаговая инструкция
![Illustration: Пошаговая инструкция](https://ai-release.net/guides/img/lora-fayntyun-na-odnoy-videokarte-en-2.jpg)

1. **Pick a model that fits your GPU.** For an 8 GB card start with Qwen3.5-0.8B or Gemma 2B; for 16–24 GB use Gemma 9B or Phi-4 14B. Expected result: you know the exact model name for the next steps. This works because smaller models leave room for gradients and optimizer states.

2. **Prepare your dataset as JSONL.** One JSON object per line with an instruction and an output. Expected result: a file like `train.jsonl` that loads without errors.

**Пример:**

```json
{"instruction": "Classify the ticket", "output": "billing"}
{"instruction": "Classify the ticket", "output": "technical"}
```

This works because the trainer expects plain text pairs, not a database.

3. **Load the base model and tokenizer.** Use `AutoModelForCausalLM` and `AutoTokenizer` from `transformers`. Expected result: the model loads to GPU without an out-of-memory error.

**Пример:**

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-0.8B")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-0.8B")
```

This works because the model is downloaded once and cached locally.

4. **Attach a LoRA adapter via `peft`.** Set a small rank (r=8 or 16) and target the attention layers. Expected result: `model.print_trainable_parameters()` shows something like "0.5% of parameters trainable". This works because the base weights stay frozen and only adapter matrices update — that's what makes single-GPU training possible.

**Пример:**

```python
from peft import LoraConfig, get_peft_model
cfg = LoraConfig(r=8, lora_alpha=16, target_modules=["q_proj", "v_proj"])
model = get_peft_model(model, cfg)
model.print_trainable_parameters()
```

5. **Run training for 1–3 epochs on your dataset.** Use a small batch size (1–4) with gradient accumulation. Expected result: the loss decreases over steps, e.g. from ~2.5 to ~1.0. This works because a narrow task needs few updates, not many epochs.

6. **Save the adapter and test it.** Call `model.save_pretrained("my-lora-adapter")`, then load the base model plus adapter and ask it a question from your domain. Expected result: answers follow your dataset's style. This works because the adapter stores only the learned delta, keeping the file small (tens of MB instead of GB).

Before shipping your fine-tuned model anywhere, check quality properly — see [How to Test and Evaluate AI Models: A QA Guide for Machine Learning](https://ai-release.net/guides/kak-testirovat-i-otsenivat-kachestvo-ii-modeley-gayd-po.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

**Checklist before you start training:**

- [ ] Model chosen and fits in VRAM with headroom
- [ ] `train.jsonl` has at least 100–300 clean examples
- [ ] `peft` and `transformers` installed
- [ ] Output directory for the adapter created
- [ ] Test questions written down in advance

## Возможные проблемы и решения
![Illustration: Возможные проблемы и решения](https://ai-release.net/guides/img/lora-fayntyun-na-odnoy-videokarte-en-3.jpg)

- **Out of memory (error like "CUDA out of memory").** Symptom: training crashes on the first steps. Fix: switch to a smaller model (e.g. from Gemma 9B to Gemma 2B or Qwen3.5-0.8B), reduce batch size to 1, or lower LoRA rank from 16 to 8.
- **Loss doesn't decrease.** Symptom: loss stays flat after hundreds of steps. Fix: check the dataset format (each line must be valid JSON with instruction/output), increase learning rate slightly, and verify the adapter actually targets the right modules — `print_trainable_parameters()` should show a non-zero trainable count.
- **Model outputs gibberish after training.** Symptom: fine-tuned model produces broken text. Fix: you likely trained too many epochs on too little data — retrain for 1–2 epochs instead, and re-check that the tokenizer used at inference matches the one used in training.

## FAQ

**Can I fine-tune Phi-4 14B on one GPU?**
Yes, but it needs a high-VRAM consumer card; for smaller GPUs start with Qwen3.5-0.8B or Gemma 2B.

**Is a fine-tuned SLM a replacement for a large LLM?**
No. SLMs have their own niche — narrow tasks and agentic systems — not a general replacement for LLMs.

**How much data do I need for LoRA?**
For a narrow task, a few hundred clean examples in JSONL format are usually enough to see a style/format change.

**Which small model should a beginner pick?**
Start with Qwen3.5-0.8B or a small Gemma (2B) — they run on consumer hardware and leave VRAM headroom for training.
