LoRA fine-tuning on a single GPU lets you adapt a small language model (SLM — a model from a few million up to 10 billion parameters that runs on consumer hardware) to your own task without retraining the whole model. This guide is for beginners with one consumer-grade GPU (8–24 GB VRAM) and a Linux or Windows machine with Python installed. It works best with compact models like Qwen3.5-0.8B, the Gemma family (2–9B), Gemini Nano-1 (1.8B) and Nano-2 (3.25B), or Phi-4 14B from Microsoft. If you haven't run a model locally yet, start with How to run AI models locally on your PC: a beginner's guide.

Пример:install the core packages:
pip install transformers peft datasets
This works because `peft` adds LoRA adapters on top of a frozen base model, so only a tiny fraction of weights is trained.

1. **Pick a model that fits your GPU.** For an 8 GB card start with Qwen3.5-0.8B or Gemma 2B; for 16–24 GB use Gemma 9B or Phi-4 14B. Expected result: you know the exact model name for the next steps. This works because smaller models leave room for gradients and optimizer states.
2. **Prepare your dataset as JSONL.** One JSON object per line with an instruction and an output. Expected result: a file like `train.jsonl` that loads without errors.
Пример:
{"instruction": "Classify the ticket", "output": "billing"}
{"instruction": "Classify the ticket", "output": "technical"}
This works because the trainer expects plain text pairs, not a database.
3. **Load the base model and tokenizer.** Use `AutoModelForCausalLM` and `AutoTokenizer` from `transformers`. Expected result: the model loads to GPU without an out-of-memory error.
Пример:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-0.8B")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-0.8B")
This works because the model is downloaded once and cached locally.
4. **Attach a LoRA adapter via `peft`.** Set a small rank (r=8 or 16) and target the attention layers. Expected result: `model.print_trainable_parameters()` shows something like "0.5% of parameters trainable". This works because the base weights stay frozen and only adapter matrices update — that's what makes single-GPU training possible.
Пример:
from peft import LoraConfig, get_peft_model
cfg = LoraConfig(r=8, lora_alpha=16, target_modules=["q_proj", "v_proj"])
model = get_peft_model(model, cfg)
model.print_trainable_parameters()
5. **Run training for 1–3 epochs on your dataset.** Use a small batch size (1–4) with gradient accumulation. Expected result: the loss decreases over steps, e.g. from ~2.5 to ~1.0. This works because a narrow task needs few updates, not many epochs.
6. **Save the adapter and test it.** Call `model.save_pretrained("my-lora-adapter")`, then load the base model plus adapter and ask it a question from your domain. Expected result: answers follow your dataset's style. This works because the adapter stores only the learned delta, keeping the file small (tens of MB instead of GB).
Before shipping your fine-tuned model anywhere, check quality properly — see How to Test and Evaluate AI Models: A QA Guide for Machine Learning.
Checklist before you start training:

Can I fine-tune Phi-4 14B on one GPU?
Yes, but it needs a high-VRAM consumer card; for smaller GPUs start with Qwen3.5-0.8B or Gemma 2B.
Is a fine-tuned SLM a replacement for a large LLM?
No. SLMs have their own niche — narrow tasks and agentic systems — not a general replacement for LLMs.
How much data do I need for LoRA?
For a narrow task, a few hundred clean examples in JSONL format are usually enough to see a style/format change.
Which small model should a beginner pick?
Start with Qwen3.5-0.8B or a small Gemma (2B) — they run on consumer hardware and leave VRAM headroom for training.