# Running the fine-tune on a rented GPU

The collection box has no GPU. Rent one by the hour, train, pull the GGUF back, cancel
the pod. A LoRA over ~10k examples on a 4B model is roughly 2-3 hours on a 4090.

## 1. Pick a pod

RunPod or Vast.ai, **RTX 4090 24GB** (~$0.35-0.50/hr) is plenty for 4B at 4-bit.
Choose a PyTorch 2.x + CUDA 12.x template. An A100 40GB works too and is faster, but
costs more per unit of work at this size.

## 2. Install

```bash
pip install unsloth
pip install --no-deps trl peft accelerate bitsandbytes
```

Install `unsloth` first and let it resolve torch/transformers — it pins a working
combination. Installing TRL first is the usual cause of a broken environment.

## 3. Upload the dataset

From the collection box:

```bash
scp -P <port> data/train.jsonl data/val.jsonl root@<pod-host>:/workspace/data/
```

Only these two files are needed. The raw corpus stays home.

## 4. Train

```bash
cd /workspace && python lora.py 2>&1 | tee train.log
```

Watch the eval loss printed every 50 steps. If it turns upward while train loss keeps
falling, you are overfitting — rerun with `--epochs 2`, and if that still overfits,
`--lr 1e-4`. With a few thousand examples this is common; it is the main reason to
budget for two or three runs rather than one.

The script prints one fully formatted example before training starts. **Read it.** If
the chat template markers don't match the `instruction_part` / `response_part` strings
in `lora.py`, the completion-only masking silently does nothing and you will train on
the briefs as well as the packs.

## 5. Bring the model home

```bash
scp -P <port> -r root@<pod-host>:/workspace/out/gguf-q4_k_m ./out/
scp -P <port> -r root@<pod-host>:/workspace/out/adapter ./out/
```

Keep the adapter. It is small, and it is what you re-quantise from later without
paying for another training run.

Then terminate the pod. Storage bills accrue even when it's stopped.

## Version drift

Unsloth and TRL move quickly. Two names change more often than the rest:

- `SFTConfig(max_seq_length=...)` has been `max_length` in some TRL releases
- `SFTTrainer(tokenizer=...)` has been `processing_class=` in some TRL releases

If the script raises a `TypeError` on one of those, swap the name and continue.
