# `data/raw/` — your source data goes here

`scripts/prepare_data.py` reads everything in this folder and produces:
- `data/sft/train.jsonl` + `data/sft/eval.jsonl` — instruction pairs for QLoRA fine-tuning.
- `data/kb/*.md` chunks — the knowledge base that RAG retrieves from.

## Accepted formats

### 1. Q&A pairs → used for **SFT** (fine-tuning behavior)
A `.jsonl` file where each line is one object. Minimal shape:

```json
{"question": "How do I file a claim for lost luggage?", "answer": "File within 30 days via the app...", "source": "claims-guide#lost-luggage"}
```

`source` is optional but recommended — it teaches the model to cite.
See `qa_pairs.jsonl` in this folder for a working example.

### 2. Policy / FAQ documents → used for **RAG** (facts)
Markdown (`.md`), plain text (`.txt`), or CSV. Markdown is preferred because headings
become citation anchors. See `policy_travel_basic.md` for an example.

- `.md` / `.txt`: chunked by the prep script; the nearest heading becomes the `source`.
- `.csv`: expects a `text` column (and optional `source` column).

> PDFs: convert to markdown/text first (e.g. `pdftotext` or the `pdf` skill), then drop here.

## The sample files
`qa_pairs.jsonl` and `policy_travel_basic.md` are **placeholders** so the pipeline runs
end-to-end today. Replace them with your real data — same formats.
