# Travel-Insurance AI Agent — Energy-Efficient (Hybrid QLoRA + RAG)

A domain-specialized conversational agent for **travel insurance** designed to run a **small model
on a single GPU** with a strong focus on **energy efficiency**.

## The idea

- **QLoRA fine-tune** teaches the model *behavior*: domain tone, answer structure, the habit of
  citing sources, and refusal/guardrail patterns.
- **RAG** supplies the *facts*: coverage limits, exclusions, claim steps — retrieved from your
  policy documents at query time, so knowledge stays fresh **without retraining**.
- A **semantic cache** in front of the model means repeated questions never touch the GPU.

```
User query
  → [1] Semantic cache lookup (hit → return, zero GPU)   ← biggest energy saver
  → [2] Retriever: embed query → vector search over policy docs → top-k chunks
  → [3] LLM (fine-tuned 3B, int4) with retrieved context → grounded answer + citations
  → [4] Guardrail/refusal check (no personalized advice, cite source, escalate to human)
  → cache the result
```

## Energy-efficiency levers (built in)

| Lever | Effect |
|-------|--------|
| 3B-class model @ int4 | Largest single lever — a fraction of a 7B–70B model's draw |
| Semantic cache | Repeated FAQs cost ~0 GPU |
| RAG (not bigger model) | Small model stays accurate on facts it never learned |
| vLLM paged KV-cache + batching | High throughput per watt |
| Capped `max_tokens` + stop sequences | No wasted generation |

## Token efficiency (minimizing cost + energy)

Fewer tokens in/out is the main cost and energy lever at the application layer. Inspired by the
KV-cache-compression / sparse-context principles in
[arXiv:2606.19348](https://arxiv.org/html/2606.19348v1) (which are model-architecture techniques we
can't graft onto Gemma 4), we apply their spirit at the layers we *do* control:

| Technique | Where | Token effect |
|-----------|-------|--------------|
| **Router** skips trivial/greeting inputs | `pipeline._route` | Zero tokens for non-questions |
| **Semantic cache** on repeated questions | `cache.py` | Zero tokens on a hit |
| **Compact system prompt** (~45 vs ~150 tok) | `prompt.SYSTEM_PROMPT` | Saved on *every* call; behavior comes from the fine-tune |
| **Context dedup + token budget** | `prompt.compress_context` | Caps input context to `retrieval.context_token_budget` |
| **Smaller chunks, tight top-k** | `config.data` / `retrieval` | Fewer, denser context tokens |
| **Capped `max_tokens`** (256) | `config.generation` | Bounds the priciest (output) tokens |
| **fp8 KV cache + prefix caching** | vLLM serve flags | Less memory/energy per token; static prompt not re-prefilled |
| **n-gram speculative decoding** | `scripts/serve_vllm.py` | Lower latency + energy *per request* (not token count); RAG-ideal |

Every response carries `prompt_tokens` / `completion_tokens` / `total_tokens` / `route`, and
`scripts/evaluate.py` reports **avg tokens per query** and **zero-token query count** so you can
track the objective and tune the knobs above.

### A note on speculative decoding (DeepSpec) and DwarfStar (ds4)

Two things people reach for that **don't reduce token usage** — worth knowing what they *do*:

- **Speculative decoding** (e.g. `deepseek-ai/DeepSpec`): a tiny "draft" proposes tokens the target
  verifies in parallel, producing the **same output** faster. We use vLLM's **n-gram / prompt-lookup**
  variant (`efficiency.speculative.method: ngram`) — it needs **no extra model** and is a great fit
  for RAG, where answers copy verbatim from the retrieved context. It lowers latency and
  energy-per-request; it does **not** change how many tokens are billed/generated. (For higher
  acceptance at scale you could train an Eagle3 draft with DeepSpec and set `method: eagle` +
  `draft_model:` — but running a second model rarely pays off for an already-tiny E2B/E4B target.)
- **`antirez/ds4` (DwarfStar)**: an inference engine **only for DeepSeek-V4** MoE models. It would
  mean dropping small Gemma for a large MoE model needing 96GB+ RAM / DGX-class hardware — the
  opposite of this project's small/low-energy goal — so we deliberately do **not** use it here.

## Quick start

```bash
# 1. Install. CPU (RAG/cache/API layer):   pip install -e .
#    GPU box (training + quant + serving):  pip install -e ".[train,quant,serve,eval]"
#    Gemma 4 is gated — accept the license on its HF page, then:  huggingface-cli login

# 2a. Drop policy PDFs/Word docs in data/source_docs/ and convert them to citable markdown
#     (headings + benefit tables preserved). Skip if your data is already md/txt/csv/jsonl.
python scripts/ingest_docs.py --input data/source_docs --config config/config.yaml

# 2b. Build the SFT set (from Q&A jsonl) + RAG knowledge base (from all docs in data/raw/):
python scripts/prepare_data.py --config config/config.yaml

# 3. Build the RAG index (works even before fine-tuning)
python scripts/build_index.py --config config/config.yaml

# 4. Fine-tune with QLoRA — Gemma 4 E4B ~17GB VRAM, E2B ~8-10GB  (GPU)
python scripts/train_qlora.py --config config/config.yaml

# 5. Merge LoRA + quantize to int4 (w4a16) for serving  (GPU)
python scripts/merge_and_quantize.py --config config/config.yaml

# 6a. Start the model server. serve_vllm.py applies the efficiency settings from config
#     (fp8 KV cache + prefix caching + n-gram speculative decoding). Add --dry-run to inspect.
python scripts/serve_vllm.py --config config/config.yaml
#    or Ollama:  ollama run gemma4:e4b-it-qat        (serves OpenAI-compat on :11434)

# 6b. Start the agent API (cache -> retrieve -> llm -> guardrails)
uvicorn app.server:app --host 0.0.0.0 --port 8000

# 7. Ask a question
curl -s localhost:8000/chat -H 'content-type: application/json' \
  -d '{"query":"Does my policy cover a cancelled flight due to a storm?"}' | jq
```

Fast start: skip steps 4–5 and point `serving.*` at Google's official QAT weights
(`ollama run gemma4:e4b-it-qat`) to get a working RAG agent immediately, then layer the fine-tune
in later for domain tone/citations.

## Run locally without a big GPU (verified)

The RAG + agent stack runs on CPU with a small local model via **Ollama** — this is the default in
`config.yaml` (`serving.base_url` = Ollama, `served_model: gemma3:4b`). Fine-tuning still needs a
16 GB+ GPU, but you can develop and use the whole agent locally:

```bash
pip install -e .                              # CPU deps (add torch<2.4-compatible transformers)
ollama pull gemma3:4b                          # or any small local model
python scripts/build_index.py --config config/config.yaml
python -m uvicorn app.server:app --port 8000
curl -s localhost:8000/chat -H 'content-type: application/json' \
  -d '{"query":"Are pre-existing conditions covered?"}'
```

Verified behavior on this path: grounded answers with a **single precise citation**, off-topic
questions escalate at **0 tokens** (retrieval score below `min_score` → LLM never called), repeated
questions served from cache at **0 tokens**, and live `prompt/completion/total` token counts from
the backend.

## Voice web UI (localhost)

A browser voice interface with a Sesame-style audio-reactive orb is served by the same FastAPI app:

```bash
python -m uvicorn app.server:app --port 8000
# open http://localhost:8000 in Chrome or Edge, click the mic, and speak
```

- **Speech-to-text runs locally** via faster-whisper (`pip install -e ".[voice]"`): the browser
  records your mic, POSTs the audio to the `/transcribe` endpoint, and it's transcribed **on this
  machine** — audio never goes to a cloud service. Tap the mic to start, tap again to stop. If the
  voice deps aren't installed, it falls back to the browser Web Speech API, then to the text box.
- **Text-to-speech runs in the browser** (Web Speech API) — no extra service needed.
- The **visualizer** is 5 chunky bars driven by the Web Audio `AnalyserNode`: while listening each
  bar tracks a band of *your* voice (teal), they bounce in a talking cadence while the agent speaks
  (pink), and breathe when idle (blue).
- Each answer shows its **route** (llm/cache/router), **token count**, and **citations** as chips.
- Must be opened over `http://localhost` (not `file://`) so the mic is allowed. Chrome/Edge have the
  best Web Speech support; other browsers fall back to the text box.

### LLM backend

Set by `serving.*` in `config.yaml`. Any OpenAI-compatible endpoint works:

| Backend | Setup |
|---------|-------|
| **OpenRouter** (default) | `setx OPENROUTER_API_KEY "sk-or-..."` (new shell), then run the server. Model: `google/gemma-4-31b-it:free` (or `google/gemma-3n-e4b-it` to mirror the deployed small model). The key is read from the env var — never stored in the repo. |
| **Ollama** (local, no key) | Uncomment the Ollama lines in `serving`; `ollama pull gemma3:4b`. |
| **vLLM** (your GPU server) | Uncomment the vLLM lines; serve your fine-tuned int4 model with `scripts/serve_vllm.py`. |

Note: on a metered backend like OpenRouter, the token-minimization work (router, cache, compact
prompt, context budget) translates **directly into lower $ per query**, not just energy.

> Privacy: with the `[voice]` deps installed, speech-to-text is fully on-device (faster-whisper),
> so mic audio never leaves the machine. Only the browser Web Speech *fallback* (used when those
> deps are missing) sends audio to the browser vendor. Model/size is set in `config.yaml` under
> `stt` (`tiny.en` → `small.en`).

## Tests

```bash
python tests/test_agent.py          # or: pytest tests/test_agent.py
```

Self-contained (no downloads/GPU): fakes the embedder/retriever/LLM and checks routing, caching,
context compression, citation mapping, and guardrails.

## Evaluate + measure energy

```bash
python scripts/evaluate.py --config config/config.yaml
# → accuracy / groundedness on data/sft/eval.jsonl
# → Wh-per-query and cache hit-rate (via codecarbon + nvidia-smi sampling)
```

## Layout

```
config/config.yaml     model ids, paths, top_k, cache thresholds, generation params
data/raw/              your source docs (Q&A, policy PDFs/markdown/CSV)
data/sft/              generated instruction JSONL (train/eval)
data/kb/               chunked knowledge base for RAG
scripts/               prepare_data · train_qlora · merge_and_quantize · build_index · evaluate
src/agent/             retriever · cache · prompt · guardrails · llm
app/server.py          FastAPI /chat endpoint
```

## Notes on model choice

Default base model is **`google/gemma-4-E4B-it`**. The "E" = *effective* parameters: Per-Layer
Embeddings keep the runtime memory footprint low (E4B ≈ 4B effective, E2B ≈ 2B), which is exactly
the energy profile we want. Gemma 4 also ships **quantization-aware-trained (QAT) w4a16 weights**
with day-0 vLLM/Ollama support, so int4 serving keeps near-fp16 quality.

| Variant | `model.base_id` | QLoRA VRAM | When |
|---------|-----------------|-----------|------|
| Leanest | `google/gemma-4-E2B-it` | ~8–10 GB | Max energy efficiency; RAG carries the facts |
| Default | `google/gemma-4-E4B-it` | ~17 GB | Better reasoning/nuance, still very efficient |

Swap `model.base_id` in `config/config.yaml` to switch. Gemma 4 is gated on Hugging Face — accept
the license and `huggingface-cli login` before downloading. Training uses **Unsloth** (day-0 Gemma 4
support, ~2× faster, lower VRAM); see `scripts/train_qlora.py` for a plain-transformers fallback.
