# DeepFeline Value

Use Gemini 3.1 Pro as the teacher: it writes value-investing analyses in the
DeepFeline persona (Buffett / Burry / Roaring Kitty, per `persona/system_prompt.txt`)
over point-in-time fundamentals, and those analyses are distilled onto Gemma 4 E4B
(same model family; ~4.5B effective params, 128K context) so the final model runs
locally.

The final model produces **analysis reports** (moat, financial forensics, intrinsic
value, contrarian setup, risks, verdict) and **screening/ranking** over candidate
lists. It is a research assistant, not a trading bot — do not wire it to real money.

---

## Hardware plan

| Stage | Where | Why | Est. cost |
|---|---|---|---|
| Corpus + dataset build | **Local** (CPU) | Just Python + API calls | $0 |
| Teacher generation (Gemini 3.1 Pro, Batch API) | **Local** — needs only GEMINI_API_KEY (billed) | Batch mode is 50% off (~$1/$6 per MTok); no GPU, no fine-tune, no seed examples | ~$15–35 for ~420 prompts × 3 samples |
| Student fine-tune (Gemma 4 E4B QLoRA) | **Google Colab** (free T4) or 1× 4090 rental | GTX 1070 (Pascal, 8GB) can't run Unsloth/bitsandbytes; Unsloth's Gemma 4 E4B notebook fits a free Colab T4 | $0–15 |
| Final inference | **Local** — GTX 1070 via Ollama | E4B @ Q4 GGUF ≈ 4.5GB | $0 |

## Pipeline

```
Phase 1  CORPUS        download/extract Buffett letters, Burry writings, RK transcripts,
                       your PDFs (Valuation, Dalio, Shiller)  →  corpus/*.txt
Phase 2  FUNDAMENTALS  point-in-time SEC EDGAR data for (ticker, as-of-date) pairs
                       →  datasets/fundamentals/*.json
Phase 3  PROMPTS       persona + fundamentals snapshot + retrieved corpus excerpts
                       →  datasets/teacher_prompts.jsonl
Phase 4  TEACHER       (local) Gemini 3.1 Pro via the Batch API writes persona
                       analyses for every prompt, 3 lenses per prompt
                       →  datasets/teacher_outputs.jsonl
Phase 4.5 VERIFY       (local) DS-STAR-style audit pass: a Gemini Flash batch job
                       checks every analysis for invented figures, lookahead
                       leakage, and broken verdict blocks  (~$3)
                       →  datasets/teacher_outputs_verified.jsonl
Phase 5  FILTER        score each analysis against actual forward returns; keep chains
                       that reasoned well AND aged well (rejection sampling)
                       →  datasets/student_sft.jsonl
Phase 6  STUDENT       (Colab) SFT Gemma 4 E4B on filtered teacher outputs
                       →  export merged model → GGUF Q4 → Ollama
Phase 7  EVAL          held-out TIME split (train ≤ 2022-12, test 2023+); backtest
                       verdicts vs SPY using the LSTM project's harness as reference
```

## Run order

```bash
pip install -r requirements.txt

# Phase 1 — corpus (local)
python src/corpus/download_buffett_letters.py          # fetches berkshirehathaway.com letters
python src/corpus/extract_pdfs.py                      # extracts your Training Data PDFs
# Burry + Roaring Kitty: drop source files into corpus/burry/ and corpus/roaring_kitty/
# (RK YouTube transcripts: yt-dlp --write-auto-sub --skip-download <url>)

# Phase 2+3 — dataset (local; set SEC_EMAIL first)
python src/data/edgar.py --tickers datasets/tickers.csv --dates datasets/asof_dates.txt
python src/data/build_prompts.py

# Phase 4 — teacher (local; set GEMINI_API_KEY first)
python src/train/generate_teacher_data.py

# Phase 4.5 — verify (local)
python src/data/verify_teacher_outputs.py

# Phase 5 — filter (local; uses the verified file automatically if present)
python src/eval/score_and_filter.py

# Phase 6 — student (CLOUD), then locally:
#   ollama create deepfeline -f Modelfile

# Phase 7 — eval (local)
python src/eval/backtest_verdicts.py
```

## Local inference — speculative decoding (optional speedup)

Ollama works fine (`ollama create deepfeline -f Modelfile`), but for ~1.5–2.5x
faster generation on the GTX 1070 use llama.cpp's server with the stock
**Gemma 4 E2B** as a draft model — it proposes tokens, the fine-tuned E4B
verifies them, output is identical to running E4B alone:

```
llama-server -m deepfeline_e4b.Q4_K_M.gguf ^
             -md gemma-4-E2B-it.Q4_K_M.gguf ^
             --ctx-size 8192 --draft-max 8 -ngl 999 -ngld 999
```

Notes:
- Draft + target at Q4 is ~7GB; if the 1070's 8GB OOMs, lower `-ngl` (target
  layers to GPU) first and keep the small draft fully on GPU (`-ngld 999`).
- Speedup is NOT free lunch elsewhere: same tokens billed/generated, same
  quality — it only cuts wall-clock latency.
- If acceptance rates are poor after fine-tuning (persona style drifts from
  stock E2B), a cheap LoRA of E2B on the same student_sft.jsonl fixes it.
- `src/eval/backtest_verdicts.py` speaks the OpenAI-compatible API, so it works
  against llama-server (port 8080) and Ollama (port 11434) alike.

## Critical rules baked into the pipeline

1. **Point-in-time only.** Every fundamentals snapshot excludes anything *filed* after
   the as-of date. Otherwise the model "predicts" the past and the backtest lies.
2. **Time-based holdout.** Never a random split. Train ≤ 2022, evaluate 2023+.
3. **Verdicts must be falsifiable.** Every generated analysis ends in a structured
   verdict block (direction, conviction 1–10, horizon) so Phase 5/7 can score it.
4. **The disclaimer stays.** The persona's "not financial advice" line is part of the
   training target, not decoration.

## Honest expectations

A 4B-effective-parameter model distilled from a 27B will faithfully reproduce the
*analysis style and process* — checklist discipline, forensic framing, contrarian
questions. It will not have an information edge and will not reliably beat an index.
Its value is as a fast, local, always-available first-pass analyst whose reasoning
you check, in the same way you'd treat a junior analyst's memo.
