# SocialmediaAi

A small, locally-run model that turns a topic brief into one short-form video concept
packaged for YouTube, Instagram and Pinterest — distilled from posts that actually
performed, across four niches: mental wellness, 3D printing, personal finance and
long-term consumer technology.

## Two decisions everything else follows from

**Real winners are the ground truth, not a teacher model.** A teacher (`gpt-5.4-mini`)
is used for exactly one thing: reconstructing the *brief* that would have produced each
high-engagement post. The post itself stays the training target. This is the cheap half
of the problem, and it means the model learns from what earned attention rather than
from a large model's impression of what sounds good.

**Research is code, not model capability.** Teaching a 4B model to browse is the
expensive, fragile part and it buys nothing here. `serve/research.py` retrieves from the
corpus of posts that already performed and assembles a brief; the model only writes. The
scraped corpus therefore pays for itself twice — training set and live retrieval index.

## Pipeline

```
YouTube Data API v3 ─┐
Apify (IG, Pinterest)┴─► data/raw/*.jsonl
                              │  build/score.py      rank against each creator's own baseline
                              ▼
                        curated.jsonl + holdout.jsonl (300, carved before the teacher runs)
                              │  collect/transcripts.py   real creator wording for hook/script
                              │  build/backtranslate.py   teacher reconstructs the brief
                              ▼
                        briefs.jsonl ──► build/pack.py ──► train.jsonl / val.jsonl
                              │                                    │
                              │                          train/lora.py on a rented GPU
                              │                                    ▼
                              │                          GGUF Q4 ──► Ollama, this box
                              ▼                                    │
                     serve/research.py ────────► brief ────────────┴──► serve/generate.py
```

## Setup

```bash
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
cp .env.example .env          # add YOUTUBE_API_KEY and APIFY_TOKEN
```

## Running it

```bash
# 1. Seed the creator lists, then prune them by hand. This file matters more than
#    any hyperparameter — twenty good channels beat two hundred noisy ones.
python collect/youtube.py --discover --niche finance
$EDITOR creators.yaml

# 2. Collect. YouTube is free; Apify is metered, so it is the one to run second.
python collect/youtube.py
python collect/apify.py --platform instagram
python collect/apify.py --platform pinterest --actor <actor-from-apify-store>

# 3. Curate, then fetch real transcripts for the winners only
python build/score.py
python collect/transcripts.py

# 4. Reconstruct briefs (Batch API, half price, up to 24h)
python build/backtranslate.py --limit 20 --sync   # taste the output first
python build/backtranslate.py
python build/backtranslate.py --split holdout     # briefs for eval
python build/pack.py

# 5. Train — see train/README.md. Upload only train.jsonl and val.jsonl.

# 6. Serve
ollama create socialmedia -f serve/Modelfile
python serve/research.py --niche printing3d --topic "bed adhesion" --list
python serve/generate.py --niche printing3d --topic "bed adhesion"

# 7. Evaluate
python eval/run.py --model socialmedia  --tag tuned
python eval/run.py --model llama3.2:1b --fewshot 3 --tag fewshot
python eval/judge.py --a out/gen-tuned.jsonl                          # vs real winners
python eval/judge.py --a out/gen-tuned.jsonl --b out/gen-fewshot.jsonl
```

Every collector is resumable and checkpoints as it goes, so a quota wall, a failed
Apify run or a Ctrl-C never costs more than the page in flight.

## What to expect

Judged against real top-performing posts, **45% is a genuinely good win rate** and 30%
is still a useful model — the comparison set is the top ~18% of each creator's output,
not average content.

The comparison that actually decides things is `tuned` vs `fewshot`. If prompting an
untuned model with retrieved examples from the same corpus matches the fine-tune, then
the retrieval pipeline is the product and the fine-tune is optional. The corpus is
valuable either way, so that outcome costs a few dollars of GPU, not the project.

## Known constraints on this box

- **Generation is slow.** 2 CPU cores, no GPU. A 1B model takes ~90s per pack; a 4B
  will take several minutes. That is fine for producing content in batches, which is
  the intended use, but it is not interactive and it is not a live API.
- **Run the eval on the GPU pod before terminating it.** 300 holdout briefs at CPU
  speed is many hours; on the pod it is minutes. Use `--limit` for spot checks here.
- **RAM is tight.** ~1GB free with `llama-server` holding 1.8GB. A 4B at Q4_K_M needs
  about 3GB resident. Free that up before `ollama create`, or export a 3B instead.
- **Instagram and Pinterest collection runs against those platforms' terms**, whoever
  executes it. Apify absorbs the operational risk of blocks and breakage, not the terms
  question. The corpus is used for style learning, and the near-duplicate gate in
  `eval/run.py` is what verifies the model writes rather than recites.

## Cost

| | |
|---|---|
| YouTube Data API | free (10k units/day) |
| Apify, IG + Pinterest | $10–25 |
| Teacher backtranslation (~1.6k tokens/example, Batch API) | $8–15 |
| GPU, 2–3 LoRA runs on a 4090 | ~$4 |
| Judging | ~$3 |
