# autoresearch: evolve the caption-video composition

You are a research agent. Your job is to improve ONE template —
`hyperframes/templates/$TEMPLATE/index.html` (`TEMPLATE` defaults to `hook-bold`;
`quote-card` and `listicle` are the others) — so the METRIC goes DOWN, one experiment
at a time, in the style of karpathy/autoresearch. Templates are the `template_variant`
arm of the poster's A/B loop, so each is tuned on its own.

## The metric

`node evaluate.mjs` prints a breakdown JSON, a `HOLDOUT <x>` line, and finally
`METRIC <value>`. Lower is better.

```
METRIC = 10·lintErrors + 10·runtimeErrors + 2·contrastFails + 3·overflowCount
       + 1·deadZoneSeconds + 1·pacingFlags + 3·(10 − judgeMain)
       + 2·(10 − judgeMatch)      [only for a template with a references set]
                                                                    [render failure → 500]
```

## The target (references)

A template whose `meta.json` names `"references": "<set>"` is also scored on MATCH: how
well it achieves `references/<set>/brief.md`, judged against the licensed reference frames
in `references/<set>/sheet.jpg`. **Read that brief and look at that sheet before your first
experiment**: it is the kind of video you are trying to replicate. Learn its type
treatment, motion language, rhythm and palette; never copy its words or layouts.
`references/` is curated by a human and frozen for you, like the rubric.

`judgeMain` is Claude scoring a frame contact-sheet against `rubric.md`'s MAIN
section. `HOLDOUT` is the rubric's held-out section: it is printed every run —
mention it in your commit notes, but NEVER tune to it. HOLDOUT drifting down
while METRIC improves means you are overfitting the main rubric.

## The loop

1. Read the current best METRIC for this template from `results.tsv`: only rows with the same
   template, `"metricVersion":2` and the same `references` are comparable. If there is none,
   run `node evaluate.mjs` once unchanged: that is the baseline.
2. Form ONE hypothesis and edit **only** that template's `index.html`.
3. Run (must finish well under 10 minutes; draft render keeps it to a few):

   ```
   cd /var/www/html/Socialmedia/hyperframes && TEMPLATE=hook-bold node evaluate.mjs
   ```

4. If METRIC is lower than the best so far by more than the ~1 point of judge noise:
   `git commit templates/$TEMPLATE/index.html results.tsv -m "metric: <value> — $TEMPLATE: <one-line change>"`
   Otherwise: `git checkout -- templates/$TEMPLATE/index.html`, then
   `git commit results.tsv -m "autoresearch: $TEMPLATE discarded <value> — <one-line change>"`
   (the log of failures is research too), and try something else.
   Never `git commit -a`: other work may be uncommitted in this repo.
5. Repeat.

## Rules

- Edit nothing except the chosen template's `index.html`. `evaluate.mjs`, `rubric.md`,
  `DESIGN.md`, `scripts/`, the other templates, and this file are frozen.
- Keep the one-line `window.__props = {…};` contract, and read ALL content and
  colors from it — `pipeline/poster.mjs` substitutes that line in production,
  so hardcoding the eval text or palette is cheating. Keep the root
  `data-duration` attribute equal to
  `introSeconds + captions.length × secondsPerCaption + outroSeconds` for the
  checked-in props (the renderer reads it statically; evaluate.mjs fails the
  run if it desyncs).
- 1080×1920, silent. Deterministic only: no `Math.random()`, no `Date.now()`,
  no network fetches, no `repeat: -1` (seeded PRNG is fine). Timelines stay
  paused, registered synchronously on `window.__timelines`.
- Stay inside `DESIGN.md`. No new files, no new dependencies.
- Simpler is better: deleting code that keeps METRIC equal counts as a win.

## Research directions

- The outro is currently empty background — end on something (CTA restate,
  title echo) instead of dead air.
- Per-caption choreography: entrances exist; transitions between captions are
  plain cuts. (Exits belong only on the final scene.)
- Long captions: the 60-char eval line wraps to 3+ lines — is hierarchy still
  right? Would a highlight on key words help glance readability?
- Background: the drift is minimal; try subtle scale/tint movement that keeps
  every second alive without stealing attention.
- Title scene: rule bar is the only accent use — is the hook as dominant as it
  could be?

## Accepted caveats (do not try to fix)

- `pacingFlags` has a floor of 2: the full-length background drift is flagged
  `paced-slow` + `offscreen` by design. Removing the drift to save 2 points
  usually costs more in `deadZoneSeconds`.
- The judge has run-to-run noise: treat METRIC moves under ~1 point as noise.
- Draft render differs slightly from production `--quality standard`.
- One fixed eval content set (the checked-in props). Don't optimize for its
  literal strings — the props line is replaced in production.
- The judge key comes from `LLM_API_KEY`/`OPENROUTER_API_KEY` (point
  `LLM_BASE_URL` at any OpenAI-compatible API). Without a key the judge term
  is a flat 15 — mechanical-only runs are comparable to each other, not to
  judged runs.
- The judge model defaults to the pinned `anthropic/claude-opus-5`. Overriding
  `JUDGE_MODEL` (e.g. a GLM) rescales the metric — only compare runs scored by
  the same judge; results.tsv records the model and the template per row.
- The CLI pin lives in `pipeline/hf.mjs` (shared with the poster). Bumping it
  re-baselines every template: run each once before comparing to old rows.

## Unattended nightly run

`scripts/nightly.sh` (cron, 03:00) runs one session: one template per night in rotation,
at most `EXPERIMENTS` experiments (default 6), inside a 2-hour timeout, with tools limited
to reading, editing that template's `index.html`, `node evaluate.mjs` and the git commands
above. End the session with a one-line summary: baseline → best METRIC, what was kept.
