Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/. 0. Inputs Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in). Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets. Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it." What Evals really does (only show these — no invented features, numbers, rankings or names) Ask: type what you want to make in the composer ("What would you like to make?"). Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct. Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session). Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B"). Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status. Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups). Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards. Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures. End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag. 1. Reverse-engineer the reference first (frames are the truth) Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions: every cut (frame number), shot order and durations, average shot length, cuts per 10s camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame) type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change. 2. Scenario Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid. 3. Design Write DESIGN.md once and make every section obey it (no per-section drift): palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0) card treatment: real UI captures, rounded 14–18px corners, soft layered shadow Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features. 4. Assets (all real, all new) Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills. Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task. Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films. Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour. New music track (detect BPM and first downbeat with a script) and a fresh SFX set. 5. Motion rules (this is the 감도 — non-negotiable) One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone. Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it. Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched). Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut. Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal. Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains. Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops. No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts. Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks). Every cut on the beat ±1 frame; visual hit and SFX on the same frame. 6. Speed, edit, impact Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames. Quantize every segment to 8th notes; round segment start frames up. Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo): zoom punch +8.5%, decaying τ = 0.09s crop shake 16px at ~22Hz, decaying τ = 0.1s rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split). Find hit frames from frame-diff peaks and by looking. 7. Sound Music bed carries the film; no ducking. SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps. At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut. Master at −14 LUFS, AAC 192k. 8. QA before you report (mandatory) motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts. strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs. Check text readability, safe margins, and that every number/name on screen exists in the real product. Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.
Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.
Discussion