Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/. 0. Inputs Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in). Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets. Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it." What Evals really does (only show these — no invented features, numbers, rankings or names) Ask: type what you want to make in the composer ("What would you like to make?"). Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct. Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session). Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B"). Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status. Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups). Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards. Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures. End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag. 1. Reverse-engineer the reference first (frames are the truth) Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions: every cut (frame number), shot order and durations, average shot length, cuts per 10s camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame) type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change. 2. Scenario Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid. 3. Design Write DESIGN.md once and make every section obey it (no per-section drift): palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0) card treatment: real UI captures, rounded 14–18px corners, soft layered shadow Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features. 4. Assets (all real, all new) Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills. Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task. Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films. Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour. New music track (detect BPM and first downbeat with a script) and a fresh SFX set. 5. Motion rules (this is the 감도 — non-negotiable) One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone. Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it. Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched). Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut. Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal. Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains. Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops. No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts. Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks). Every cut on the beat ±1 frame; visual hit and SFX on the same frame. 6. Speed, edit, impact Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames. Quantize every segment to 8th notes; round segment start frames up. Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo): zoom punch +8.5%, decaying τ = 0.09s crop shake 16px at ~22Hz, decaying τ = 0.1s rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split). Find hit frames from frame-diff peaks and by looking. 7. Sound Music bed carries the film; no ducking. SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps. At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut. Master at −14 LUFS, AAC 192k. 8. QA before you report (mandatory) motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts. strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs. Check text readability, safe margins, and that every number/name on screen exists in the real product. Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.
Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.
Open one to inspect its creator, recorded recipe and run record.
Task outputs, votes and discussion
Questions and reactions stay attached to the task, across every output.
Discussion