Evals landscape demo-video prompt (approved "감도") Copy everything below the line into a new session. Make a ~60s landscape product demo video for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. Horizontal 16:9, 1920×1080, 30fps (for YouTube / website / X timeline). Not a vertical Short/Reel — no 9:16 output, no vertical crops, no phone safe-zones. Music-locked, no narration — on-screen type tells the story while the real product is demoed end to end. Build it in HyperFrames (HTML + one paused GSAP timeline) under /Users/evan/workspace/video-generator/videos/evals-demo3/. 0. Inputs Product: https://evals.ag (app: https://evals-app.vercel.app — use the logged-in Claude in Chrome session if connected; otherwise public pages). Motion reference (local file, already downloaded): /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its motion feel exactly (it is also 16:9); change only scenario, design and assets. Its 26s cut-down is just the reference for feel — this video runs the full ~60s. Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it." Run it autonomously — do not ask me for input Work in Claude Code with the repo /Users/evan/workspace/video-generator as the working directory. Every path below is absolute; everything you need is already on this machine. Do NOT stop to ask questions or request inputs. Every input is specified here; for anything not specified, pick a sensible default, write the decision into DECISIONS.md, and keep going until the final file is rendered. If something is unavailable, use the fallback and continue: Chrome not connected or a page needs sign-in → use the public pages (https://evals-app.vercel.app home, /compare, leaderboards, explore) with Playwright, and record outputs from their public "Open preview" links. Can't run a new Battle → use public outputs that share the same prompt (the /compare pages show pairs). Music: pick an unused track from /Users/evan/workspace/video-generator/videos/evals-hype/music3/ or music2/. SFX: /Users/evan/workspace/video-generator/videos/evals-hype/sfx2/, ref2/sfx3/, ref2/sfx4/. QA scripts: /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py and strip.sh (also beats.py for BPM). If they fail, write equivalents. Render: npx --yes hyperframes@0.8.101 render -o <out.mp4> --quiet from the project dir (--fps 95/4 etc. for speed re-render). API calls via curl (python has no SSL certs here). What Evals really does (only show these — no invented features, numbers, rankings or names) Ask: type what you want to make in the composer ("What would you like to make?"). Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct. Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session). Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B"). Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status. Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups). Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards. Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures. End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag. 1. Reverse-engineer the reference first (frames are the truth) Study the local reference file frame by frame (ffmpeg frame dumps + frame-diff). Its design notes are in /Users/evan/workspace/video-generator/videos/evals-hype/ref2/SPEC.md and MOTION_FIX.md — read them, but design and copy must be new. Write SPEC.md with measured numbers, not impressions: every cut (frame number), shot order and durations, average shot length, cuts per 10s camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame) type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change. 2. Scenario Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid. 3. Design Write DESIGN.md once and make every section obey it (no per-section drift): palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0) card treatment: real UI captures, rounded 14–18px corners, soft layered shadow Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features. 4. Assets (all real, all new) Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills. Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task. Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films. Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour. New music track (detect BPM and first downbeat with a script) and a fresh SFX set. 5. Motion rules (this is the 감도 — non-negotiable) One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone. Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it. Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched). Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut. Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal. Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains. Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops. No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts. Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks). Every cut on the beat ±1 frame; visual hit and SFX on the same frame. 6. Speed, edit, impact Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames. Quantize every segment to 8th notes; round segment start frames up. Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo): zoom punch +8.5%, decaying τ = 0.09s crop shake 16px at ~22Hz, decaying τ = 0.1s rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split). Find hit frames from frame-diff peaks and by looking. 7. Sound Music bed carries the film; no ducking. SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps. At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut. Master at −14 LUFS, AAC 192k. 8. QA before you report (mandatory) python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts. sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs. Check text readability, safe margins, and that every number/name on screen exists in the real product. Deliver one 1920×1080 file renders/evals-demo3-v1.mp4 (no vertical version) plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.
Evals landscape demo-video prompt (approved "감도")
Copy everything below the line into a new session.
Make a ~60s landscape product demo video for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. Horizontal 16:9, 1920×1080, 30fps (for YouTube / website / X timeline). Not a vertical Short/Reel — no 9:16 output, no vertical crops, no phone safe-zones. Music-locked, no narration — on-screen type tells the story while the real product is demoed end to end. Build it in HyperFrames (HTML + one paused GSAP timeline) under /Users/evan/workspace/video-generator/videos/evals-demo3/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use the logged-in Claude in Chrome session if connected; otherwise public pages).
Motion reference (local file, already downloaded): /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its motion feel exactly (it is also 16:9); change only scenario, design and assets. Its 26s cut-down is just the reference for feel — this video runs the full ~60s.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
Run it autonomously — do not ask me for input
Work in Claude Code with the repo /Users/evan/workspace/video-generator as the working directory. Every path below is absolute; everything you need is already on this machine.
Do NOT stop to ask questions or request inputs. Every input is specified here; for anything not specified, pick a sensible default, write the decision into DECISIONS.md, and keep going until the final file is rendered.
If something is unavailable, use the fallback and continue:
Chrome not connected or a page needs sign-in → use the public pages (https://evals-app.vercel.app home, /compare, leaderboards, explore) with Playwright, and record outputs from their public "Open preview" links.
Can't run a new Battle → use public outputs that share the same prompt (the /compare pages show pairs).
Music: pick an unused track from /Users/evan/workspace/video-generator/videos/evals-hype/music3/ or music2/. SFX: /Users/evan/workspace/video-generator/videos/evals-hype/sfx2/, ref2/sfx3/, ref2/sfx4/.
QA scripts: /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py and strip.sh (also beats.py for BPM). If they fail, write equivalents.
Render: npx --yes hyperframes@0.8.101 render -o <out.mp4> --quiet from the project dir (--fps 95/4 etc. for speed re-render). API calls via curl (python has no SSL certs here).
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Study the local reference file frame by frame (ffmpeg frame dumps + frame-diff). Its design notes are in /Users/evan/workspace/video-generator/videos/evals-hype/ref2/SPEC.md and MOTION_FIX.md — read them, but design and copy must be new. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver one 1920×1080 file renders/evals-demo3-v1.mp4 (no vertical version) plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.
Open one to inspect its creator, recorded recipe and run record.
Task outputs, votes and discussion
Open one to inspect its creator, recorded recipe and run record.