Loading tasks… LLM Gen AI Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first. Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan Cha1 turn·6.9M tokens Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan Cha
GPT-6 Astra/hyperframes 1 turn·331.9K tokens Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan ChaQwen3.8 Max (0902)/hyperframes 1 turn·843.2K tokens Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first. Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan Cha
GPT-6.1 Sol/hyperframes 1 turn·3.8M tokens Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan ChaQwen3.8 Max (0902)/hyperframes 1 turn·5.4M tokens Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first. Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan Cha
GPT-6.1 Sol/hyperframes 1 turn·3.2M tokens Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first. Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.EEvan Cha
GPT-6.1 Sol/hyperframes 1 turn·5.9M tokens START NOW. Do not ask me any questions and do not wait for input — every input is below. Make every unspecified decision yourself, log it in DECISIONS.md, and keep going until ONE finished video file exists. The only deliverable is that one .mp4.
Task
Make ONE ~60s landscape demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
Output: /Users/evan/workspace/video-generator/videos/evals-demo3/renders/evals-demo3.mp4, 1920×1080, 30fps, H.264 + AAC 192k, −14 LUFS. Nothing else to deliver (no vertical version, no variants).
Work dir: repo /Users/evan/workspace/video-generator (Claude Code on this Mac). Build in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-demo3/. Render: npx --yes hyperframes@0.8.101 render -o <out> --quiet from the project dir. Use curl for HTTP/API calls (python has no SSL certs). ElevenLabs key: ELEVENLABS_API_KEY in /Users/evan/workspace/video-generator/.env.
Story (English voiceover + on-screen type)
Hook, first 3s must be instantly clear: real AI-made GAME/VIDEO outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
Type a real game/video task into the Evals composer → Battle → up to 4 setups (Model · Skill · Environment, each landing with real values and moving UI).
Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
Blind compare: two moving outputs, "Which would you choose?" → vote ("I prefer A").
Run record of that same output (real LLM, skill, environment, turn time, tokens).
Setup leaderboard ("Setups, not just models.").
Community wall of real game/video/shorts outputs.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
Content rules
Show ONLY videos, games, 3D scenes, motion videos, shorts made in Evals. No websites / landing pages.
Main thread = the real 4-setup Sonic battle: prompt "Build a polished playable Sonic-style 3D platform game for the browser.", clips /Users/evan/workspace/video-generator/videos/evals-launch/assets/clips/run_a.mp4 … run_d.mp4, run https://evals-app.vercel.app/runs/run_597Q2CRE4FXNJ2R7HZXMVF4Z03.
Capture fresh real UI from https://evals-app.vercel.app with Playwright at 2× DPR: composer, Run-mode menu, setups stepper, compare page, output run record, Leaderboards, Explore → Outputs. If a page needs sign-in, use the public pages.
Record extra community game/video outputs from their public "Open preview" links. Screencast 1600×900, 30fps, 6–7s, with keys pressed or orbit so they move. Reuse the pattern in /Users/evan/workspace/video-generator/videos/evals-launch/scripts/rec.mjs.
No real-person likeness, no invented features/numbers/names. Only show model names you see in captures (e.g. Claude Opus 5.5, GPT-6 Astra).
Voiceover (must flow continuously)
ElevenLabs voice Will bIHbv24MWmeRgasZH58o, model eleven_v4, stability 0.15, similarity 0.8, style 0.8, speed 1.1, with emotion tags in the text ([excited] [confident] [curious] [emphatic] [quick]). Use the /with-timestamps endpoint.
Write ~165–180 words so the voice fills the whole film. Generate it as ONE take and lay it unspliced — never cut it into phrases or time-stretch pieces. Longest silence inside the VO ≤ 0.35s. No clipped words.
The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
Verify with Whisper that every word is there and no tags are spoken. If "Evals" sounds like "evils", respell it "E-vals".
Music & sound
New modern premium BGM, not cheesy. Generate 2–3 candidates with ElevenLabs Music: "modern minimal tech launch, punchy kick, deep 808, crisp hats, glassy synth plucks, Apple/Linear keynote feel, no cheesy EDM risers, no corporate/stock vibe, 120 BPM". Pick the cleanest; detect BPM and the first downbeat.
VO leads at about −16 LUFS. Music ducks about 6dB only while he speaks (smooth envelope). SFX bus ~3dB under the music, above it only on hits.
At 4–6 big hits, layer reverse swell + boom (HP 60Hz) + transient + glitch. Light ticks and whooshes elsewhere — never on every cut.
Look
Fresh premium design system, written once in DESIGN.md:
one sans font, one accent colour
real UI captures as rounded cards (16–22px) with soft shadows
dark/light section switches
Never: outlines/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups.
Motion (non-negotiable — this is the feel)
Reference for feel: /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4; rules in /Users/evan/workspace/video-generator/videos/evals-hype/MOTION_FIX.md.
One camera = a pure function of time with velocity-continuous easing. Deterministic: no timers or randomness.
Two easing families only: snappy ≈25% of remaining distance per frame, floaty ≈12%.
Exits accelerate into the cut; the next shot carries the velocity through it.
Real whip-pans: both shots travel together over 0.3–0.45s with velocity blur.
Overlays are screen-locked.
Overlapping action, no start-stop chains.
Every hold drifts 1–3%. No dead stops, no frozen frames, no double cuts.
Word-by-word reveals 2–4 frames apart; the newest word flashes the accent.
Cuts on the beat ±1 frame.
Speed by re-rendering: render at --fps 30/speedup (e.g. --fps 95/4 = 1.263×), play back at 30fps.
Impact FX in post (ffmpeg) on 4–6 big hits only:
+8.5% zoom punch (τ 0.09s)
16px shake at ~22Hz (τ 0.1s)
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
A lighter version on secondary cuts.
QA (do it, then deliver)
python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat>: no dead stops, no frozen holds, cuts on the beat.
sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> at every transition — look at each strip.
VO RMS gap check (≤ 0.35s) plus a Whisper pass on the final mix.
Text readable and inside the margins; every name and number on screen is real.
Fix anything that fails, re-render, and finish with the single file videos/evals-demo3/renders/evals-demo3.mp4. START NOW. Do not ask me any questions and do not wait for input — every input is below. Make every unspecified decision yourself, log it in DECISIONS.md, and keep going until ONE finished video file exists. The only deliverable is that one .mp4.
Task
Make ONE ~60s landscape demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
Output: /Users/evan/workspace/video-generator/videos/evals-demo3/renders/evals-demo3.mp4, 1920×1080, 30fps, H.264 + AAC 192k, −14 LUFS. Nothing else to deliver (no vertical version, no variants).
Work dir: repo /Users/evan/workspace/video-generator (Claude Code on this Mac). Build in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-demo3/. Render: npx --yes hyperframes@0.8.101 render -o <out> --quiet from the project dir. Use curl for HTTP/API calls (python has no SSL certs). ElevenLabs key: ELEVENLABS_API_KEY in /Users/evan/workspace/video-generator/.env.
Story (English voiceover + on-screen type)
Hook, first 3s must be instantly clear: real AI-made GAME/VIDEO outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
Type a real game/video task into the Evals composer → Battle → up to 4 setups (Model · Skill · Environment, each landing with real values and moving UI).
Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
Blind compare: two moving outputs, "Which would you choose?" → vote ("I prefer A").
Run record of that same output (real LLM, skill, environment, turn time, tokens).
Setup leaderboard ("Setups, not just models.").
Community wall of real game/video/shorts outputs.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
Content rules
Show ONLY videos, games, 3D scenes, motion videos, shorts made in Evals. No websites / landing pages.
Main thread = the real 4-setup Sonic battle: prompt "Build a polished playable Sonic-style 3D platform game for the browser.", clips /Users/evan/workspace/video-generator/videos/evals-launch/assets/clips/run_a.mp4 … run_d.mp4, run https://evals-app.vercel.app/runs/run_597Q2CRE4FXNJ2R7HZXMVF4Z03.
Capture fresh real UI from https://evals-app.vercel.app with Playwright at 2× DPR: composer, Run-mode menu, setups stepper, compare page, output run record, Leaderboards, Explore → Outputs. If a page needs sign-in, use the public pages.
Record extra community game/video outputs from their public "Open preview" links. Screencast 1600×900, 30fps, 6–7s, with keys pressed or orbit so they move. Reuse the pattern in /Users/evan/workspace/video-generator/videos/evals-launch/scripts/rec.mjs.
No real-person likeness, no invented features/numbers/names. Only show model names you see in captures (e.g. Claude Opus 5.5, GPT-6 Astra).
Voiceover (must flow continuously)
ElevenLabs voice Will bIHbv24MWmeRgasZH58o, model eleven_v4, stability 0.15, similarity 0.8, style 0.8, speed 1.1, with emotion tags in the text ([excited] [confident] [curious] [emphatic] [quick]). Use the /with-timestamps endpoint.
Write ~165–180 words so the voice fills the whole film. Generate it as ONE take and lay it unspliced — never cut it into phrases or time-stretch pieces. Longest silence inside the VO ≤ 0.35s. No clipped words.
The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
Verify with Whisper that every word is there and no tags are spoken. If "Evals" sounds like "evils", respell it "E-vals".
Music & sound
New modern premium BGM, not cheesy. Generate 2–3 candidates with ElevenLabs Music: "modern minimal tech launch, punchy kick, deep 808, crisp hats, glassy synth plucks, Apple/Linear keynote feel, no cheesy EDM risers, no corporate/stock vibe, 120 BPM". Pick the cleanest; detect BPM and the first downbeat.
VO leads at about −16 LUFS. Music ducks about 6dB only while he speaks (smooth envelope). SFX bus ~3dB under the music, above it only on hits.
At 4–6 big hits, layer reverse swell + boom (HP 60Hz) + transient + glitch. Light ticks and whooshes elsewhere — never on every cut.
Look
Fresh premium design system, written once in DESIGN.md:
one sans font, one accent colour
real UI captures as rounded cards (16–22px) with soft shadows
dark/light section switches
Never: outlines/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups.
Motion (non-negotiable — this is the feel)
Reference for feel: /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4; rules in /Users/evan/workspace/video-generator/videos/evals-hype/MOTION_FIX.md.
One camera = a pure function of time with velocity-continuous easing. Deterministic: no timers or randomness.
Two easing families only: snappy ≈25% of remaining distance per frame, floaty ≈12%.
Exits accelerate into the cut; the next shot carries the velocity through it.
Real whip-pans: both shots travel together over 0.3–0.45s with velocity blur.
Overlays are screen-locked.
Overlapping action, no start-stop chains.
Every hold drifts 1–3%. No dead stops, no frozen frames, no double cuts.
Word-by-word reveals 2–4 frames apart; the newest word flashes the accent.
Cuts on the beat ±1 frame.
Speed by re-rendering: render at --fps 30/speedup (e.g. --fps 95/4 = 1.263×), play back at 30fps.
Impact FX in post (ffmpeg) on 4–6 big hits only:
+8.5% zoom punch (τ 0.09s)
16px shake at ~22Hz (τ 0.1s)
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
A lighter version on secondary cuts.
QA (do it, then deliver)
python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat>: no dead stops, no frozen holds, cuts on the beat.
sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> at every transition — look at each strip.
VO RMS gap check (≤ 0.35s) plus a Whisper pass on the final mix.
Text readable and inside the margins; every name and number on screen is real.
Fix anything that fails, re-render, and finish with the single file videos/evals-demo3/renders/evals-demo3.mp4.EEvan Cha
Gemini 3.8 Flash/hyperframes 1 turn·27.5M tokens Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results. 
Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.EEvan Cha
GPT-6.1 Sol/hyperframes 1 turn·289.4K tokens Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.EEvan Cha1 turn·12.9M tokens Evals landscape demo-video prompt (approved "감도")
Copy everything below the line into a new session.
Make a ~60s landscape product demo video for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. Horizontal 16:9, 1920×1080, 30fps (for YouTube / website / X timeline). Not a vertical Short/Reel — no 9:16 output, no vertical crops, no phone safe-zones. Music-locked, no narration — on-screen type tells the story while the real product is demoed end to end. Build it in HyperFrames (HTML + one paused GSAP timeline) under /Users/evan/workspace/video-generator/videos/evals-demo3/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use the logged-in Claude in Chrome session if connected; otherwise public pages).
Motion reference (local file, already downloaded): /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its motion feel exactly (it is also 16:9); change only scenario, design and assets. Its 26s cut-down is just the reference for feel — this video runs the full ~60s.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
Run it autonomously — do not ask me for input
Work in Claude Code with the repo /Users/evan/workspace/video-generator as the working directory. Every path below is absolute; everything you need is already on this machine.
Do NOT stop to ask questions or request inputs. Every input is specified here; for anything not specified, pick a sensible default, write the decision into DECISIONS.md, and keep going until the final file is rendered.
If something is unavailable, use the fallback and continue:
Chrome not connected or a page needs sign-in → use the public pages (https://evals-app.vercel.app home, /compare, leaderboards, explore) with Playwright, and record outputs from their public "Open preview" links.
Can't run a new Battle → use public outputs that share the same prompt (the /compare pages show pairs).
Music: pick an unused track from /Users/evan/workspace/video-generator/videos/evals-hype/music3/ or music2/. SFX: /Users/evan/workspace/video-generator/videos/evals-hype/sfx2/, ref2/sfx3/, ref2/sfx4/.
QA scripts: /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py and strip.sh (also beats.py for BPM). If they fail, write equivalents.
Render: npx --yes hyperframes@0.8.101 render -o <out.mp4> --quiet from the project dir (--fps 95/4 etc. for speed re-render). API calls via curl (python has no SSL certs here).
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Study the local reference file frame by frame (ffmpeg frame dumps + frame-diff). Its design notes are in /Users/evan/workspace/video-generator/videos/evals-hype/ref2/SPEC.md and MOTION_FIX.md — read them, but design and copy must be new. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver one 1920×1080 file renders/evals-demo3-v1.mp4 (no vertical version) plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results. Evals landscape demo-video prompt (approved "감도")
Copy everything below the line into a new session.
Make a ~60s landscape product demo video for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. Horizontal 16:9, 1920×1080, 30fps (for YouTube / website / X timeline). Not a vertical Short/Reel — no 9:16 output, no vertical crops, no phone safe-zones. Music-locked, no narration — on-screen type tells the story while the real product is demoed end to end. Build it in HyperFrames (HTML + one paused GSAP timeline) under /Users/evan/workspace/video-generator/videos/evals-demo3/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use the logged-in Claude in Chrome session if connected; otherwise public pages).
Motion reference (local file, already downloaded): /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its motion feel exactly (it is also 16:9); change only scenario, design and assets. Its 26s cut-down is just the reference for feel — this video runs the full ~60s.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
Run it autonomously — do not ask me for input
Work in Claude Code with the repo /Users/evan/workspace/video-generator as the working directory. Every path below is absolute; everything you need is already on this machine.
Do NOT stop to ask questions or request inputs. Every input is specified here; for anything not specified, pick a sensible default, write the decision into DECISIONS.md, and keep going until the final file is rendered.
If something is unavailable, use the fallback and continue:
Chrome not connected or a page needs sign-in → use the public pages (https://evals-app.vercel.app home, /compare, leaderboards, explore) with Playwright, and record outputs from their public "Open preview" links.
Can't run a new Battle → use public outputs that share the same prompt (the /compare pages show pairs).
Music: pick an unused track from /Users/evan/workspace/video-generator/videos/evals-hype/music3/ or music2/. SFX: /Users/evan/workspace/video-generator/videos/evals-hype/sfx2/, ref2/sfx3/, ref2/sfx4/.
QA scripts: /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py and strip.sh (also beats.py for BPM). If they fail, write equivalents.
Render: npx --yes hyperframes@0.8.101 render -o <out.mp4> --quiet from the project dir (--fps 95/4 etc. for speed re-render). API calls via curl (python has no SSL certs here).
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Study the local reference file frame by frame (ffmpeg frame dumps + frame-diff). Its design notes are in /Users/evan/workspace/video-generator/videos/evals-hype/ref2/SPEC.md and MOTION_FIX.md — read them, but design and copy must be new. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver one 1920×1080 file renders/evals-demo3-v1.mp4 (no vertical version) plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.EEvan Cha
GPT-6.1 Sol/hyperframes 1 turn·4.4M tokens Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results. 
Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/.
0. Inputs
Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in).
Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets.
Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it."
What Evals really does (only show these — no invented features, numbers, rankings or names)
Ask: type what you want to make in the composer ("What would you like to make?").
Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct.
Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session).
Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B").
Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status.
Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups).
Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards.
Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures.
End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag.
1. Reverse-engineer the reference first (frames are the truth)
Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions:
every cut (frame number), shot order and durations, average shot length, cuts per 10s
camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame)
type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration
card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts
sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change.
2. Scenario
Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid.
3. Design
Write DESIGN.md once and make every section obey it (no per-section drift):
palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0)
card treatment: real UI captures, rounded 14–18px corners, soft layered shadow
Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features.
4. Assets (all real, all new)
Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills.
Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task.
Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films.
Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour.
New music track (detect BPM and first downbeat with a script) and a fresh SFX set.
5. Motion rules (this is the 감도 — non-negotiable)
One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone.
Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it.
Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched).
Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut.
Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal.
Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains.
Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops.
No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts.
Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks).
Every cut on the beat ±1 frame; visual hit and SFX on the same frame.
6. Speed, edit, impact
Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames.
Quantize every segment to 8th notes; round segment start frames up.
Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo):
zoom punch +8.5%, decaying τ = 0.09s
crop shake 16px at ~22Hz, decaying τ = 0.1s
rgbashift 16px for 2 frames, then 7px for 2 frames
1–2 frames of white flash
Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split).
Find hit frames from frame-diff peaks and by looking.
7. Sound
Music bed carries the film; no ducking.
SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps.
At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut.
Master at −14 LUFS, AAC 192k.
8. QA before you report (mandatory)
motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts.
strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs.
Check text readability, safe margins, and that every number/name on screen exists in the real product.
Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.EEvan Cha
GPT-6.1 Sol/hyperframes 1 turn·221.7K tokens ttamburins2 outputs I want the whole human story in one shot — from the first people to the AI we use now.
15 seconds, 16:9, no narration, no on-screen text until the very last frame.
One continuous move, never cutting. The camera pushes forward and the world changes
around it: firelight on a cave wall, a hand pressed in ochre, then fields and the first
walls, then a printing press, then steel and smoke, then a wall of city light, then a
room of humming machines, then the light of a screen on a face in the dark.
Each era arrives while the last one is still leaving — the fire becomes a furnace,
the cave painting becomes a page, the page becomes a screen. Hand things off on shape:
something round stays round, something bright stays in the same corner of the frame.
Keep a human in almost every era, always small in the frame, so the scale of the thing
reads against a person.
No famous faces, no logos, no company names, no invented alphabets on signage.
It should feel like one long exhale, not a slideshow.
The last frame can carry one line of English text, my choice later — leave it clean.
Give me the finished mp4, a cover frame, and the prompts you used per shot. I want the whole human story in one shot — from the first people to the AI we use now.
15 seconds, 16:9, no narration, no on-screen text until the very last frame.
One continuous move, never cutting. The camera pushes forward and the world changes
around it: firelight on a cave wall, a hand pressed in ochre, then fields and the first
walls, then a printing press, then steel and smoke, then a wall of city light, then a
room of humming machines, then the light of a screen on a face in the dark.
Each era arrives while the last one is still leaving — the fire becomes a furnace,
the cave painting becomes a page, the page becomes a screen. Hand things off on shape:
something round stays round, something bright stays in the same corner of the frame.
Keep a human in almost every era, always small in the frame, so the scale of the thing
reads against a person.
No famous faces, no logos, no company names, no invented alphabets on signage.
It should feel like one long exhale, not a slideshow.
The last frame can carry one line of English text, my choice later — leave it clean.
Give me the finished mp4, a cover frame, and the prompts you used per shot.ttamburins
GPT-6 AstraNo skill 1 turn·440.7K tokens I want the whole human story in one shot — from the first people to the AI we use now.
15 seconds, 16:9, no narration, no on-screen text until the very last frame.
One continuous move, never cutting. The camera pushes forward and the world changes
around it: firelight on a cave wall, a hand pressed in ochre, then fields and the first
walls, then a printing press, then steel and smoke, then a wall of city light, then a
room of humming machines, then the light of a screen on a face in the dark.
Each era arrives while the last one is still leaving — the fire becomes a furnace,
the cave painting becomes a page, the page becomes a screen. Hand things off on shape:
something round stays round, something bright stays in the same corner of the frame.
Keep a human in almost every era, always small in the frame, so the scale of the thing
reads against a person.
No famous faces, no logos, no company names, no invented alphabets on signage.
It should feel like one long exhale, not a slideshow.
The last frame can carry one line of English text, my choice later — leave it clean.
Give me the finished mp4, a cover frame, and the prompts you used per shot. · 2ttamburins
Claude Opus 5.5No skill 1 turn·2.8M tokens