Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind. ## Story 1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing." 2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI. 3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side. 4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A". 5. Run record of that same output: LLM, skill, environment, turn time, tokens. 6. Setup leaderboard: "Setups, not just models." 7. Community wall of real game / video / shorts outputs. 8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag. ## Content - Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages. - Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra). - No real-person likeness. ## Voice - One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps. - ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals". - The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates. ## Music & sound - Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe. - Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope. - SFX sit just under the music and rise above it only on hits. - 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut. ## Look - One sans font, one accent colour. - Real UI as rounded cards (16–22px radius) with soft shadows. - Sections alternate dark and light. - Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI. ## Motion: this is the feel - Premium, fast, confident: an Apple / Linear keynote cut at shorts speed. - One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%). - Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame. - Exits accelerate into the cut; the next shot carries that velocity through it. - Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout. - Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains. - Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s. - On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat. - Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts. - Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts. ## Quality bar Before calling it done, watch it like a motion director who didn't build it: - no dead stops, frozen frames or stutters - every cut on the beat - every word audible - every name and number real Fix the 5 worst moments first.
Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind.
## Story
1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing."
2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI.
3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side.
4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A".
5. Run record of that same output: LLM, skill, environment, turn time, tokens.
6. Setup leaderboard: "Setups, not just models."
7. Community wall of real game / video / shorts outputs.
8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag.
## Content
- Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages.
- Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra).
- No real-person likeness.
## Voice
- One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps.
- ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals".
- The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates.
## Music & sound
- Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe.
- Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope.
- SFX sit just under the music and rise above it only on hits.
- 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut.
## Look
- One sans font, one accent colour.
- Real UI as rounded cards (16–22px radius) with soft shadows.
- Sections alternate dark and light.
- Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI.
## Motion: this is the feel
- Premium, fast, confident: an Apple / Linear keynote cut at shorts speed.
- One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%).
- Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame.
- Exits accelerate into the cut; the next shot carries that velocity through it.
- Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout.
- Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains.
- Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s.
- On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat.
- Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts.
- Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts.
## Quality bar
Before calling it done, watch it like a motion director who didn't build it:
- no dead stops, frozen frames or stutters
- every cut on the beat
- every word audible
- every name and number real
Fix the 5 worst moments first.
Open one to inspect its creator, recorded recipe and run record.
Task outputs, votes and discussion
Questions and reactions stay attached to the task, across every output.
Discussion