VideoOutput 3 / 3EEvan ChaOct 8make a demo video about https://evals.ag/ In the video, all the reference example vid must be 3d game, 2d game, videos like hooking in 3 seconds. follow the demo video skill!LLMClaude Opus 5.5Gen AI—Not RecordedSkilldemo-video-creatorEnv—Not Recorded
WebsiteOutput 1 / 4EEvan ChaOct 8Build a playable 3D hover-bike racing game in the browser (Three.js). Neon cyberpunk city at night, rain, glowing billboards, bloom. Race through traffic with boost pads, ramps and near-miss sparks; speedometer, lap timer and a dramatic crash + replay screen. Make it look like a premium console game trailer the moment it loads. · 4LLMGrok 4.7Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 4 / 4EEvan ChaOct 8Build a playable 3D flying game in the browser (Three.js): pilot a paper plane through floating islands at golden hour. Soft volumetric clouds, waterfalls spilling off the islands, flocks of birds, glowing rings to fly through, wind trails, a cinematic intro fly-by and a score. Hand-painted storybook look, gorgeous from the first frame.LLMDeepSeek V4.1 FlashGen AI—Not RecordedSkillNo skillEnv—Not Recorded
VideoOutput 4 / 4EEvan ChaOct 8Create a 12-second cinematic 3D product reveal film in the browser (Three.js) for a fictional smartwatch called ORBIT. Dark studio, dramatic rim lighting, slow camera orbit and dolly, liquid-metal reflections, a particle burst on the reveal, bold kinetic typography, logo end card. Auto-plays and loops like an Apple keynote launch film.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 4 / 4EEvan ChaOct 8Build a playable 3D monster-truck smash game in the browser (Three.js + physics). Drive a giant monster truck through a colourful toy city and smash everything: towers of blocks topple, cars flip, glass shatters into debris with real physics, slow-motion on big crashes, dust and sparks, a destruction score. Should feel explosive and satisfying from the first second.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
Loading answer…DocumentOutput 3 / 3EEvan ChaOct 7Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind. ## Story 1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing." 2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI. 3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side. 4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A". 5. Run record of that same output: LLM, skill, environment, turn time, tokens. 6. Setup leaderboard: "Setups, not just models." 7. Community wall of real game / video / shorts outputs. 8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag. ## Content - Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages. - Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra). - No real-person likeness. ## Voice - One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps. - ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals". - The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates. ## Music & sound - Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe. - Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope. - SFX sit just under the music and rise above it only on hits. - 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut. ## Look - One sans font, one accent colour. - Real UI as rounded cards (16–22px radius) with soft shadows. - Sections alternate dark and light. - Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI. ## Motion: this is the feel - Premium, fast, confident: an Apple / Linear keynote cut at shorts speed. - One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%). - Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame. - Exits accelerate into the cut; the next shot carries that velocity through it. - Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout. - Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains. - Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s. - On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat. - Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts. - Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts. ## Quality bar Before calling it done, watch it like a motion director who didn't build it: - no dead stops, frozen frames or stutters - every cut on the beat - every word audible - every name and number real Fix the 5 worst moments first.LLMQwen3.8 Max (0902)Gen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoOutput 1 / 2EEvan ChaOct 7https://www.instagram.com/reel/Dd4VIgTq8e7/ reference this video! you must reference the asset of the video. music, sound effect etc... and what u need to describe is https://evals.ag/ site! · 2LLMGemini 3.8 FlashGen AI—Not RecordedSkillproduct-launch-videoEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 7Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind. ## Story 1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing." 2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI. 3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side. 4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A". 5. Run record of that same output: LLM, skill, environment, turn time, tokens. 6. Setup leaderboard: "Setups, not just models." 7. Community wall of real game / video / shorts outputs. 8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag. ## Content - Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages. - Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra). - No real-person likeness. ## Voice - One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps. - ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals". - The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates. ## Music & sound - Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe. - Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope. - SFX sit just under the music and rise above it only on hits. - 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut. ## Look - One sans font, one accent colour. - Real UI as rounded cards (16–22px radius) with soft shadows. - Sections alternate dark and light. - Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI. ## Motion: this is the feel - Premium, fast, confident: an Apple / Linear keynote cut at shorts speed. - One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%). - Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame. - Exits accelerate into the cut; the next shot carries that velocity through it. - Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout. - Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains. - Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s. - On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat. - Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts. - Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts. ## Quality bar Before calling it done, watch it like a motion director who didn't build it: - no dead stops, frozen frames or stutters - every cut on the beat - every word audible - every name and number real Fix the 5 worst moments first.LLMQwen3.8 Max (0902)Gen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoEEvan ChaOct 7Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind. ## Story 1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing." 2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI. 3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side. 4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A". 5. Run record of that same output: LLM, skill, environment, turn time, tokens. 6. Setup leaderboard: "Setups, not just models." 7. Community wall of real game / video / shorts outputs. 8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag. ## Content - Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages. - Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra). - No real-person likeness. ## Voice - One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps. - ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals". - The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates. ## Music & sound - Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe. - Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope. - SFX sit just under the music and rise above it only on hits. - 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut. ## Look - One sans font, one accent colour. - Real UI as rounded cards (16–22px radius) with soft shadows. - Sections alternate dark and light. - Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI. ## Motion: this is the feel - Premium, fast, confident: an Apple / Linear keynote cut at shorts speed. - One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%). - Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame. - Exits accelerate into the cut; the next shot carries that velocity through it. - Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout. - Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains. - Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s. - On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat. - Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts. - Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts. ## Quality bar Before calling it done, watch it like a motion director who didn't build it: - no dead stops, frozen frames or stutters - every cut on the beat - every word audible - every name and number real Fix the 5 worst moments first.LLMGPT-6.1 SolGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoEEvan ChaOct 7make a demo video about https://evals.ag/LLMGPT-6 AstraGen AI—Not RecordedSkilldemo-video-creatorEnv—Not Recorded
VideoOutput 3 / 4EEvan ChaOct 7reference this video for motion graphic demo https://x.com/samgrows/status/2107571800020222393?s=20 you must describe this site https://evals.ag/LLMClaude Opus 5.5Gen AI—Not RecordedSkillagent-browserEnv—Not Recorded
VideoEEvan ChaOct 7Make a ~60s 16:9 product demo video for Evals (evals.ag): battle AI setups on your own task and vote blind. ## Story 1. Hook (first 3s must be instantly clear): real AI-made game/video outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing." 2. A real game task is typed into the Evals composer: "Build a polished playable Sonic-style 3D platform game for the browser." → Battle → up to 4 setups (Model · Skill · Environment), each landing with real values and moving UI. 3. Send → velocity-matched zoom-through into the big hit → the 4 results play side by side. 4. Blind compare: two moving outputs, "Which would you choose?" → "I prefer A". 5. Run record of that same output: LLM, skill, environment, turn time, tokens. 6. Setup leaderboard: "Setups, not just models." 7. Community wall of real game / video / shorts outputs. 8. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag. ## Content - Show only games, videos, 3D scenes, motion pieces and shorts. No websites, no landing pages. - Real Evals UI only, never mocked up. Every name and number on screen is real; only model names that exist in Evals (e.g. Claude Opus 5.5, GPT-6 Astra). - No real-person likeness. ## Voice - One energetic, confident English male narrator: excited on the hook, curious on the question, emphatic on the tagline, quick through the UI steps. - ~170 words in one continuous take that fills the whole film. No pause longer than ~0.35s, no clipped words. Say "Evals" clearly as "E-vals". - The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates. ## Music & sound - Modern minimal tech-launch track, 120 BPM: punchy kick, deep 808, crisp hats, glassy synth plucks. Apple / Linear keynote feel. No cheesy EDM risers, no corporate stock vibe. - Voice always leads. Music dips ~6dB only while the voice speaks, on a smooth envelope. - SFX sit just under the music and rise above it only on hits. - 4–6 big hits get a layered reverse swell + sub boom + transient + glitch. Light ticks and whooshes elsewhere, never on every cut. ## Look - One sans font, one accent colour. - Real UI as rounded cards (16–22px radius) with soft shadows. - Sections alternate dark and light. - Never: outlines or hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI. ## Motion: this is the feel - Premium, fast, confident: an Apple / Linear keynote cut at shorts speed. - One continuous camera with velocity-continuous easing. Two easing personalities only: snappy (~25% of remaining distance per frame) and floaty (~12%). - Nothing ever stops dead. Every hold keeps drifting 1–3%. Fast moves decelerate over at least 6 frames. The end card keeps pushing until the last frame. - Exits accelerate into the cut; the next shot carries that velocity through it. - Real whip-pans: both shots travel together over 0.3–0.45s, motion blur peaking mid-move. One whip direction throughout. - Overlapping action: the next element starts while the previous one is ~60% done. No start-stop chains. - Text is screen-locked and arrives together with its container. Words reveal one by one, 2–4 frames apart, and the newest word flashes the accent. In ≤0.35s, readable ≥0.7s, out ≤0.25s. - On busy footage the camera stays calm (slow push ≤3%); the energy comes from cutting on the beat. - Every cut lands on the beat (±1 frame), with the visual hit and the SFX on the same frame. No double cuts. - Big hits (4–6 only): +8.5% zoom punch, ~16px shake at ~22Hz settling in 0.1s, RGB split 16px → 7px over 4 frames, 1–2 frames of white flash. A lighter version on secondary cuts. ## Quality bar Before calling it done, watch it like a motion director who didn't build it: - no dead stops, frozen frames or stutters - every cut on the beat - every word audible - every name and number real Fix the 5 worst moments first.LLMGPT-6.1 SolGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoEEvan ChaOct 7reference this video for motion graphic demo https://x.com/samgrows/status/2107571800020222393?s=20 you must describe this site https://evals.ag/LLMGPT-6.1 SolGen AI—Not RecordedSkillNo skillEnv—Not Recorded
VideoEEvan ChaOct 6could you make a demo video reference this site https://x.com/jordanarchivess/status/2107536262894915762?s=20. what you need to make it site is https://evals.ag. you must use tts with elevenlabs v4 apiLLMGPT-6.1 SolGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoEEvan ChaOct 6START NOW. Do not ask me any questions and do not wait for input — every input is below. Make every unspecified decision yourself, log it in DECISIONS.md, and keep going until ONE finished video file exists. The only deliverable is that one .mp4. Task Make ONE ~60s landscape demo video for Evals (evals.ag): battle AI setups on your own task and vote blind. Output: /Users/evan/workspace/video-generator/videos/evals-demo3/renders/evals-demo3.mp4, 1920×1080, 30fps, H.264 + AAC 192k, −14 LUFS. Nothing else to deliver (no vertical version, no variants). Work dir: repo /Users/evan/workspace/video-generator (Claude Code on this Mac). Build in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-demo3/. Render: npx --yes hyperframes@0.8.101 render -o <out> --quiet from the project dir. Use curl for HTTP/API calls (python has no SSL certs). ElevenLabs key: ELEVENLABS_API_KEY in /Users/evan/workspace/video-generator/.env. Story (English voiceover + on-screen type) Hook, first 3s must be instantly clear: real AI-made GAME/VIDEO outputs of the SAME prompt slam in side by side, moving, full-bleed → big type "Which one is better?" → rapid A/B/C/D flips on the beats → "Same prompt. Different setups." → "You can't tell by guessing." Type a real game/video task into the Evals composer → Battle → up to 4 setups (Model · Skill · Environment, each landing with real values and moving UI). Send → velocity-matched zoom-through into the big hit → the 4 results play side by side. Blind compare: two moving outputs, "Which would you choose?" → vote ("I prefer A"). Run record of that same output (real LLM, skill, environment, turn time, tokens). Setup leaderboard ("Setups, not just models."). Community wall of real game/video/shorts outputs. End card: Evals mark + wordmark, "Stop guessing. Battle it.", evals.ag. Content rules Show ONLY videos, games, 3D scenes, motion videos, shorts made in Evals. No websites / landing pages. Main thread = the real 4-setup Sonic battle: prompt "Build a polished playable Sonic-style 3D platform game for the browser.", clips /Users/evan/workspace/video-generator/videos/evals-launch/assets/clips/run_a.mp4 … run_d.mp4, run https://evals-app.vercel.app/runs/run_597Q2CRE4FXNJ2R7HZXMVF4Z03. Capture fresh real UI from https://evals-app.vercel.app with Playwright at 2× DPR: composer, Run-mode menu, setups stepper, compare page, output run record, Leaderboards, Explore → Outputs. If a page needs sign-in, use the public pages. Record extra community game/video outputs from their public "Open preview" links. Screencast 1600×900, 30fps, 6–7s, with keys pressed or orbit so they move. Reuse the pattern in /Users/evan/workspace/video-generator/videos/evals-launch/scripts/rec.mjs. No real-person likeness, no invented features/numbers/names. Only show model names you see in captures (e.g. Claude Opus 5.5, GPT-6 Astra). Voiceover (must flow continuously) ElevenLabs voice Will bIHbv24MWmeRgasZH58o, model eleven_v4, stability 0.15, similarity 0.8, style 0.8, speed 1.1, with emotion tags in the text ([excited] [confident] [curious] [emphatic] [quick]). Use the /with-timestamps endpoint. Write ~165–180 words so the voice fills the whole film. Generate it as ONE take and lay it unspliced — never cut it into phrases or time-stretch pieces. Longest silence inside the VO ≤ 0.35s. No clipped words. The edit follows the voice: each shot change lands on the beat nearest the start of the phrase it illustrates. Verify with Whisper that every word is there and no tags are spoken. If "Evals" sounds like "evils", respell it "E-vals". Music & sound New modern premium BGM, not cheesy. Generate 2–3 candidates with ElevenLabs Music: "modern minimal tech launch, punchy kick, deep 808, crisp hats, glassy synth plucks, Apple/Linear keynote feel, no cheesy EDM risers, no corporate/stock vibe, 120 BPM". Pick the cleanest; detect BPM and the first downbeat. VO leads at about −16 LUFS. Music ducks about 6dB only while he speaks (smooth envelope). SFX bus ~3dB under the music, above it only on hits. At 4–6 big hits, layer reverse swell + boom (HP 60Hz) + transient + glitch. Light ticks and whooshes elsewhere — never on every cut. Look Fresh premium design system, written once in DESIGN.md: one sans font, one accent colour real UI captures as rounded cards (16–22px) with soft shadows dark/light section switches Never: outlines/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups. Motion (non-negotiable — this is the feel) Reference for feel: /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4; rules in /Users/evan/workspace/video-generator/videos/evals-hype/MOTION_FIX.md. One camera = a pure function of time with velocity-continuous easing. Deterministic: no timers or randomness. Two easing families only: snappy ≈25% of remaining distance per frame, floaty ≈12%. Exits accelerate into the cut; the next shot carries the velocity through it. Real whip-pans: both shots travel together over 0.3–0.45s with velocity blur. Overlays are screen-locked. Overlapping action, no start-stop chains. Every hold drifts 1–3%. No dead stops, no frozen frames, no double cuts. Word-by-word reveals 2–4 frames apart; the newest word flashes the accent. Cuts on the beat ±1 frame. Speed by re-rendering: render at --fps 30/speedup (e.g. --fps 95/4 = 1.263×), play back at 30fps. Impact FX in post (ffmpeg) on 4–6 big hits only: +8.5% zoom punch (τ 0.09s) 16px shake at ~22Hz (τ 0.1s) rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash A lighter version on secondary cuts. QA (do it, then deliver) python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat>: no dead stops, no frozen holds, cuts on the beat. sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> at every transition — look at each strip. VO RMS gap check (≤ 0.35s) plus a Whisper pass on the final mix. Text readable and inside the margins; every name and number on screen is real. Fix anything that fails, re-render, and finish with the single file videos/evals-demo3/renders/evals-demo3.mp4.LLMGemini 3.8 FlashGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoOutput 1 / 2EEvan ChaOct 6Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/. 0. Inputs Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in). Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets. Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it." What Evals really does (only show these — no invented features, numbers, rankings or names) Ask: type what you want to make in the composer ("What would you like to make?"). Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct. Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session). Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B"). Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status. Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups). Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards. Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures. End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag. 1. Reverse-engineer the reference first (frames are the truth) Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions: every cut (frame number), shot order and durations, average shot length, cuts per 10s camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame) type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change. 2. Scenario Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid. 3. Design Write DESIGN.md once and make every section obey it (no per-section drift): palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0) card treatment: real UI captures, rounded 14–18px corners, soft layered shadow Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features. 4. Assets (all real, all new) Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills. Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task. Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films. Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour. New music track (detect BPM and first downbeat with a script) and a fresh SFX set. 5. Motion rules (this is the 감도 — non-negotiable) One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone. Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it. Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched). Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut. Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal. Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains. Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops. No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts. Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks). Every cut on the beat ±1 frame; visual hit and SFX on the same frame. 6. Speed, edit, impact Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames. Quantize every segment to 8th notes; round segment start frames up. Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo): zoom punch +8.5%, decaying τ = 0.09s crop shake 16px at ~22Hz, decaying τ = 0.1s rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split). Find hit frames from frame-diff peaks and by looking. 7. Sound Music bed carries the film; no ducking. SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps. At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut. Master at −14 LUFS, AAC 192k. 8. QA before you report (mandatory) motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts. strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs. Check text readability, safe margins, and that every number/name on screen exists in the real product. Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.LLMGrok 4.7Gen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 6Make a viral demo video reference https://x.com/titouangillet_/status/2107016017259921464 You must reference this site https://evals.ag/. You must use font sanserif. Motion must be speedy. Never make ai slop.LLMClaude Opus 5.5Gen AIseedance_2_5SkillhyperframesEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 6Make a viral demo video reference https://x.com/titouangillet_/status/2107016017259921464 You must reference this site https://evals.ag/. You must use font sanserif. Motion must be speedy. Never make ai slop.LLMClaude Opus 5.5Gen AIseedance_2_5SkillhyperframesEnv—Not Recorded
VideoEEvan ChaOct 6Evals landscape demo-video prompt (approved "감도") Copy everything below the line into a new session. Make a ~60s landscape product demo video for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. Horizontal 16:9, 1920×1080, 30fps (for YouTube / website / X timeline). Not a vertical Short/Reel — no 9:16 output, no vertical crops, no phone safe-zones. Music-locked, no narration — on-screen type tells the story while the real product is demoed end to end. Build it in HyperFrames (HTML + one paused GSAP timeline) under /Users/evan/workspace/video-generator/videos/evals-demo3/. 0. Inputs Product: https://evals.ag (app: https://evals-app.vercel.app — use the logged-in Claude in Chrome session if connected; otherwise public pages). Motion reference (local file, already downloaded): /Users/evan/workspace/video-generator/videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its motion feel exactly (it is also 16:9); change only scenario, design and assets. Its 26s cut-down is just the reference for feel — this video runs the full ~60s. Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it." Run it autonomously — do not ask me for input Work in Claude Code with the repo /Users/evan/workspace/video-generator as the working directory. Every path below is absolute; everything you need is already on this machine. Do NOT stop to ask questions or request inputs. Every input is specified here; for anything not specified, pick a sensible default, write the decision into DECISIONS.md, and keep going until the final file is rendered. If something is unavailable, use the fallback and continue: Chrome not connected or a page needs sign-in → use the public pages (https://evals-app.vercel.app home, /compare, leaderboards, explore) with Playwright, and record outputs from their public "Open preview" links. Can't run a new Battle → use public outputs that share the same prompt (the /compare pages show pairs). Music: pick an unused track from /Users/evan/workspace/video-generator/videos/evals-hype/music3/ or music2/. SFX: /Users/evan/workspace/video-generator/videos/evals-hype/sfx2/, ref2/sfx3/, ref2/sfx4/. QA scripts: /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py and strip.sh (also beats.py for BPM). If they fail, write equivalents. Render: npx --yes hyperframes@0.8.101 render -o <out.mp4> --quiet from the project dir (--fps 95/4 etc. for speed re-render). API calls via curl (python has no SSL certs here). What Evals really does (only show these — no invented features, numbers, rankings or names) Ask: type what you want to make in the composer ("What would you like to make?"). Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct. Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session). Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B"). Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status. Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups). Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards. Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures. End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag. 1. Reverse-engineer the reference first (frames are the truth) Study the local reference file frame by frame (ffmpeg frame dumps + frame-diff). Its design notes are in /Users/evan/workspace/video-generator/videos/evals-hype/ref2/SPEC.md and MOTION_FIX.md — read them, but design and copy must be new. Write SPEC.md with measured numbers, not impressions: every cut (frame number), shot order and durations, average shot length, cuts per 10s camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame) type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change. 2. Scenario Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid. 3. Design Write DESIGN.md once and make every section obey it (no per-section drift): palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0) card treatment: real UI captures, rounded 14–18px corners, soft layered shadow Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features. 4. Assets (all real, all new) Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills. Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task. Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films. Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour. New music track (detect BPM and first downbeat with a script) and a fresh SFX set. 5. Motion rules (this is the 감도 — non-negotiable) One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone. Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it. Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched). Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut. Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal. Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains. Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops. No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts. Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks). Every cut on the beat ±1 frame; visual hit and SFX on the same frame. 6. Speed, edit, impact Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames. Quantize every segment to 8th notes; round segment start frames up. Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo): zoom punch +8.5%, decaying τ = 0.09s crop shake 16px at ~22Hz, decaying τ = 0.1s rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split). Find hit frames from frame-diff peaks and by looking. 7. Sound Music bed carries the film; no ducking. SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps. At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut. Master at −14 LUFS, AAC 192k. 8. QA before you report (mandatory) python3 /Users/evan/workspace/video-generator/videos/evals-hype/scripts/motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts. sh /Users/evan/workspace/video-generator/videos/evals-hype/scripts/strip.sh <render> <t> <out.png> (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs. Check text readability, safe margins, and that every number/name on screen exists in the real product. Deliver one 1920×1080 file renders/evals-demo3-v1.mp4 (no vertical version) plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.LLMGPT-6.1 SolGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
OtherEEvan ChaOct 6Make a ~60s product launch film for Evals (evals.ag) — the place to battle AI setups on your own task and vote blind. 16:9, 1920×1080, 30fps, music-locked, no narration — on-screen type tells the story. Build it in HyperFrames (HTML + one paused GSAP timeline) under videos/evals-<name>/. 0. Inputs Product: https://evals.ag (app: https://evals-app.vercel.app — use my logged-in Chrome session for anything behind sign-in). Motion reference: videos/evals-hype/renders/evals-ref2-short-fast2-impact.mp4 (approved). Match its feel exactly; change only scenario, design and assets. Core message: "Same prompt, different setups — battle them and vote blind. Stop guessing. Battle it." What Evals really does (only show these — no invented features, numbers, rankings or names) Ask: type what you want to make in the composer ("What would you like to make?"). Run modes: Battle (blind), Side-by-side comparison (pick models and skills), Direct. Battle: up to 4 setups. A setup = model + skill + environment (e.g. Claude Opus 5.5 / GPT-6 Astra, skills like hyperframes / frontend-design / web-design-guidelines, environments like Native Claude / Recorded session). Compare, blind: same prompt, outputs side by side as playable live previews, names hidden until you vote ("I prefer A / B"). Inspect: every output keeps its full run record — LLM, gen AI, skill, software, environment, services contacted, turn time, turn tokens, status. Rank: votes roll up into leaderboards of setups, not just models ("Setups, not just models." — same model, different setups). Explore: community Outputs, Tasks, Skills, Compare, Community, Leaderboards. Real model names visible in the product: Claude Opus 5.5, GPT-6 Astra. Show only names/numbers you actually see in your captures. End card: Evals mark + wordmark, "Stop guessing. Battle it.", URL evals.ag. 1. Reverse-engineer the reference first (frames are the truth) Download the reference and study it frame by frame. Write SPEC.md with measured numbers, not impressions: every cut (frame number), shot order and durations, average shot length, cuts per 10s camera moves per shot (start/end position, scale, curve), easing measured per frame (fraction of remaining distance covered each frame) type-on speed (frames per char), caret behaviour, word-by-word reveal spacing, new-word accent flash duration card/UI choreography (fly-in paths, stacking, depth), light↔dark section switches, split-screen layouts sound map: where every hit/whoosh/click lands relative to the cut Then rebuild shot for shot at the same rhythm. Only brand, copy, design and assets change. 2. Scenario Write SCRIPT.md before building: a fresh narrative arc true to what Evals really does (hook: a new SOTA model every week — which one is actually best at your work? → type a task → Battle up to 4 setups → blind side-by-side, vote → full run record → setup leaderboards → "Stop guessing. Battle it." evals.ag). Don't reuse the previous films' copy. Short punchy lines (2–5 words per beat), English. Every beat gets a timestamp on the music grid. 3. Design Write DESIGN.md once and make every section obey it (no per-section drift): palette (one accent per frame), type pairing (one sans for headlines; weights/sizes in px; tracking −0.02em, line-height 1.0) card treatment: real UI captures, rounded 14–18px corners, soft layered shadow Never: outline/hairline boxes, rules, underlines, highlight boxes behind words, cheap gradients, emoji, fake UI mockups, invented numbers/rankings/features. 4. Assets (all real, all new) Capture the real Evals UI with Playwright at 2–3× DPR: home composer (empty / typed), Run-mode menu (Battle / Side-by-side / Direct), setups stepper (1–4), a live run page (4 agents generating → "Your options are ready"), Compare (blind A/B with "I prefer A/B"), output run record panel, Leaderboards (setup rows), Explore Outputs/Tasks/Skills. Run one real 4-setup Battle on a visual prompt (e.g. a playable browser game) so the four results are genuinely the same task. Record real Evals outputs as clips from their live previews (screencast, 1600×900, 30fps, 6–7s each; press start/arrow keys or orbit so they move). Use outputs not used in earlier Evals films. Typing scenes: put a text layer exactly over the real input field, matching its font, size and colour. New music track (detect BPM and first downbeat with a script) and a fresh SFX set. 5. Motion rules (this is the 감도 — non-negotiable) One camera = one pure function of time cam(t) built from keyframes with velocity-continuous easing; applied from the master timeline. Deterministic: no timers, no randomness, any frame renders alone. Two easing families only: snappy ≈25% of remaining distance per frame (90% in 8–9 frames) and floaty ≈12% (18–20 frames). No overshoot unless the reference has it. Exits accelerate into the cut; the next shot carries the motion through it (velocity-matched). Real whip-pans: outgoing and incoming shots travel together over 0.30–0.45s, with velocity-driven motion blur that peaks mid-move. Never blur switched on over a hard cut. Overlays are screen-locked (never inherit the camera whip). Text arrives with its container in one mask reveal. Overlapping action: the next element starts at ~60% of the previous tween. No start-stop chains. Nothing freezes: every hold carries a 1–3% slow drift (sine.inOut across the whole hold). Fast moves decelerate over ≥6 frames — no dead stops. No double cuts (two visual changes on adjacent frames). Busy footage gets only a slow push; energy comes from the cuts. Word reveals 2–4 frames apart; the newest word flashes the accent, then settles 3–12 frames later. Hero type-on 1 char / 2 frames; in-product typing 0.6–0.9 char/frame; caret solid (never blinks). Every cut on the beat ±1 frame; visual hit and SFX on the same frame. 6. Speed, edit, impact Speed by re-rendering, never frame-dropping: render at --fps 30/speedup and play back at 30fps (e.g. --fps 95/4 → 1.263×). Pick the speed-up so one beat = a whole number of frames. Quantize every segment to 8th notes; round segment start frames up. Impact FX in post (ffmpeg) on 4–6 big hits only (cut to dark, reveal word, flash, key word, logo): zoom punch +8.5%, decaying τ = 0.09s crop shake 16px at ~22Hz, decaying τ = 0.1s rgbashift 16px for 2 frames, then 7px for 2 frames 1–2 frames of white flash Lighter version on secondary cuts (3.5% punch, 6px shake, 1 frame of split). Find hit frames from frame-diff peaks and by looking. 7. Sound Music bed carries the film; no ducking. SFX bus ~3dB under the music overall, above it only on hits and in the music's own gaps. At each big hit, layer reverse swell + boom (high-passed at 60Hz) + short transient + glitch synced to the RGB split. Use a modest total (~40–45 sounds); never spread SFX evenly on every cut. Master at −14 LUFS, AAC 192k. 8. QA before you report (mandatory) motion.py <render> <bpm> <first_beat> must show: no dead stops, no frozen holds, every cut within ±1.5 frames of a beat, no adjacent-frame cuts. strip.sh (16 consecutive frames) at every transition, and look at each one. A good whip shows a smooth velocity ramp with both shots visible mid-move. Contact sheets alone hide motion bugs. Check text readability, safe margins, and that every number/name on screen exists in the real product. Deliver renders/evals-<name>-v1.mp4 plus a report: script, design summary, asset list, shot list with times, BPM and speed factor, hit frames with FX, SFX map, QA results.LLMGPT-6.1 SolGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 6Make a viral demo video reference https://x.com/titouangillet_/status/2107016017259921464 You must reference this site https://evals.ag/. You must use font sanserif. Motion must be speedy. Never make ai slop.LLMClaude Opus 5.5Gen AIminimax_h3_maxSkillhyperframesEnv—Not Recorded
VideoOutput 2 / 2ttamburinsOct 6Make me a 15-second trailer for an F1 movie starring Ariana GrandeLLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2ttamburinsOct 6내가 지금 레퍼런스로 넣은 이미지들의 무드와 비슷하게 cozy한느낌의 브랜드 필름을 만들어줘LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
VideoEEvan ChaOct 6make a viral video in tiktok trending!LLMGPT-6.1 SolGen AI—Not RecordedSkillai-video-generationEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 6make a viral demo video reference https://x.com/titouangillet_/status/2107016017259921464?s=20 what you make it is https://evals.ag/ siteLLMClaude Opus 5.5Gen AI—Not RecordedSkillhyperframesEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 6Build a single-page landing site for a coffee subscription startup with a hero, pricing table and signup form.LLMGPT-6.1 SolGen AI—Not RecordedSkilldesign-taste-frontendEnv—Not Recorded
VideoOutput 1 / 2EEvan ChaOct 6make a demo intro video reference this video https://x.com/rajsinghfirst/status/2107161801322549725?s=20 demo site is https://evals.ag ! this siteLLMGPT-6 AstraGen AI—Not RecordedSkillhyperframesEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 6make a viral video reference this thing https://x.com/motion_conquest/status/2106786068712558848?s=20 what you need to make it is https://evals.ag/ site! demo videLLMClaude Opus 5.5Gen AI—Not RecordedSkillai-video-generationEnv—Not Recorded
VideoOutput 1 / 2EEvan ChaOct 6make a demo intro video reference this video https://x.com/rajsinghfirst/status/2107161801322549725?s=20 demo site is https://evals.ag ! this siteLLMClaude Opus 5.5Gen AI—Not RecordedSkillhyperframesEnv—Not Recorded
ImageOutput 2 / 2EEvan ChaOct 6make a ai influencer for x viralLLMClaude Opus 5.5Gen AI—Not RecordedSkillai-avatar-videoEnv—Not Recorded
VideoOutput 2 / 2EEvan ChaOct 6make a viral video in xLLMGPT-6 AstraGen AI—Not RecordedSkillai-video-generationEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 6Build a landing page for a neighborhood bakery: menu, opening hours, map and a reservation button. It should look great on phones.LLMClaude Opus 5.5Gen AI—Not RecordedSkilldesign-taste-frontendEnv—Not Recorded
WebsiteEEvan ChaOct 4make a 2d kawai samurai game.LLMGPT-6 AstraGen AIautospriteSkillNo skillEnv—Not Recorded
WebsiteOutput 4 / 4EEvan ChaOct 1Build a polished playable Sonic-style 3D platform game for the browser. Include a fast character, rings to collect, ramps, jumping, hazards, checkpoints, score and restart. Make speed and platforming feel satisfying. Deliver a working browser artifact.LLMClaude Opus 5.5Gen AI—Not RecordedSkillthreejs-scene-setupEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Build a polished playable Sonic-style 3D platform game for the browser. Include a fast character, rings to collect, ramps, jumping, hazards, checkpoints, score and restart. Make speed and platforming feel satisfying. Deliver a working browser artifact.LLMClaude Opus 5.5Gen AI—Not RecordedSkillplatformerEnv—Not Recorded
WebsiteOutput 1 / 2KKeith RyuOct 11+1?LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
VideoOutput 2 / 2ttamburinsOct 1Make me a 15-second ad with MapleStory characters built in 3DLLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2ttamburinsOct 1Make MapleStory as a 3D game I can play in the browserLLMClaude Opus 5.5Gen AI—Not RecordedSkillthreejs-scene-setupEnv—Not Recorded
VideoOutput 1 / 2ttamburinsOct 1I want the whole human story in one shot — from the first people to the AI we use now. 15 seconds, 16:9, no narration, no on-screen text until the very last frame. One continuous move, never cutting. The camera pushes forward and the world changes around it: firelight on a cave wall, a hand pressed in ochre, then fields and the first walls, then a printing press, then steel and smoke, then a wall of city light, then a room of humming machines, then the light of a screen on a face in the dark. Each era arrives while the last one is still leaving — the fire becomes a furnace, the cave painting becomes a page, the page becomes a screen. Hand things off on shape: something round stays round, something bright stays in the same corner of the frame. Keep a human in almost every era, always small in the frame, so the scale of the thing reads against a person. No famous faces, no logos, no company names, no invented alphabets on signage. It should feel like one long exhale, not a slideshow. The last frame can carry one line of English text, my choice later — leave it clean. Give me the finished mp4, a cover frame, and the prompts you used per shot. · 2LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
WebsiteEEvan ChaOct 1https://x.com/notdwd/status/2105373513104392552?s=20 reference this video for making production of https://evals-app.vercel.app/ site. launching demo video!LLMClaude Opus 5.5Gen AI—Not RecordedSkillhyperframesEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create the Eiffel Tower and its surroundings in Paris as a detailed interactive 3D browser scene. Show recognizable lattice structure, all platforms, Champ de Mars paths and greenery, neighboring Parisian buildings, atmospheric lighting and orbit controls. Deliver a working browser artifact. Source: https://x.com/EnvolDev/status/2103535619603567054 (66K views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteEEvan ChaOct 1Build the highest-quality playable PUBG-style battle royale prototype you can in the browser. Include an original island map, movement, aiming, collectible equipment, simple computer opponents, a shrinking safe zone, health, win/lose and restart. Deliver a working game artifact with clear controls. Source: https://x.com/emmanuel_2m/status/2102717178223235228 (58K views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create a detailed visual 3D presentation of an AH-64E Apache Guardian in Indian Air Force livery. Emphasize its recognizable silhouette, cockpit, rotor geometry, materials and dramatic lighting. Add orbit controls and rotor animation in a working browser artifact. This is a visual modeling exercise. Source: https://x.com/srikanthvaluri/status/2100935811059593547 (39K views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create a 20-second concept launch video for ChatGPT-6 Sol and Luna. Use only these model names as factual copy; invent no performance, pricing or availability claims. Build a polished motion-design treatment with expressive typography, contrasting visual identities, elegant transitions and a final two-model lockup. Label it an unofficial concept. Deliver a playable animation and MP4 if available. Source: https://x.com/israelfemiojo/status/2102642002613481939 (196K views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Build a polished playable Sonic-style 3D platform game for the browser. Include a fast character, rings to collect, ramps, jumping, hazards, checkpoints, score and restart. Make speed and platforming feel satisfying. Deliver a working browser artifact. Adapt the source Godot task to the browser. Source: https://x.com/AiBattle_/status/2095994051354919049 (1.5M views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteEEvan ChaOct 1Build a Night train. Create an immersive animated browser scene with a moving passenger train at night, lit windows, passing scenery, convincing wheels and tracks, and atmospheric light. Include a cinematic camera and pause/replay controls. Deliver a working browser artifact. Source: https://x.com/EnvolDev/status/2103282586055213355 (111K views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Build an interactive educational 3D jet engine in the browser. Include cutaway and x-ray views, animated airflow and rotating fan/compressor/turbine stages, clear labels, an exploded view and orbit controls. Explain the airflow visually and mark simplifications. Deliver the working browser artifact. Source: https://x.com/YouWareAI/status/2101271938706685991 (80K views observed 2026-10-01). Adapted recreation brief based on the public post, not a verbatim original prompt or replication of its model versions. Use the models selected in Evals.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteEEvan ChaOct 1Build a detailed interactive 3D Roman Colosseum in the browser. Include the elliptical amphitheater, stacked arcades, recognizable ruined outer wall, arena floor and surrounding terrain. Add atmospheric lighting, orbit/zoom camera controls and an automatic cinematic tour. Deliver the working browser artifact. Source: https://x.com/LuminaBench/status/2104995514433564800 (65K views, 433 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteEEvan ChaOct 1Create a bluetooth smart ring flex circuit using tscircuit. Deliver the tscircuit source, a rendered circuit/PCB preview and a concise component list. Explain any assumptions and validation limits; distinguish a concept design from manufacturing-ready hardware. Source: https://x.com/seveibar/status/2105043390576590978 (29.5K views, 153 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create a playable 3D high-speed ride through a cyberpunk city in the browser. Include neon architecture, a convincing sense of speed, a controllable vehicle, obstacles, collision feedback, a score and restart. Support keyboard and visible touch controls, and deliver the working browser artifact. Source: https://x.com/higgsfield_ai/status/2105087030464245944 (37K views, 247 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteEEvan ChaOct 1Build a beautiful interactive 3D voxel pagoda scene in the browser, with tiered curved roofs, red columns, detailed wooden structure, stone paths, trees and water in a coherent miniature landscape. Use voxel geometry, atmospheric lighting, and orbit/zoom controls. Deliver the working browser artifact. Source: https://x.com/LuminaBench/status/2104598576160456967 (40K views, 364 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Build a playable futuristic anti-gravity racing game inspired by the F-Zero-style racing demo in the source post. Create an original hovering vehicle, a twisting elevated track, boost, steering, lap timing and restart with a strong sense of speed. Deliver a browser-playable artifact. This adaptation targets the browser instead of the original Unreal environment. Source: https://x.com/ForwardEditor/status/2104633533981491575 (9.3K views, 128 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteEEvan ChaOct 1Build a playful 3D pool party game with a colorful swimming pool, inflatable floats, animated water, and a controllable character that can move and jump into the pool with splash feedback. Add a simple collectible challenge, score and reset. Deliver a browser-playable artifact. This adaptation targets the browser instead of the original Blender/Godot workflow. Source: https://x.com/bijanbowen/status/2105113620891726041 (31.6K views, 354 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create an interactive 3D gummy dice experience in the browser: colorful translucent dice with rounded edges, correctly arranged pips, glossy highlights, and soft stretchy jelly deformation as they drop, bounce and settle. Add a Roll button and appealing studio lighting. Deliver the working browser artifact. Source: https://x.com/vib3coded/status/2104828270789492913 (30K views, 559 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create an interactive fall foliage simulator as a polished browser experience. Show a grove of trees whose leaves transition from green through gold, orange and red, then fall with believable wind and accumulate on the ground. Include season progression and wind controls, pause/reset, and a beautiful coherent composition. Deliver a working HTML artifact, not just an explanation. Source inspiration: https://x.com/claudeai/status/2104674987164782598 (Sonnet 5 vs Sonnet 5.5 fall foliage simulator; observed 687K views, 7.8K likes on 2026-10-01). This is an adapted reproduction brief based on the public post description, not the original verbatim prompt. The current comparison uses the models selected in Evals; do not claim to reproduce the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create a visually engaging animated history of dinosaurs in under one minute using JavaScript. Cover their emergence in the Triassic, the Jurassic giants, Cretaceous diversity, and the end-Cretaceous extinction, with clear English labels, a timeline, recognizable moving dinosaurs and smooth scene transitions. Deliver a playable HTML artifact with play/pause and replay controls. Source: https://x.com/claudeai/status/2104674993212702803 (273K views, 684 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
WebsiteOutput 2 / 2EEvan ChaOct 1Create a cinematic luxury car rental website with a striking premium hero, beautifully presented vehicle collection, subtle motion, refined typography, and a functional booking interaction with dates, vehicle choice and a clear price summary. Make the layout responsive and accessible. Use fictional branding and demo pricing. Deliver a working browser artifact. Source: https://x.com/viktoroddy/status/2105226806492283286 (25K views, 404 likes observed 2026-10-01). Adapted reproduction brief based on the public post; not the original verbatim prompt. Compare the current Evals-selected models, not the original model versions.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnv—Not Recorded
VideoOutput 1 / 2EEvan ChaSep 30https://www.skills.ag/skills/skill_01KRJP787GB1PGW20M61REAJ69 이 사이트 보고 제품 설명 줬잖아 인스타 릴스 광고 만들어줘LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 1 / 2EEvan ChaSep 30make a advertisment which can upload to instgram reels.LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2ttamburinsSep 30Make me a 15-second trailer for an F1 movie starring Ariana GrandeLLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2ttamburinsSep 30내가 지금 레퍼런스로 넣은 이미지들의 무드와 비슷하게 cozy한느낌의 브랜드 필름을 만들어줘LLMGPT-6 AstraGen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2ttamburinsSep 30Please create a video showing the process of making a Rolex watch by an artisan,LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2ttamburinsSep 30make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it's your showreel for a résumé. go all out.LLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2ttamburinsSep 30https://www.youtube.com/watch?v=o7DnBq7Fn74 The Powerpuff Girls x NewJeans music video starts at 1:20 in this video, and it combines the production process and graphics? Taking that feeling into account, Ain't In LA Song • ADÉLA, make a music video of the highlight part of the song for about 15 secondsLLMClaude Opus 5.5Gen AI—Not RecordedSkillNo skillEnvHiggsfield workflow
VideoOutput 2 / 2EEvan ChaSep 30Make a 10-second cinematic promo video for Jeju Island: coastline, volcanic landscape and tangerine farms, ending on the text "Jeju. Slow down."LLMGPT-6 AstraGen AISeedance 2.5 (fal)Skillgeneral-videoEnvHiggsfield workflow
WebsiteOutput 2 / 2EEvan ChaSep 30Build an interactive personal finance dashboard web app: monthly spending by category, a budget progress view, recent transactions, and a savings goal tracker, using realistic example data.LLMGPT-6 AstraGen AI—Not RecordedSkillfrontend-designEnvBrowser + code workspace
VideoOutput 2 / 2EEvan ChaSep 30Make a 15-second animated announcement video for a coding bootcamp called "Launchpad" opening enrollment for its spring cohort, with the dates and a call to action.LLMGPT-6 AstraGen AI—Not RecordedSkillgeneral-videoEnvHiggsfield workflow
WebsiteOutput 2 / 2EEvan ChaSep 30Create a generative artwork inspired by a city at night seen from above: grids of streets, clusters of light, and traffic flow. Deliver it as an interactive piece that renders in the browser.LLMGPT-6 AstraGen AI—Not RecordedSkillalgorithmic-artEnvBrowser + code workspace
VideoOutput 2 / 2ttamburinsSep 30Make a 12-second animated video of a cute Norwegian Forest cat making a pretty strawberry drink at a caféLLMClaude Opus 5.5Gen AI—Not RecordedSkillfind-skillsEnvHiggsfield workflow
WebsiteOutput 2 / 2EEvan ChaSep 30Build the homepage for "Pixel Harbor", a three-person indie game studio. Show their upcoming game with a trailer placeholder, past titles, a devlog section, and a newsletter signup. It should feel like the studio has a strong personality.LLMGPT-6 AstraGen AI—Not RecordedSkillfrontend-designEnvBrowser + code workspace
VideoOutput 2 / 2EEvan ChaSep 30Make a 15-second product teaser video for "Aura Buds", new wireless earbuds with 40-hour battery life and adaptive noise cancelling. End on the product name and "Available October 15".LLMGPT-6 AstraGen AI—Not RecordedSkillhyperframesEnvHiggsfield workflow
VideoOutput 2 / 2EEvan ChaSep 30Make a 30-second explainer video on why the sky is blue, for curious 12-year-olds, with clear visuals and on-screen captions.LLMGPT-6 AstraGen AI—Not RecordedSkillfaceless-explainerEnvBrowser + code workspace