This article as a 2 minute explainer, made with the setup it describes.

The best way I've found to explain an idea isn't another article. It's a 90-second video: a narrator walking you through it while the diagram builds itself on screen. People who would skim a 2,000-word doc will watch that to the end, and they come away actually understanding it.

This has been a real aha moment for me, because I've made plenty of videos the old way: in After Effects for BARUCHealth, and editing video back in my days at National Grid. A decent explainer meant keyframing every box and arrow, cutting a timeline, recording voiceover, re-recording every flubbed line, then redoing half of it when the plan changed. Hours of production for a minute and a half of video.

That's gone, and it still feels a little incredible. Describe the video to Claude Code or Codex and it builds it: the animation, the timing, the edit, the captions and the narration. Product tours come out polished in a fraction of the time they used to take. And AI narration has gotten good enough that it doesn't feel weird to watch. It sounds like someone explaining, not a machine reading.

The video at the top is this article, made with the setup it describes. Below are more examples, then how the setup works, with the starter code so you can build your own on Linux, Mac, or Windows.

A few I've made

Another explainer, then product tours and trailers for apps I've built. In the tours, the footage is always the real app, not a mockup. Everything around it, the cuts, camera moves, captions, music and narration, came from a description of what the video should show.

How it works: a video is just a web page

Each video is an HTML page with one function, draw(t), that can paint any moment of the video on demand. A small script opens the page in headless Chrome, asks for frame 0, frame 1, frame 2, all the way to the end, and pipes every frame to ffmpeg next to a narration track that was mixed once, up front.

The slide video pipeline: the script is voiced once, then every frame is drawn exactly and joined with the pre-mixed audioscript.mdone line per scenenarrate.pyKokoro, local voicenarration.jstext, duration, audioindex.htmldraw(t) + score()render.mjsheadless ChromeAudio trackmixed offline, onceFramesdraw(f / 30), each oneffmpegH.264 + AACvideo.mp44 to 10 MB
The narration is voiced once per line. The browser then draws each frame on request, so render speed never affects timing.

Screen recording looks like the obvious approach, and it breaks in an unexpected way. Real-time capture drifts whenever the machine is busy: one 4-minute video came out 30 seconds longer than its own audio. Rendering frame by frame can't drift, because every frame is drawn at its exact timestamp no matter how slow the machine is. It's also faster than real time; a 3.5-minute video renders in under 2 minutes.

Two other choices are what make editing cheap:

  • The narration sets the pace. Each scene lasts a short lead-in, its line of narration, and a short tail. Rewrite a line and only that scene stretches or shrinks. Nothing else needs fixing by hand.
  • Animation follows the words. A helper, at(scene, 'phrase'), returns the moment the narrator reaches a phrase. A box can appear exactly as its name is spoken, and it stays in sync when the sentence changes.

Two kinds of video

There are two kinds, and the choice depends on whether there's a real product to show.

Slide explainerApp film
What's on screenAnimated diagrams and text, one scene per line of narrationA real website or prototype, driven by a script
What builds itOne HTML canvas page and two small scriptsfilm-kit, a small wrapper around Playwright
Best forStrategy, plans, architectureProduct walkthroughs and demos
Usual length1 to 4 minutes1 to 2 minutes

If you're new to this, start with a slide explainer. Product tours build on the same ideas, and they get their own section below.

What you need

Five tools, all free, plus a voice model, and rclone if you want uploads handled for you. Expect about 15 minutes and roughly 1 GB of downloads.

ToolWhy
Node 22 or newerRuns the render script (Node 22 ships a built-in WebSocket client)
Python 3.12 and uvRuns Kokoro, the voice model, in its own environment
ffmpegEncodes the final MP4
Google ChromeDraws the frames, headless
espeak-ng (recommended)Helps Kokoro pronounce words it doesn't know
rclone (optional)Uploads finished videos to Google Drive from the command line

Linux (Fedora shown; swap in apt or pacman). Install the CPU build of PyTorch first. The default Linux build drags in several GB of GPU libraries that a model this small doesn't need.

sudo dnf install nodejs ffmpeg espeak-ng google-chrome-stable
uv venv --python 3.12 ~/.venvs/kokoro
uv pip install --python ~/.venvs/kokoro/bin/python torch --index-url https://download.pytorch.org/whl/cpu
uv pip install --python ~/.venvs/kokoro/bin/python "kokoro>=0.9" soundfile numpy

Mac

brew install node [email protected] uv ffmpeg espeak-ng
brew install --cask google-chrome
uv venv --python 3.12 ~/.venvs/kokoro
uv pip install --python ~/.venvs/kokoro/bin/python "kokoro>=0.9" soundfile numpy

Windows (PowerShell). espeak-ng comes from the installer on its GitHub releases page.

winget install OpenJS.NodeJS.LTS
winget install Python.Python.3.12
winget install astral-sh.uv
winget install Gyan.FFmpeg
winget install Google.Chrome
uv venv --python 3.12 $HOME\.venvs\kokoro
uv pip install --python $HOME\.venvs\kokoro\Scripts\python.exe "kokoro>=0.9" soundfile numpy

The first time Kokoro runs, it downloads its 82-million-parameter model and whichever voices you use from Hugging Face. After that it works fully offline.

Every video lives in one folder, ~/videos, with a subfolder per video and a starter/ folder to copy from. Each video folder holds everything needed to rebuild it, so fixing a typo six months later is a two-minute job. Keep that folder out of iCloud, Dropbox, or Drive sync. Drafts stay local, and finished videos get uploaded one at a time, on purpose.

The starter files below handle Linux, Mac and Windows paths, and every piece is cross-platform. If a first render hiccups, paste the error into Claude Code or Codex. It's usually a one-line fix.

Your first video, step by step

A slide explainer is one folder with four files. The full code is in the appendix at the end; here's what each one does and how they fit together.

FileWhat it does
script.mdThe narration, one paragraph per scene
narrate.pyVoices each paragraph with Kokoro and writes narration.js: the text, its duration, and the audio
index.htmlA 1920x1080 canvas. draw(t) paints any moment; score() lays out the voice and a soft music bed
render.mjsMixes the audio once, draws every frame at its exact time, and hands both to ffmpeg

The loop:

  1. Copy the starter into a new folder, say ~/videos/my-video, and write your script: one paragraph per scene.
  2. Voice it with ~/.venvs/kokoro/bin/python narrate.py script.md narration.js. On Windows the Python lives at $HOME\.venvs\kokoro\Scripts\python.exe. Rerun this only when a line changes.
  3. Serve the folder from a second terminal: python3 -m http.server 8781 --bind 127.0.0.1 (python on Windows).
  4. Check stills. Open http://127.0.0.1:8781/index.html?t=12.5 in any browser to see the video at 12.5 seconds. Every moment of the video is a URL, so you can check the key beat of every scene before rendering anything.
  5. Render: node render.mjs my-video.mp4 8781.
  6. Watch it with sound. Pacing and pronunciation problems only show up in playback, never in stills.

The output is H.264 video with AAC audio, set up to start playing before it finishes downloading. A 3.5-minute video lands around 9.5 MB, small enough to drop into a wiki, a slide deck, or a chat thread.

To grow the starter, add a paragraph to script.md and a matching drawing function to the SCENES list in index.html. That's the whole model: one paragraph, one function, one scene.

Making it sound good

The voice

Kokoro handles nearly all the narration. It's free, it runs offline, and it re-voices a line in seconds, which is what keeps edits quick. Good starting voices are af_heart and af_bella (female) and am_michael (male). Voices are just tensors, so you can blend them: .85 * am_michael + .15 * af_heart gives a lighter male voice. Speed 1.2 to 1.3 suits explainers. Switching voices or blends between videos keeps a series from sounding like one person.

For a final cut that needs a more expressive read, ElevenLabs on eleven_v4_turbo is worth it. It costs half the credits of eleven_v4 and still sounds very good. Two habits save money there. First, voice ten seconds in two or three candidate voices before committing. Second, turbo ignores the speed setting, so if it reads slower than the Kokoro draft, re-time the edit around it instead of paying to re-voice. Only narrate.py changes: swap the Kokoro call for the ElevenLabs API and write the same narration.js.

ElevenLabs can also clone your own voice on a paid plan. Clone only yours, and treat it like a password: anyone holding the API key can make it say anything.

Pronunciation

Fix it in the text before synthesis, not in the audio after. Test a line first: Kokoro gets more right than you'd expect (it reads "uv", "ffmpeg" and even "sauna" correctly), so only fix what it actually misses. narrate.py keeps a short list of spoken forms for the rest: CLAUDE.md loses its "dot" unless it's written "Claude dot M D". For heteronyms and borrowed words, an inline phoneme forces the right sound: "deploy it live" comes out as "liv" until it's written [live](/lˈIv/), and the Finnish sauna stove [kiuas](/kˈiuɑs/) stops sounding like "KIH-yoo-us". The screen still shows the normal spelling.

Music and sound effects

Keep music well under the voice. The starter synthesizes a soft chord per scene right in the page, ducked under the narration, so there's nothing to license. For a real track, Pixabay Music is free for commercial use with no attribution required; save the license page next to the video's files.

For sound effects, the Sonniss GDC Game Audio Bundle is years of professional effects, free and royalty-free (its license does forbid AI training). It ships as enormous zip files, so let the agent search the track lists, download only the part you need, and keep the clips in a shared sfx/ folder. Use them sparingly: a soft click on a cue, a whoosh between scenes, always quieter than the voice.

Filming a real app

Slides explain an idea. For a product, the best explainer is the product itself, actually running, not a mockup and not a screen recording.

The tours at the top of this post were made this way, from HCM Intelligence, Brain Games Arcade, Sauna Rounds and Mini Block Builder, each one the real, working app.

Same rule as before: never record in real time. Playwright drives headless Chromium through the app the way a person would, tapping, typing, scrolling. After each frame renders, it takes a screenshot and advances the page's clock by exactly one frame. Nothing happens in between, so motion stays smooth and no frame is ever dropped. On Linux, Chrome's beginFrame control makes this fast. On Mac and Windows, a virtual-time clock does the same job, just slower.

I wrapped this in a small kit called film-kit. Each shot is a short script of actions keyed to frame numbers, so re-shooting after a design change is one command:

film new ~/code/my-app-film --format 16:9    # also 4:5, 9:16 or 1:1
film capture                                 # shoot every scripted shot
film vo && film build && film preview 0,5,10 # narration, edit, stills
film audio && film render && film deliver && film verify

The kit adds narration with word-timed captions, the app's own sounds, and a music bed, then renders a full-size master and a phone-sized preview. The last step checks the result: audio and video streams, color tags, loudness, frame accuracy, and that the first frame isn't black, because the first frame is your thumbnail.

What works:

  1. Learn the app first. Run it locally and map the routes, flows and screens worth filming.
  2. Pick the data. Film the app as it really is, or, when real records shouldn't be on screen, seed a dozen believable ones (nothing named "test") and reset them at the start of every shot.
  3. Script, then shoot. About 13 lines and 180 words fill 75 seconds. Review a contact sheet of every shot before editing.
  4. Preview before rendering. A still every two seconds catches most problems.

Common gotchas:

  • Animation libraries like Framer Motion run on the Web Animations API and ignore the virtual clock. They need pinning to it.
  • Lazy-loaded routes flash a loading state mid-shot. Visit the route first, or navigate inside the app.
  • Fast-forwarding the page clock desyncs it from the server, so server-computed numbers can come back as zero. Show stored records instead.
  • Anything random, like a shuffle: seed Math.random per shot so re-shoots match the edit.
  • Network: allow only localhost traffic while shooting, and serve fonts locally.

Letting Claude or Codex drive

Once the setup exists, very little is done by hand. One sentence starts a video:

Make a narrated video of <doc or page> for <audience>, about <length>.

The agent drafts the script and shows it for review before rendering anything; a voice or tone preference goes in the same sentence. That works because the conventions are saved in its instructions file (CLAUDE.md for Claude Code, AGENTS.md for Codex): where videos live, which starter to copy, and the rules below. If you make videos often, ask it to turn the whole workflow into a skill.

The setup also keeps getting better. When the agent finds something better mid-project (a pronunciation fix, a smoother transition, a faster render), it proposes the change and the reason, and updates the starter once approved. Every video leaves the next one a little easier.

Lessons from reviewing:

  • Lock the story first. On one product film, the thing that changed most in review was the story, not the visuals: it went from a feature tour to the user's point of view. Story changes after shooting mean re-shooting.
  • Stills for layout, playback for pacing. Stills catch layout problems for free. Only a full watch with sound catches a rushed scene or a mispronounced name.
  • Leave rebuild notes. Every video folder gets a short README with its sources, voice, speed, music, and the exact commands to rebuild it.

And the ground rules the agent gets up front:

  • Keep other people's data off screen. Film the real app, but never with someone else's personal records showing.
  • Hosted voices see your script. Kokoro never leaves your machine; ElevenLabs does. Keep anything confidential out of scripts you send to a hosted voice.
  • Draft until approved. The agent builds locally. Uploading, sharing and publishing wait for a yes.
  • Render frame by frame, never screen-record.

Sharing with rclone

For sharing, the MP4 goes to Google Drive and gets embedded wherever people already read: Confluence, Notion, Google Sites, or whatever your team uses for its wiki. Drive streams the video in its own player, so the page stays light and one file can sit on several pages. In most of these tools you paste the Drive link on its own line and pick the embed option. Set the file's sharing to anyone in your organization with the link first, or readers get a "request access" screen instead of the video.

The upload itself doesn't need a browser. rclone is a free command-line tool that syncs files with Google Drive and most other cloud storage, so the agent can upload a finished video with one command and confirm the copy matches.

Install it (brew install rclone, sudo dnf install rclone, or winget install Rclone.Rclone), then connect it to Drive once with rclone config, which signs you in through your browser. After that, an upload is:

rclone copy my-video.mp4 "gdrive:Demo Videos"
rclone check . "gdrive:Demo Videos" --one-way --include my-video.mp4

The second line compares checksums, so you know the file in Drive is exactly the one you rendered. Drive has one more handy trick: upload a new version of the same file and the link doesn't change, so every page that embeds it picks up the new cut.

Try it

The fastest way in is to let an agent do the setup. This prompt works in Claude Code or Codex; paste it as is:

Read https://www.danbaruch.com/writing/making-videos/, including the appendix.
Set up this machine to make videos that way. Detect whether it's Linux, Mac or Windows
and use the matching commands.

1. Install whatever is missing: Node 22+, Python 3.12, uv, ffmpeg, Google Chrome,
   espeak-ng. Ask me before anything that needs admin rights.
2. Create the Kokoro environment at ~/.venvs/kokoro (CPU-only PyTorch on Linux).
3. Create ~/videos/starter with the four starter files from the appendix.
4. Copy it to ~/videos/test-video, voice it, render it, check the MP4 with ffprobe
   (H.264 + AAC, length about equal to the narration), and show me a still frame.
5. Add a short Videos section to your instructions file (CLAUDE.md for Claude Code,
   AGENTS.md for Codex) so that "make a video of X" works next time: where videos
   live, the starter, and these rules:
   - Kokoro for drafts; ask before any paid voice.
   - Keep other people's personal data out of videos and scripts.
   - Render frame by frame, never screen-record.
   - Never upload or publish without asking me.

Then pick a doc you wish more people had read, and ask for a 90-second video of it.

Appendix: the starter files

These four files make a working two-scene video that renders end to end. render.mjs finds Chrome on Mac, Linux or Windows, or uses the CHROME environment variable if you set it.

script.md

Making videos with Claude Code or Codex. One scene for every line of narration.

Every frame is rendered exactly, never recorded from the screen.

narrate.py

# One Kokoro clip per scene -> narration.js (text, duration, base64 WAV) for index.html.
# Usage: <kokoro venv python> narrate.py script.md narration.js
# script.md = one paragraph per scene. Runs fully offline after the first model download.
import base64, io, json, re, sys
import numpy as np, soundfile as sf
from kokoro import KPipeline
SPEED, SR = 1.2, 24000
pipe = KPipeline(lang_code='a', repo_id='hexgrad/Kokoro-82M')  # downloads model and voices on first run
VOICE = pipe.load_voice('af_heart')  # blend: .85 * pipe.load_voice('am_michael') + .15 * pipe.load_voice('af_heart')
# Spoken forms, applied before synthesis only. Inline phonemes fix heteronyms: [live](/lˈIv/)
SAY = [(r'\bCLAUDE\.md\b', 'Claude dot M D'), (r'\bgo live\b', 'go [live](/lˈIv/)'), (r'\bkiuas\b', '[kiuas](/kˈiuɑs/)')]
lines = [p.strip() for p in open(sys.argv[1], encoding='utf-8').read().split('\n\n') if p.strip()]
clips = []
for line in lines:
    text = re.sub(r'\[([^\]]+)\]\(/[^)]+/\)', r'\1', line)  # what the page shows
    say = line
    for a, b in SAY: say = re.sub(a, b, say)
    audio = np.concatenate([np.asarray(a) for _, _, a in pipe(say, voice=VOICE, speed=SPEED)])
    loud = np.nonzero(np.abs(audio) > .01)[0]  # trim edge silence, keep ~80 ms of air
    audio = audio[max(0, loud[0] - 2000):loud[-1] + 2000]
    buf = io.BytesIO(); sf.write(buf, audio, SR, format='WAV', subtype='PCM_16')
    clips.append({'text': text, 'dur': round(len(audio) / SR, 3), 'wav': base64.b64encode(buf.getvalue()).decode()})
    print(f'{len(audio) / SR:5.2f}s  {text}')
with open(sys.argv[2] if len(sys.argv) > 2 else 'narration.js', 'w', encoding='utf-8') as f:
    f.write('const NARRATION = ' + json.dumps(clips) + ';\n')

index.html

<!doctype html>
<meta charset="utf-8">
<style>html,body{margin:0;height:100%;background:#0E1430;display:grid;place-items:center}
canvas{width:min(100vw,177.78vh);aspect-ratio:16/9}</style>
<canvas id="c" width="1920" height="1080"></canvas>
<script src="narration.js"></script>
<script>
// Starter canvas video: one scene per narration line. draw(t) paints any moment; ?t=4.2 shows a still.
const cv = document.getElementById('c'), ctx = cv.getContext('2d'), FONT = '"Helvetica Neue", Arial, sans-serif';
const clamp = (x, a = 0, b = 1) => Math.min(b, Math.max(a, x)), seg = (t, a, b) => clamp((t - a) / (b - a));
const ease = x => x < .5 ? 4 * x * x * x : 1 - Math.pow(-2 * x + 2, 3) / 2;
// Timeline: each scene lasts LEAD + its narration + TAIL, so editing a line re-times its scene.
const VO = NARRATION, N = VO.length, LEAD = .3, TAIL = .6;
const RB = [0]; for (let k = 0; k < N; k++) RB.push(RB[k] + LEAD + VO[k].dur + TAIL);
const REAL_END = RB[N];
function at(k, phrase) { // seconds into scene k when the narrator reaches `phrase` (by character position)
  const s = VO[k].text.toLowerCase(), i = s.indexOf(phrase.toLowerCase());
  if (i < 0) console.warn('cue not found', k, phrase);
  return LEAD + VO[k].dur * Math.max(0, i) / s.length;
}
function txt(s, x, y, size, color, alpha) {
  ctx.save(); ctx.globalAlpha = alpha; ctx.font = `700 ${size}px ${FONT}`; ctx.fillStyle = color;
  ctx.textAlign = 'center'; ctx.textBaseline = 'middle'; ctx.fillText(s, x, y); ctx.restore();
}
const SCENES = [
  (u, fa) => txt('Making videos with Claude Code or Codex', 960, 500, 72, '#1B2653', fa(0, .8)),
  (u, fa) => { const c = at(1, 'rendered'); txt('Every frame', 960, 440, 64, '#1B2653', fa(0, .6));
               txt('rendered exactly', 960, 560, 48, '#FB194B', fa(c, c + .6)); },
];
function draw(t) {
  ctx.fillStyle = '#F4F6FB'; ctx.fillRect(0, 0, 1920, 1080);
  for (let k = 0; k < N; k++) {
    if (t < RB[k] || t > RB[k + 1]) continue;
    const u = t - RB[k], out = 1 - ease(seg(t, RB[k + 1] - .35, RB[k + 1]));
    SCENES[k](u, (a, b) => Math.min(out, ease(seg(u, a, b))));  // fa(a, b): fade in over [a, b], out at scene end
  }
}
// Audio: render.mjs creates `ac`, `out` (music bus) and `voice`, fills `voBufs`, then calls score(0).
let ac, out, voice, voBufs = [];
async function score(t0) {
  (await Promise.all(voBufs)).forEach((b, k) => { const n = ac.createBufferSource(); n.buffer = b; n.connect(voice); n.start(t0 + RB[k] + LEAD); });
  for (let k = 0; k < N; k++) { // a soft pad chord per scene; add melody, ticks on cues, etc. the same way
    [261.63, 329.63, 392].forEach(f => { const o = ac.createOscillator(), g = ac.createGain(); o.frequency.value = f;
      g.gain.setValueAtTime(0, t0 + RB[k]); g.gain.linearRampToValueAtTime(.03, t0 + RB[k] + .8);
      g.gain.linearRampToValueAtTime(0, t0 + RB[k + 1] + .4); o.connect(g).connect(out); o.start(t0 + RB[k]); o.stop(t0 + RB[k + 1] + .5); });
  }
}
draw(+(new URLSearchParams(location.search).get('t') || 0));
</script>

render.mjs

// Frame-exact render of a canvas video page: mixes the audio offline, draws every frame at its exact
// time, and pipes both to ffmpeg. Load-proof, and faster than real time.
// Usage: serve the video folder (python3 -m http.server 8781), then: node render.mjs out.mp4 [port]
// Page contract (index.html): canvas `cv`, `draw(t)`, `REAL_END` (seconds), narration buffers `voBufs`,
// and `score(t0)`, which schedules voice and music on the globals `ac`, `out` and `voice`.
import { spawn } from 'node:child_process';
import { writeFileSync, mkdtempSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
const [OUT = 'video.mp4', PORT = '8781'] = process.argv.slice(2), FPS = 30, DBG = 9339;
const CHROME = process.env.CHROME || {
  darwin: '/Applications/Google Chrome.app/Contents/MacOS/Google Chrome',
  win32: 'C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe',
}[process.platform] || '/usr/bin/google-chrome';
const dir = mkdtempSync(join(tmpdir(), 'render-')), WAV = join(dir, 'mix.wav');
const chrome = spawn(CHROME, ['--headless=new',
  `--remote-debugging-port=${DBG}`, `--user-data-dir=${dir}`, '--window-size=1920,1080',
  `http://127.0.0.1:${PORT}/index.html?t=0`], { stdio: 'ignore' });
const sleep = ms => new Promise(r => setTimeout(r, ms));
async function connect() { // minimal Chrome DevTools Protocol client (Node 22+ has a global WebSocket)
  let t;
  for (let i = 0; i < 60 && !t; i++) {
    try { t = (await (await fetch(`http://127.0.0.1:${DBG}/json/list`)).json()).find(x => x.type === 'page'); } catch {}
    if (!t) await sleep(400);
  }
  const ws = new WebSocket(t.webSocketDebuggerUrl), waiters = new Map(); let id = 0;
  await new Promise(r => ws.addEventListener('open', r, { once: true }));
  ws.addEventListener('message', e => { const m = JSON.parse(e.data); waiters.get(m.id)?.(m); waiters.delete(m.id); });
  return async expression => {
    const i = ++id, m = await new Promise(r => { waiters.set(i, r); ws.send(JSON.stringify({ id: i,
      method: 'Runtime.evaluate', params: { expression, returnByValue: true, awaitPromise: true } })); });
    if (m.error || m.result.exceptionDetails) throw new Error(JSON.stringify(m.error || m.result.exceptionDetails).slice(0, 400));
    return m.result.result.value;
  };
}
try {
  const ev = await connect(); await sleep(2500); // let fonts and narration load
  const wav = await ev(`(async () => {
    const SR = 44100, len = Math.ceil((REAL_END + 1) * SR);
    ac = new OfflineAudioContext(1, len, SR); const comp = ac.createDynamicsCompressor();
    out = ac.createGain(); out.gain.value = .28; voice = ac.createGain(); voice.gain.value = 1.15;
    out.connect(comp); voice.connect(comp); comp.connect(ac.destination);
    voBufs = NARRATION.map(c => ac.decodeAudioData(Uint8Array.from(atob(c.wav), ch => ch.charCodeAt(0)).buffer));
    await score(0);
    const d = (await ac.startRendering()).getChannelData(0), b = new DataView(new ArrayBuffer(44 + d.length * 2));
    const w = (o, s) => [...s].forEach((c, i) => b.setUint8(o + i, c.charCodeAt(0)));
    w(0, 'RIFF'); b.setUint32(4, 36 + d.length * 2, true); w(8, 'WAVEfmt '); b.setUint32(16, 16, true);
    b.setUint16(20, 1, true); b.setUint16(22, 1, true); b.setUint32(24, SR, true); b.setUint32(28, SR * 2, true);
    b.setUint16(32, 2, true); b.setUint16(34, 16, true); w(36, 'data'); b.setUint32(40, d.length * 2, true);
    for (let i = 0; i < d.length; i++) b.setInt16(44 + i * 2, Math.max(-1, Math.min(1, d[i])) * 32767, true);
    const u = new Uint8Array(b.buffer); let s = '';
    for (let i = 0; i < u.length; i += 0x8000) s += String.fromCharCode(...u.subarray(i, i + 0x8000));
    return btoa(s); })()`);
  writeFileSync(WAV, Buffer.from(wav, 'base64')); console.log('audio mixed');
  const n = Math.ceil((await ev('REAL_END')) * FPS);
  const ff = spawn('ffmpeg', ['-loglevel', 'error', '-y', '-f', 'image2pipe', '-framerate', String(FPS), '-i', '-',
    '-i', WAV, '-c:v', 'libx264', '-preset', 'slow', '-crf', '24', '-pix_fmt', 'yuv420p', '-c:a', 'aac', '-b:a', '128k',
    '-shortest', '-movflags', '+faststart', OUT], { stdio: ['pipe', 'inherit', 'inherit'] });
  const done = new Promise(r => ff.on('close', r));
  for (let f = 0; f < n; f += 15) { // 15 frames per round trip
    const ts = Array.from({ length: Math.min(15, n - f) }, (_, i) => (f + i) / FPS);
    const imgs = await ev(`${JSON.stringify(ts)}.map(t => (draw(t), cv.toDataURL('image/jpeg', .92).slice(23)))`);
    for (const im of imgs) if (!ff.stdin.write(Buffer.from(im, 'base64'))) await new Promise(r => ff.stdin.once('drain', r));
    if (f % 900 === 0) console.log(`frame ${f}/${n}`);
  }
  ff.stdin.end(); await done; console.log('wrote', OUT);
} finally { chrome.kill(); await sleep(500); rmSync(dir, { recursive: true, force: true }); }