An agent skill that turns a short brief into a directed motion-graphics MP4.
Your AI agent researches only what you did not supply and writes a direction. It builds the scenes in HTML, CSS and SVG, and renders them frame-exact with headless Chromium and ffmpeg. Sound effects are synthesized from the motion on screen, so every hit lands on its frame.
| Video | Example request |
|---|---|
| Product teaser | "A 30-second teaser for the app in this repo." |
| Feature explainer | "Explain how our invoice scanner works, 45 seconds, from these screenshots." |
| App demo or tutorial | "A walkthrough of the sign-up flow, portrait, for Instagram." |
| Lesson or explainer | "A 60-second video that explains how DNS resolves a domain." |
| Data story | "Turn these quarterly numbers into a short animated story." |
| Event or announcement | "Announce our meetup on 12 March, with the venue photo." |
| Music-synced edit | "Cut these screenshots to this track, on the beat." |
| Kinetic typography | "Our manifesto as kinetic type, no UI." |
Give it whatever you have: a repo, a URL, screenshots, a logo, footage, music, a voiceover. From a one-line brief, the agent researches the subject and puts only real facts on screen.
Ask your AI agent:
Install the motion-video skill from https://github.com/farhan-syah/motion-video-skill for yourself, then run its setup.
For the agent doing the install:
- Clone:
git clone https://github.com/farhan-syah/motion-video-skill <skills-folder>/motion-video. Use the folder your tool loads skills from, else any folder the user picks. - Install:
cd <skills-folder>/motion-video/scripts && bun install, ornpm installwithout Bun. - Check: run
node scripts/video.mjs doctorin the user's folder. Fix what it marks missing, or tell the user. - Speech (optional): the first
speakortranscribeinstalls the local speech runtime and downloads its models once."models"in~/.config/motion-video/config.jsonpoints at a folder the user keeps. - Report the installed version and anything
doctorfound missing.
Manual install: git clone https://github.com/farhan-syah/motion-video-skill motion-video into your agent's skills folder, then cd motion-video/scripts && bun install. An agent without a skills feature must read motion-video/SKILL.md before it makes a video.
Requirements:
- Bun (https://bun.sh) installs dependencies and fetches fonts. npm works as a fallback.
- Node.js 20+.
- ffmpeg and ffprobe on
PATH. - Chrome or Chromium, found automatically. Otherwise run
bunx playwright install chromium-headless-shellinscripts/. - GPU (optional): Chromium draws on any available GPU, and an NVIDIA GPU also encodes with NVENC. Otherwise rendering runs on the CPU.
- Installed version:
node scripts/video.mjs version, also printed bydoctor.CHANGELOG.mdlists what each version adds. - When to check: when the user asks, or when
doctorsays one is due. It is due when none ran yet, or the last is over 30 days old. The skill never uses the network for this unasked. - Check:
node scripts/video.mjs update --checkshows what a newer version adds. - Update: with the user's consent,
node scripts/video.mjs updatepulls the latest version and reinstalls the dependencies. On local edits or a branch other thanmain, it changes nothing and says what to do. - Not a git clone: replace the folder with a fresh clone, then run the install step again.
Ask for a video in plain words. The agent follows SKILL.md. It checks your machine and files, gathers facts, and asks a few questions with defaults (say "one-shot" to skip them). Then it writes direction.md, builds one HTML file per scene, lints each scene, and renders.
The kit holds UI parts, 3D props, atmospheres, maps from Natural Earth data, 1854 icons and any Fontsource font. Scenes can also use canvas, WebGL (three.js), GSAP, Lottie and footage.
Output: out/video.mp4 (master) and out/video-compressed.mp4, about a quarter the size, for chat apps and social uploads.
Sound: music only when you supply it. Otherwise each motion gets a synthesized sound effect.
Voiceover: your recording, a TTS file, subtitles or a script (plain or timed per line), used as given. A local Whisper model times every word, so captions and scenes follow the speech. For a script without a recording, the agent picks a local text-to-speech engine that fits your language and machine:
| Engine | Languages | Machine |
|---|---|---|
| VoxCPM2 | 30, with described or cloned voices | NVIDIA GPU with 8 GB, and uv |
| Kokoro | English | Any, CPU only included |
| Any Hugging Face model | The model's, when its license allows commercial use | Any, CPU only included, and uv |
| Any TTS you already use | The tool's | Wherever it runs |
The default is the best one your machine runs: VoxCPM2 with that GPU, else Kokoro. node scripts/video.mjs doctor shows which of these your machine runs. Languages, hardware and adding your own voice or model: references/tts.md.
The agent runs node scripts/video.mjs <command> from the project folder. You can run it too. Every command, video.json field and environment variable: references/cli.md.
| Path | Contents |
|---|---|
SKILL.md |
The workflow the agent follows |
CHANGELOG.md |
What each version adds |
references/ |
Direction, concepts, building, craft thresholds, styles, narration, text to speech, CLI, icon index |
templates/ |
What init copies into a project: CSS kit, runtime, fonts, icons, scene and direction templates |
scripts/ |
The CLI, renderer, linter, sound synthesis and audit |
scripts/dev/ |
Tools for editing the sound catalog |
See CONTRIBUTING.md for what the skill accepts and how to send a change.
- Tests:
node --test scripts/lib/ - Sound catalog review:
references/craft.md, Changing a sound recipe
MIT. See LICENSE.
Bundled:
- Lucide icons (ISC):
templates/icons.LICENSE.txt - Inter and JetBrains Mono fonts (SIL OFL 1.1):
templates/fonts/OFL.txt
Installed as dependencies:
- Playwright (Apache 2.0)
- three.js and lottie-web (MIT)
- GSAP, under its own no-charge license: https://gsap.com/standard-license
- world-atlas, d3-geo and topojson-client (ISC), with Natural Earth data (public domain)
Installed on first use of speak or transcribe:
- transformers.js (Apache 2.0) and ONNX Runtime (MIT)
- Kokoro-82M (Apache 2.0) and Whisper (MIT) models
- VoxCPM2 (Apache 2.0), only with
speak --engine voxcpm: its Python environment throughuv, and the model unless already on disk - A Hugging Face model named with
speak --model, under its own license. A model without ONNX weights also installs a Python environment throughuv.
Fonts fetched with font keep their own licenses, copied next to each font.