Show it. Say it. Your AI gets it.
Record your screen, think out loud — your coding agent turns it into bug tickets, idea boards and todos.
中文 · works with Claude Code, Codex and any agent that runs skills
blurt-promo.mp4
▶ 80 seconds, sound on 🐔 · made entirely in code — source
Voice alone is blind. Screenshots plus typing is slow. The most natural way to tell an AI what you mean is the way you'd tell a colleague sitting next to you: point at the screen and talk. blurt records exactly that, and your agent does the rest:
- splits a rambling half hour into separate items, even when you jump between topics and correct yourself
- picks the right frame for each one and boxes the exact spot, using where your mouse actually was
- writes it up: actual vs. expected, repro steps, owner, and the code that's probably responsible
- lets you triage it in a review page, one item at a time with the keyboard
- files it to Feishu/Lark, GitHub, Linear or Markdown, or goes straight to fixing it
What used to take two days of screenshots, red boxes and spreadsheet rows is now a 30-minute walkthrough.
| 🐞 Polish a vibe-coded product | An agent built it overnight; now you walk through every page and rant. You get a clean bug list with code pointers, ready for the agent to fix. This is the fastest way to push AI to the finish line. |
| 💡 Capture ideas while browsing | "I like how this site does onboarding… and this pricing page…" You get an idea board: each idea, why you had it, where it came from, the next step, plus a one-page digest. |
| 🤝 Hand off from anyone | PMs, designers, ops, clients: anyone can record with the Blurt app, no repo needed, and send the video. The developer's agent processes it with the code at hand. |
| 🔎 Research and walkthroughs | Competitor tours, UX research, "how this works": you get notes and findings with the frames to prove them. |
| 🖥️ Not just web apps | Terminals, TUIs and desktop apps record and get boxed the same way. For phones, use the built-in screen recorder with the mic on; for hardware, film it with your phone. Hand the video to your agent: "process this video". |
One recording can mix all of these. The agent decides what each item is, and teams can add their own
lenses (e.g. ux-research, sales-call, sop).
npx skills add AGIHunt/blurtClaude Code: /plugin marketplace add AGIHunt/blurt · or just paste this repo's URL to your agent and ask it to install the skill.
Then, in any project, tell your agent "start blurt" / 「开始口喷」. The first run picks a speech model for your machine and installs the Blurt menu-bar app.
- Draw the area to record: drag a region, click a window, or go full screen. Tabs, bookmarks and everything else stay out of the video.
- 3-2-1, then talk. A tiny floating bar shows time and mic level, with ⏸ pause, ↺ redo and Finish. The bar
itself is never recorded. Shortcuts:
⌥⇧Ppause,⌥⇧Sfinish. - Click Finish and get back to work. Your agent transcribes locally, writes the items, and opens the review page.
- Triage like a feed:
Akeep ·Xdrop ·J/Knext/prev ·Zundo ·Glist ·Voverview. Then export, or say "fix them".
Always on: the Blurt app lives in your menu bar. ⌥⇧R starts a recording from anywhere and ⌥⇧R again
finishes it; ⌥⇧B opens the menu. Recordings go to a workspace: ~/Blurt by default, or a project you bind, so
your agent can process them with the code. Turn on After recording → Claude Code / Codex and every recording gets
processed in the background, with the review page popping up when it's ready.
- Solo: record → your agent in the same repo processes and fixes. No forms, no copy-paste.
- Team: teammates without a repo just download Blurt for macOS
(unzip, then right-click → Open the first time) and press
⌥⇧R. Anyone records with the app, then uses Recent recordings → Copy video and pastes it into Slack/Feishu. Or bind a shared project folder. Items keep the recorder's name, and exports land in the team's existing tables with your column names.
- Local-first speech recognition. SenseVoice via sherpa-onnx is about 240 MB, very fast on any CPU, and handles mixed Chinese/English well. Whisper on Apple Silicon or NVIDIA. You can also bring your own Groq, OpenAI or DashScope key. By default nothing leaves your machine.
- Native recorder. ScreenCaptureKit on macOS with audio and video in sync to within one frame. Turn on system audio for calls and demos: it goes on its own track, so the transcript tells your words from the other side's. Tk + ffmpeg on Windows.
- Deterministic tools, flexible model. Scripts handle the recording, speech-to-text, frames, the review page and exports. All the judgement is left to your agent (see SKILL.md), so it adapts to your product, your language and your team.
- Any language in, same language out. SenseVoice covers Chinese, English, Japanese, Korean and Cantonese; for German, French, Spanish and the rest (99 languages) use Whisper locally or a cloud key. The first run picks one for the language you speak, and items come back in that language.
Does it burn a lot of tokens? The video is never fed to the model. Speech is transcribed locally (free), the agent reads the text, and it looks at a few frames only for the moments that become items. Cost follows how many things you talk about, not how long you record; silence and clicking around cost nothing. Quality follows the model: use one with vision, the stronger the better. Two real sessions (Claude Code, Opus 5.5, a production web app):
| recording | speech | items | transcription (local) | recording → review page | new input / output tokens | at API prices |
|---|---|---|---|---|---|---|
| 7.0 min | 2.8 min | 9 | 4 s | ~4 min | ~142k / ~19k | ~$2 |
| 8.6 min | 5.0 min | 20 | 5 s | ~5 min | ~150k / ~20k | ~$2 |
Plus cache reads of the ongoing conversation (~4–5M tokens, 1/20 of the input price, included above), which depend on how long your chat already is.
I don't do frontend. Is it for me? Yes, if you can see the problem on screen or film it with a phone. See Not just web apps above.
Browser capture (console errors and network failures lined up with the video) · circle-to-highlight gestures · Windows tray app · a hosted speech API · more lenses and exporters · toward a personal assistant that watches, listens and keeps your projects moving. See TODO.md.
MIT · made by AGI Hunt · 🐔 if blurt saved you a day, a ⭐ helps others find it
