← Back
AGIHunt

AGIHunt/blurt

Show it. Say it. Your AI gets it. Record your screen and talk — your coding agent turns it into bugs, ideas and to-dos, with the exact frames marked. 口喷鸡:边看边喷,AI 全懂。Agent Skill + macOS menu-bar app for Claude Code / Codex.

View on GitHub ↗
agent-skillsai-agentsbug-reportclaude-codecodexfeedbackfeishumacosscreen-recordingspeech-recognitionvibe-codingvoice
Stars
40
Forks
6
Watchers
40
Open issues
5
Contributors
2
Language
Python
License
MIT License
Default branch
main
Created Sep 25, 2026Updated Sep 30, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

blurt

blurt 🐔

Show it. Say it. Your AI gets it.
Record your screen, think out loud — your coding agent turns it into bug tickets, idea boards and todos.
中文 · works with Claude Code, Codex and any agent that runs skills

blurt-promo.mp4

▶ 80 seconds, sound on 🐔 · made entirely in code — source


Voice alone is blind. Screenshots plus typing is slow. The most natural way to tell an AI what you mean is the way you'd tell a colleague sitting next to you: point at the screen and talk. blurt records exactly that, and your agent does the rest:

  • splits a rambling half hour into separate items, even when you jump between topics and correct yourself
  • picks the right frame for each one and boxes the exact spot, using where your mouse actually was
  • writes it up: actual vs. expected, repro steps, owner, and the code that's probably responsible
  • lets you triage it in a review page, one item at a time with the keyboard
  • files it to Feishu/Lark, GitHub, Linear or Markdown, or goes straight to fixing it

What used to take two days of screenshots, red boxes and spreadsheet rows is now a 30-minute walkthrough.

What people blurt

🐞 Polish a vibe-coded product An agent built it overnight; now you walk through every page and rant. You get a clean bug list with code pointers, ready for the agent to fix. This is the fastest way to push AI to the finish line.
💡 Capture ideas while browsing "I like how this site does onboarding… and this pricing page…" You get an idea board: each idea, why you had it, where it came from, the next step, plus a one-page digest.
🤝 Hand off from anyone PMs, designers, ops, clients: anyone can record with the Blurt app, no repo needed, and send the video. The developer's agent processes it with the code at hand.
🔎 Research and walkthroughs Competitor tours, UX research, "how this works": you get notes and findings with the frames to prove them.
🖥️ Not just web apps Terminals, TUIs and desktop apps record and get boxed the same way. For phones, use the built-in screen recorder with the mic on; for hardware, film it with your phone. Hand the video to your agent: "process this video".

One recording can mix all of these. The agent decides what each item is, and teams can add their own lenses (e.g. ux-research, sales-call, sop).

Get it

npx skills add AGIHunt/blurt

Claude Code: /plugin marketplace add AGIHunt/blurt · or just paste this repo's URL to your agent and ask it to install the skill.

Then, in any project, tell your agent "start blurt" / 「开始口喷」. The first run picks a speech model for your machine and installs the Blurt menu-bar app.

How it feels

  1. Draw the area to record: drag a region, click a window, or go full screen. Tabs, bookmarks and everything else stay out of the video.
  2. 3-2-1, then talk. A tiny floating bar shows time and mic level, with ⏸ pause, ↺ redo and Finish. The bar itself is never recorded. Shortcuts: ⌥⇧P pause, ⌥⇧S finish.
  3. Click Finish and get back to work. Your agent transcribes locally, writes the items, and opens the review page.
  4. Triage like a feed: A keep · X drop · J/K next/prev · Z undo · G list · V overview. Then export, or say "fix them".

Always on: the Blurt app lives in your menu bar. ⌥⇧R starts a recording from anywhere and ⌥⇧R again finishes it; ⌥⇧B opens the menu. Recordings go to a workspace: ~/Blurt by default, or a project you bind, so your agent can process them with the code. Turn on After recording → Claude Code / Codex and every recording gets processed in the background, with the review page popping up when it's ready.

Solo or team

  • Solo: record → your agent in the same repo processes and fixes. No forms, no copy-paste.
  • Team: teammates without a repo just download Blurt for macOS (unzip, then right-click → Open the first time) and press ⌥⇧R. Anyone records with the app, then uses Recent recordings → Copy video and pastes it into Slack/Feishu. Or bind a shared project folder. Items keep the recorder's name, and exports land in the team's existing tables with your column names.

Under the hood

  • Local-first speech recognition. SenseVoice via sherpa-onnx is about 240 MB, very fast on any CPU, and handles mixed Chinese/English well. Whisper on Apple Silicon or NVIDIA. You can also bring your own Groq, OpenAI or DashScope key. By default nothing leaves your machine.
  • Native recorder. ScreenCaptureKit on macOS with audio and video in sync to within one frame. Turn on system audio for calls and demos: it goes on its own track, so the transcript tells your words from the other side's. Tk + ffmpeg on Windows.
  • Deterministic tools, flexible model. Scripts handle the recording, speech-to-text, frames, the review page and exports. All the judgement is left to your agent (see SKILL.md), so it adapts to your product, your language and your team.
  • Any language in, same language out. SenseVoice covers Chinese, English, Japanese, Korean and Cantonese; for German, French, Spanish and the rest (99 languages) use Whisper locally or a cloud key. The first run picks one for the language you speak, and items come back in that language.

FAQ

Does it burn a lot of tokens? The video is never fed to the model. Speech is transcribed locally (free), the agent reads the text, and it looks at a few frames only for the moments that become items. Cost follows how many things you talk about, not how long you record; silence and clicking around cost nothing. Quality follows the model: use one with vision, the stronger the better. Two real sessions (Claude Code, Opus 5.5, a production web app):

recording speech items transcription (local) recording → review page new input / output tokens at API prices
7.0 min 2.8 min 9 4 s ~4 min ~142k / ~19k ~$2
8.6 min 5.0 min 20 5 s ~5 min ~150k / ~20k ~$2

Plus cache reads of the ongoing conversation (~4–5M tokens, 1/20 of the input price, included above), which depend on how long your chat already is.

I don't do frontend. Is it for me? Yes, if you can see the problem on screen or film it with a phone. See Not just web apps above.

Roadmap

Browser capture (console errors and network failures lined up with the video) · circle-to-highlight gestures · Windows tray app · a hosted speech API · more lenses and exporters · toward a personal assistant that watches, listens and keeps your projects moving. See TODO.md.

MIT · made by AGI Hunt · 🐔 if blurt saved you a day, a ⭐ helps others find it