← Back
leemysw

leemysw/yovoice

Open-source voice creation for macOS and Windows. Local TTS, voice cloning, and emotion control — no cloud APIs or per-character fees.

View on GitHub ↗https://yovoice.leemysw.com ↗
audiodesktop-appmacostext-to-speechttsvoicevoice-cloningwindows
Stars
317
Forks
38
Watchers
317
Open issues
7
Contributors
2
Language
TypeScript
License
Apache License 2.0
Default branch
main
Created Sep 15, 2026Updated Oct 2, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

yovoice

yovoice

Give your words a voice.

Version 0.1.6 License: Apache-2.0 macOS 14+ (Apple Silicon) Windows 10/11 (x64)

English · 简体中文


yovoice is an open-source voice creation tool for macOS and Windows that turns text into natural, expressive speech locally. With voice cloning, emotion control, audio project management, and a choice of TTS models, it requires no cloud API calls and incurs no per-character charges, giving you more control over narration, voiceovers, and audio content creation at a lower cost.

yovoice creation workspace — browser preview


Features

  • Voice and expression — use a reference voice, match a reference performance, adjust emotions, or describe the delivery in words.
  • A complete audio workflow — import or record reference audio, trim clips, preview speech, and export your work.
  • Voice design and cloning — VoxCPM2 offers text-guided voice design, controllable cloning, and transcript-assisted cloning with automatic multilingual handling, and 48 kHz output.
  • Local models — run IndexTTS 2.0 / 2.5, VoxCPM2, OmniVoice, Qwen3-TTS and Kokoro through audio.cpp, with resumable model downloads and GGUF import.
  • Hardware acceleration — Metal on Apple Silicon; CPU, NVIDIA CUDA, and experimental Vulkan on Windows.
  • Agent Skill — ask your AI agent to set up local speech generation and create voiceovers from text and reference audio.

Supported Models

Model Parameters Model file size by precision Core capabilities
IndexTTS 2.0 — Q8 · 3.63 GB
F16 · 4.65 GB
ORIG · 8.08 GB
Chinese/English voice cloning, emotion control, reference performance
IndexTTS 2.5 — Q8 · 3.50 GB
F16 · 4.55 GB
ORIG · 7.89 GB
Multilingual voice cloning, emotion control, pronunciation editing
VoxCPM2 2B Q8 · 2.96 GB
BF16 · 4.77 GB
ORIG · 4.96 GB
Text-guided voice design, voice cloning, transcript-assisted cloning
OmniVoice 0.6B Q8 · 1.35 GB
BF16 · 1.64 GB
F16 · 1.64 GB
Attribute-based voice design, voice cloning, non-verbal sound tags
Qwen3-TTS Base 0.6B Q8 · 1.99 GB
BF16 · 2.52 GB
Reference voice cloning, optional transcript guidance, multilingual speech
Qwen3-TTS Base 1.7B Q8 · 2.70 GB
BF16 · 4.20 GB
ORIG · 4.54 GB
Reference voice cloning, optional transcript guidance, multilingual speech
Qwen3-TTS CustomVoice 1.7B Q8 · 2.82 GB
BF16 · 4.18 GB
9 built-in voices, text-guided style and emotion
Qwen3-TTS VoiceDesign 1.7B Q8 · 2.82 GB
BF16 · 4.18 GB
Voice design from natural-language descriptions, no reference audio required
Kokoro-82M 1.0 Official 82M Q8 · 189.55 MB
BF16 · 211.95 MB
49 built-in voices, multilingual speech excluding Japanese
Kokoro-82M 1.0 82M Q8 · 932.66 MB 54 built-in voices, full multilingual resources including Japanese; import only
Kokoro-82M 1.1-zh 82M Q8 · 255.32 MB 100 Chinese and 3 English voices; experimental, import only

All models are available in the App and CLI. See model capabilities for precisions and parameters. OmniVoice weights use the CC-BY-NC license and are restricted to non-commercial use.


Installation

Choose the package for your platform from the repository’s Releases tab.

Platform Package Install
macOS 14+ · Apple Silicon .dmg Open the disk image and drag yovoice to Applications
Windows 10/11 · x64 -setup.exe Run the installer, then open yovoice from the Start menu

The Windows installer downloads and installs WebView2 Runtime if needed. Models are downloaded inside the app; uninstalling preserves user data in ~/.yovoice.

Check for updates from the app menu on macOS or Windows. Updates are also checked and downloaded in the background; restart to install when ready.

Install yovoice on macOS


Quick Start

  1. Set up a model. Download a model in Settings.
  2. Create a project. Choose Story dubbing or Speech generation.
  3. Add content. Enter text or import subtitles, then choose a character and adjust the voice.
  4. Generate audio. Select Generate speech, then preview, edit and export.

CLI & Agent Skill

Give your agent the yovoice Skill link and ask it to install the Skill and set up the local CLI, engine, and model:

Install this Skill and set up yovoice for local speech generation: https://github.com/leemysw/yovoice/tree/main/skills/yovoice

Then describe what you want:

Read narration.txt using voice.wav as the reference voice, with a calm delivery, and save it as narration.wav.

The agent runs the standalone CLI without opening the desktop app. See the CLI guide for details.


Development

make install
make app-run

See the development guide for prerequisites. It also covers browser preview, tests, and project structure.


Contributing

Bug reports, feature suggestions, and pull requests are welcome. Include your platform, model, and steps to reproduce when reporting a problem. Run make check before submitting code changes.


Acknowledgements

  • audio.cpp by ShugoAI — the local audio inference engine.
  • VoxCPM — voice design and cloning models, licensed under Apache-2.0.
  • IndexTTS — the speech synthesis models behind yovoice.
  • OmniVoice — voice design and cloning models.
  • Qwen3-TTS — voice cloning, built-in voices, and text-guided voice design.
  • Kokoro — lightweight speech synthesis with built-in voices: Kokoro-82M and Kokoro-82M-v1.1-zh.

License

Apache-2.0. Dependencies retain their original licenses; IndexTTS models are covered separately by their model license.