Local, full-stack voice generation and cinematic dubbing. No API keys. No cloud. Just run it. Built on the open-source OmniVoice 600-language zero-shot diffusion model.
- 🎬 Video Dubbing — transcribe, translate, re-voice, and mux back into MP4 with selective track export.
- 🎧 Vocal Isolation — built-in
demucsautomatically splits speech from music, keeping original background audio perfectly preserved. - 🧬 Voice Cloning & Design — Clone specific voices from just a 3-second audio clip, or design completely new studio profiles with tags like
female, british accent, excited. - ⚡ Cross-Platform Native Execution — Auto-detects and accelerates inference using Apple Silicon (MPS), NVIDIA (CUDA), AMD (ROCm), or standard CPU.
- 🔊 Per-Segment Mixing — Fine-grained volume/gain control per dubbed segment (0–200%) for broadcast-quality audio balancing.
- ⌨️ Keyboard-Driven Workflow —
⌘+Enterto generate,⌘+Sto save,⌘+Z/⌘+Shift+Zfor undo/redo. - 📡 Live Model Telemetry — Real-time CPU/RAM/VRAM stats + model warm-up indicator (idle → loading → ready).
The easiest way to run OmniVoice Studio locally or on a cloud VM is via Docker. Our environment utilizes an optimized pytorch/pytorch configuration which seamlessly enables zero-config GPU passthrough if your host supports it.
git clone https://github.com/debpalash/OmniVoice-Studio.git
cd OmniVoice-Studio
docker compose up --build -dThat's it! Open http://localhost:8000 in your browser.
Tip
Windows/WSL Users: Make sure your NVIDIA drivers are up to date. Docker Desktop automatically passes GPU capabilities to this container!
Cloud VMs (AWS, RunPod): The image inherently supports CUDA 12.1. As long as nvidia-container-toolkit is installed on your host, --gpus all binds natively.
Quickly get OmniVoice Studio running natively on your hardware if you want to develop or modify code.
Prerequisites: Ensure ffmpeg is installed on your system.
Install standard modern web tooling: Bun and uv.
git clone https://github.com/debpalash/OmniVoice-Studio.git
cd OmniVoice-Studio
# Boot the Backend
uv sync
uv run uvicorn backend.main:app
# Boot the Frontend (in a separate terminal)
bun install
bun run devOmniVoice Studio launches exactly two micro-services:
| Service | Protocol | Details |
|---|---|---|
| Frontend | http://localhost:5173 |
The real-time React UI — spanning cloning, design, and audio workspace. |
| Backend | http://localhost:8000 |
The FastAPI server handling model inference, translation pipelines, transcriber tasks. |
Note
First run optimization: Model weights (approx. 1.2 GB) automatically download from HuggingFace the first time you execute a generation sequence. Subsequent launches trigger instantly from cache. (Tip: Set HF_TOKEN in your environment for faster, authenticated downloads!)
The studio is highly functional today, but we are aggressively expanding. Watch the roadmap to see what's shipping next:
- Zero-shot voice cloning & complex voice design.
- Full video cinematic dubbing pipeline (transcribe → translate → synthesize → mux).
- Vocal isolation utilizing demucs alongside background audio retention.
- Embedded waveform timeline editor for micro-segment-level audio manipulation.
- Live system telemetry tracking (CPU, RAM, GPU VRAM usage).
- Targeted multi-speaker diarization — auto-assign unique voice profiles per active speaker.
- Studio project persistence — save, load, and cache multi-track projects seamlessly via local SQLite.
- Production SRT/VTT subtitle export packaged alongside the dubbed
.mp4video output. - Selective track export — choose exactly which language tracks (Original, DE, ES, etc.) to include in final MP4.
- Per-segment volume/gain control with real-time mixing (0–200%).
- Undo/redo system for all segment edits with 50-action history depth.
- Keyboard shortcuts:
⌘+Entergenerate,⌘+Ssave,⌘+Z/⌘+Shift+Zundo/redo. - Drag-and-drop file uploads for both video and clone audio sources.
- Model warm-up indicator with live status pill (idle/loading/ready).
- Confirmation dialogs for all destructive actions (delete project/history/profile).
- UI preferences persistence (sidebar state, zoom, active tab) across sessions.
- Polished glassmorphism design system with micro-animations, focus rings, and custom scrollbars.
- Real Speaker Diarization — ML-based diarization via pyannote.audio for true multi-speaker identification.
- A/B Voice Comparison — Side-by-side voice audition for casting decisions.
- Scene-Aware Dubbing — FFmpeg scene detection to auto-split segments at visual cuts.
- Lip-Sync Scoring — Analyze dubbed audio duration against original speaker timing with color-coded badges.
- Batch Processing — Centralized async task queue ensuring sequential GPU execution with reconnectable SSE streams.
- Advanced Export Suite — VTT subtitles, per-segment WAV ZIP, compressed MP3, and stem export (vocals + background separate).
- Streaming TTS — Chunked WAV streaming with progressive download and auto-playback.
- Native Desktop Applications — Dedicated client apps for macOS, Windows, and Linux.
- One-Click Deployment — Docker image packages engineered for zero-config GPU passthrough.
- Selective Track Export: Choose exactly which audio tracks to include in the final MP4. Uncheck Original, keep only German — get a single-track export. Full per-track checkbox UI with dynamic FFmpeg stream index remapping.
- Undo/Redo System: Full
⌘+Z/⌘+Shift+Zundo/redo for all segment edits (text, voice, volume, delete). 50-action deep history stack. - Per-Segment Volume Control: Inline gain slider (0–200%) per segment row in the dub table. Backend applies gain during audio assembly with safe clamping.
- Keyboard Shortcuts:
⌘+Enterto generate,⌘+Sto save project. Browser default overrides prevented. - Model Status Indicator: Live status pill in the header showing model warm-up state (idle → loading → ready). New
/model/statusbackend endpoint. - Drag-and-Drop Everywhere: Video upload already supported drop — now clone audio upload does too, with pink highlight on hover.
- Confirmation Dialogs: All destructive actions (delete project, profile, history item, clear all history) now require confirmation.
- Session Persistence: Sidebar collapsed state, active tab, and zoom level now persist across browser sessions via localStorage.
- CSS Design System Overhaul: Anti-aliased text, input focus glow rings, button hover shimmer, progress bar shimmer animation, fade-in on history items, selection color branding, Firefox scrollbar support,
tabular-numsfor timestamp columns. - AudioContext Pooling:
playPing()synthesis notification reuses a single AudioContext instead of creating one per call (browsers cap at ~6).
- The Cinematic Studio Interface: Exhaustively re-engineered the UI to prioritize a high-density, real-estate optimized workflow featuring a dynamic UI zoom scalar (
Small,Normal,Max). We minimized dead space and overhauled the widget layout keeping crucial tuning metrics immediately accessible. - Multi-Track Timeline: Deeply integrated a multi-layered waveform sequence interface supporting precision audio segment positioning, unmuted live preview playback, localized track timing, and unconstrained draggable positioning manipulation.
- Persistent Local Projects: Put a complete stop to ephemeral state loss. All workspace metrics are successfully wrapped into
Projectslogged directly within a native embeddedSQLitedatabase. Workflows reliably survive browser shutdowns or server API reboots. - AI Cast Diarization: Dropped in an offline
Pyannote+WhisperXfusion pipeline evaluating multi-speaker metadata and categorizing overlapping, distinct speakers. Rapidly "cast" clone overrides seamlessly over complex dialogue tracks. - Polishing & Asset Control: Cleaned cross-stack filename parsing and exported media rendering via
ffmpeg, stabilizing codec dependencies, and deployed a unified customOmniVoice Studioscalable aesthetic asset system.
