Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
-
Updated
Sep 10, 2026 - Python
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
A PyTorch-based Speech Toolkit
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
End-to-End Speech Processing Toolkit
On-device Speech AI for Apple Silicon
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
A python package to build AI-powered real-time audio applications
A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness
A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.
This is the library for the Unbounded Interleaved-State Recurrent Neural Network (UIS-RNN) algorithm, corresponding to the paper Fully Supervised Speaker Diarization.
Fun-ASR speech recognition models, with native Hugging Face Transformers support for Fun-ASR-Nano and separate FunASR, vLLM and llama.cpp deployment paths.
Research and Production Oriented Speaker Verification, Recognition and Diarization Toolkit
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
turnkey self-hosted offline transcription and diarization service with llm summary
一站式全自动字幕生成软件,下载、转录、翻译、压制全流程覆盖,无需人工介入 / One-stop automated subtitle generator. Handles downloading, transcription, translation, and hardcoding—zero human intervention required.
To associate your repository with the speaker-diarization topic, visit your repo's landing page and select "manage topics."