ASRX
Universal ASR Forced Alignment & Diarization Engine
🏆 Published on PyPI · pip install asrx
A model-agnostic speech alignment engine that synchronizes audio with transcripts from any ASR model or LLM down to millisecond precision.
What it does
- Built and published an open-source Python engine that aligns transcripts from any ASR model or LLM with audio to generate millisecond-level word timestamps, confidence scores, VAD segments, and speaker turns.
- Pluggable backend architecture supporting Wav2Vec2, Meta MMS (1,107+ languages), NVIDIA NeMo, Silero VAD, and Pyannote with unified streaming and batch CLI/API interfaces.
- Eliminates vendor lock-in by decoupling alignment from speech-to-text models, enabling word-level synchronization for Whisper, Cohere, Deepgram, and open LLMs.