Dec 2025 – Aug 2026Remote, Saudi Arabia
Built an end-to-end production AI voice call system (VAD, STT, LLM, TTS) optimized for real-time conversation latency.
- Built and deployed an end-to-end AI voice call system (VAD, STT, LLM, TTS), optimizing the pipeline to achieve ∼400 - 800ms inference latency for natural conversations.
- Curated and preprocessed Saudi dialect audio data, building a full pipeline for fine-tuning open-source TTS and ASR models for high-fidelity speech synthesis.
- Deployed and maintained large language models in production using GCP for low-latency streaming inference.

