About the role: We're hiring a Research Scientist to advance real-time speech and audio understanding for conversational AI. You'll design and train models for automatic speech recognition, speaker diarization, and paralinguistic understanding (emotion, intent, turn-taking) that run under tight latency budgets in live, two-way voice conversations, and publish your work at top venues.
Responsibilities: Lead research on streaming ASR and end-to-end speech models; develop methods for speaker diarization, robustness to far-field and noisy audio, and low-latency inference; build evaluation for real-time conversational quality; work with product and infrastructure to ship research into a production voice pipeline; publish in venues such as Interspeech, ICASSP, and NeurIPS.
Requirements: PhD in speech processing, machine learning, or a related field, or equivalent research experience; strong publication record in speech/audio ML; expertise in PyTorch and modern sequence models (transformers, conformers, RNN-T); experience with streaming or low-latency inference.
Preferred: self-supervised speech representation learning (wav2vec/HuBERT-style), text-to-speech, or multilingual/low-resource ASR.
Application
Apply for Research Scientist — Speech & Audio Understanding
Zero