Open role

Research Scientist — Speech & Audio Understanding

Apply now
ZeroRemoteFull-time

About the role: We're hiring a Research Scientist to advance real-time speech and audio understanding for conversational AI. You'll design and train models for automatic speech recognition, speaker diarization, and paralinguistic understanding (emotion, intent, turn-taking) that run under tight latency budgets in live, two-way voice conversations, and publish your work at top venues.

Responsibilities: Lead research on streaming ASR and end-to-end speech models; develop methods for speaker diarization, robustness to far-field and noisy audio, and low-latency inference; build evaluation for real-time conversational quality; work with product and infrastructure to ship research into a production voice pipeline; publish in venues such as Interspeech, ICASSP, and NeurIPS.

Requirements: PhD in speech processing, machine learning, or a related field, or equivalent research experience; strong publication record in speech/audio ML; expertise in PyTorch and modern sequence models (transformers, conformers, RNN-T); experience with streaming or low-latency inference.

Preferred: self-supervised speech representation learning (wav2vec/HuBERT-style), text-to-speech, or multilingual/low-resource ASR.

Application

Apply for Research Scientist — Speech & Audio Understanding

Zero

Zero Hiring processes the information you provide (including your resume and interview responses) on behalf of Zero to assess your application. Your application may be screened and scored using automated / AI systems. We keep your data as described in our Privacy Policy and you can exercise your privacy rights at privacy@zerohiring.com.

Some screening decisions are made with the help of automated systems. You have the right to request human review of a decision that affects you — contact privacy@zerohiring.com.

Read our Privacy Policy

Research Scientist — Speech & Audio Understanding at Zero