Centific AI Research is seeking an AI Engineer specializing in Speech/Audio to develop Large Audio Language Models (LALMs), Speech-to-Speech systems, and audio reasoning models. The role combines research and production, requiring expertise in speech representation, LLM fine-tuning, and audio pipeline optimization.
Responsibilities
Design, develop, and deploy Large Audio Language Models (LALMs) for native audio understanding, reasoning, and generation.
Build Large Audio Reasoning Models for complex chain-of-thought reasoning over speech and audio inputs.
Contribute to Speech-to-Speech system development including speech understanding, dialogue management, and speech synthesis.
Implement alignment mechanisms between speech encoders and LLM backbones using adapters, LoRA, and fine-tuning strategies.
Design speech tokenization and temporal compression for long-form audio and multi-turn dialogue.
Build evaluation frameworks for audio reasoning capabilities and benchmarks.
Optimize inference pipelines for low-latency streaming applications.
Collaborate with cross-functional teams to transfer research into production systems.
Contribute to technical documentation and publications at top venues.
Requirements
Master's degree (required) or Ph.D. (preferred) in CS, EE, or related field with focus on speech/audio ML or multimodal learning.
WFH.team provides remote job intelligence for candidates and employers, offering confirmed remote job listings and resume-based matching. They also provide various employer-facing hiring tools and public resources for remote work workflows.