Apple is seeking a Senior Machine Learning Engineer to build intelligent search and proactive AI experiences across its ecosystem. This role focuses on designing, training, and deploying transformer-based large language models, semantic retrieval, and ranking systems for on-device and personalized search. The engineer will research model compression, quantization, and low-latency inference, and partner with cross-functional teams to ship AI features from research to production.
Responsibilities
Design, train, fine-tune, and optimize transformer-based language models for on-device deployment.
Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems to improve search quality and AI-powered experiences.
Develop models for query understanding, intent prediction, personalization, and ranking.
Research new approaches to model compression, quantization, and low-latency inference.
Partner with engineers, researchers, product managers, and designers to bring AI capabilities from research into production, driving technical strategy and leading projects from exploration to large-scale deployment.
Requirements
Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
5+ years of industry or research experience developing machine learning systems.
Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
Apple designs consumer electronics, software, and related services. Apple’s corporate communications highlight its large global workforce and its operations centered around Apple Park in Cupertino.
Experience training, fine-tuning, or deploying transformer-based models and large language models.
Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow.
Nice to Have
Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
Experience optimizing machine learning models for resource-constrained environments, including model compression, quantization, pruning, and knowledge distillation.
Experience with on-device machine learning or mobile inference frameworks.
Experience building retrieval-augmented generation, vector search, embedding retrieval, or semantic search systems.
Experience working with transformer architectures such as BERT, T5, Llama, Gemma, Mistral, or other foundation models.
Experience evaluating language models, designing AI quality metrics, and building offline evaluation pipelines.
Experience building large-scale production search, recommendation, or personalization systems.