Senior Machine Learning Engineer, Proactive
Back to JobsApple Get Smart Job AI Coach in the appFree on iOS and Android
Senior Machine Learning Engineer, Proactive
Location
Santa Clara, California, 95050, United States
Experience
Senior
Posted
Jul 30, 2026
Apply by
August 29, 2026
Applicants
0
Early applicantEasy applyFull-timeWork from Office
Job Description
At Apple, machine learning powers experiences that anticipate what people need before they ask. We're looking for a Machine Learning Engineer to help build the next generation of intelligent search and AI experiences technology that understands user intent, context, and personal information while preserving privacy. In this role, you'll design, train, optimize, and deploy large language models, semantic retrieval systems, and ranking models that power relevant, personalized, context-aware search across Apple's ecosystem. You'll work at the intersection of search, retrieval, natural language processing, on-device AI, and generative AI to shape the future of intelligent assistants and proactive experiences..
## Description
You'll design, train, fine-tune, and optimize transformer-based language models for on-device deployment, and build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems that improve search quality and AI-powered experiences. You'll develop models for query understanding, intent prediction, personalization, and ranking, while researching new approaches to model compression, quantization, and low-latency inference. You'll partner with engineers, researchers, product managers, and designers to bring new AI capabilities from research into production driving technical strategy and leading projects from early exploration through large-scale deployment. This is an opportunity to explore new applications of foundation models, multimodal AI, and agentic retrieval, shaping the next generation of proactive, intelligent user experiences.
## Responsibilities
Design, train, fine-tune, and optimize transformer-based language models for efficient on-device deployment.
Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, along with models for query understanding, intent prediction, personalization, and ranking.
Research and prototype approaches for on-device generative AI, including model compression, quantization, knowledge distillation, and low-latency inference.
Analyze search relevance and user behavior to design evaluation methodologies, offline benchmarks, and online metrics that measure retrieval quality, ranking, and language model performance.
Partner with engineers, researchers, product managers, and designers to bring AI capabilities from research into production, driving technical strategy across projects and exploring new applications of foundation models, multimodal AI, and agentic retrieval.
## Minimum qualifications
Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
5+ years of industry or research experience developing machine learning systems.
Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
Experience training, fine-tuning, or deploying transformer-based models and large language models.
Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow.
## Preferred qualifications
Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
Experience optimizing machine learning models for resource-constrained environments, including model compression, quantization, pruning, and knowledge distillation.
Experience with on-device machine learning or mobile inference frameworks.
Experience building retrieval-augmented generation, vector search, embedding retrieval, or semantic search systems.
Experience working with transformer architectures such as BERT, T5, Llama, Gemma, Mistral, or other foundation models.
Experience evaluating language models, designing AI quality metrics, and building offline evaluation pipelines.
Experience building large-scale production search, recommendation, or personalization systems.
Ability to prototype ideas, solve ambiguous problems, and deliver production-quality machine learning solutions.
Key Responsibilities
- Design, train, fine-tune, and optimize transformer-based language models for on-device deployment.
- Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems.
- Research and prototype approaches for on-device generative AI, including model compression and low-latency inference.
- Analyze search relevance and user behavior to design evaluation methodologies and metrics.
- Partner with cross-functional teams to bring AI capabilities from research into production.
Requirements
- Bachelor's degree in Computer Science
- Machine Learning
- Artificial Intelligence
- or a related field
Skills Required
PythonC/C++PyTorchJAXTensorFlowTransformer-based modelsLarge language modelsNatural language processingInformation retrievalSearchRecommender systemsGenerative AICollaborationTechnical strategyProject leadershipModel compressionQuantizationPruningKnowledge distillationOn-device machine learningMobile inference frameworksVector searchEmbedding retrievalSemantic search systemsBERTT5LlamaGemmaMistralFoundation modelsOffline evaluation pipelinesLarge-scale production search systemsRecommendation systemsPersonalization systemsPrototypingProblem solving
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.


