Machine Learning Engineer - AI Evaluation & LLM Systems

Back to Jobs

Machine Learning Engineer - AI Evaluation & LLM Systems

Apple

Location

Cupertino, California, 95014, United States

Experience

Entry

Posted

Jul 30, 2026

Apply by

August 29, 2026

Applicants

0

Early applicantEasy applyFull-timeWork from Office

Job Description

Join the team building the evaluation systems that enable Apple’s next generation of AI experiences. As a Machine Learning Engineer, you will develop scalable infrastructure, intelligent evaluators, and data-driven methodologies that measure and improve the quality of large language models and multimodal AI systems used across Apple products. You’ll partner closely with ML researchers, software engineers, and product teams to design novel evaluation techniques, analyze model behavior, and translate research into production-ready systems. This role requires strong engineering fundamentals, a passion for machine learning, and the curiosity to solve challenging problems at the intersection of AI, data, and software engineering. If you’re excited about building the tools that help define the future of AI quality at Apple, we’d love to hear from you. ## Description As a Machine Learning Engineer, you will build the systems that measure and improve the quality of AI experiences used by millions of people. You will develop machine learning models, evaluation frameworks, and scalable infrastructure that enable teams to understand model behavior, identify regressions, and accelerate the development of large language models and multimodal AI. Working closely with researchers, software engineers, and product teams, you will transform cutting-edge research into production-ready solutions, analyze large-scale datasets, and develop new approaches for benchmarking and improving AI quality. This is a unique opportunity to solve challenging technical problems at the intersection of machine learning, software engineering, and data, while helping shape the future of AI at Apple. ## Responsibilities Design, develop, and deploy evaluation systems and scalable software that improve the quality of AI experiences. Build robust infrastructure to support model training, benchmarking, and large-scale evaluation. Analyze model performance, identify quality issues, and develop innovative techniques to measure and improve AI behavior. Collaborate with machine learning researchers, software engineers, and product teams to translate research into production-ready solutions. Drive technical excellence by contributing to architecture, code reviews, experimentation, and engineering best practices. ## Minimum qualifications MS, or PhD in Computer Science, Machine Learning, Electrical Engineering, or a related technical field, or equivalent practical experience. 1–2 years of industry experience, or equivalent academic or internship experience, developing machine learning or AI solutions. Proficiency in Python and familiarity with C++ or another object-oriented programming language. Experience with one or more machine learning frameworks such as PyTorch, TensorFlow, or JAX. Understanding of machine learning fundamentals, including supervised learning, model evaluation, and statistical analysis. Experience working with data processing, model training, or experimentation through coursework, research, internships, or industry projects. Strong analytical, problem-solving, and communication skills with the ability to collaborate effectively in a team environment. ## Preferred qualifications Experience with large language models (LLMs), multimodal AI, or generative AI through internships, research, or personal projects. Experience building software or machine learning projects using modern engineering practices (Git, testing, CI/CD). Familiarity with distributed computing, cloud platforms, or large-scale data processing. Publications, open-source contributions, or participation in machine learning competitions. MS or PhD specializing in Machine Learning, Artificial Intelligence, or a related field.

Key Responsibilities

  • Design, develop, and deploy evaluation systems and scalable software to improve AI quality.
  • Build robust infrastructure for model training, benchmarking, and large-scale evaluation.
  • Analyze model performance and develop techniques to measure and improve AI behavior.
  • Collaborate with researchers and engineers to translate research into production-ready solutions.
  • Contribute to architecture, code reviews, and engineering best practices.

Requirements

  • MS or PhD in Computer Science
  • Machine Learning
  • Electrical Engineering
  • or a related technical field

Skills Required

PythonC++PyTorchTensorFlowJAXSupervised learningModel evaluationStatistical analysisData processingModel trainingAnalytical skillsProblem-solvingCommunicationCollaborationLarge language models (LLMs)Multimodal AIGenerative AIGitTestingCI/CDDistributed computingCloud platformsLarge-scale data processing

App exclusive · Free

Smart Job AI Coach

Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.

Interview readiness

See how prepared you are and what to improve for each role.

Personalized tips

Actionable suggestions based on your profile and the job.

After you apply

Keep coaching momentum from job detail through application success.

Get Smart Job AI Coach in the appFree on iOS and Android