Machine Learning Engineer - Agentic AI Evaluation Frameworks
Generate a McCoy IQ challenge in 30 seconds.
See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.
Key details
What makes this role novel
Evaluation frameworks and quality signals for generative AI and LLM systems represent a distinct job category that emerged only in the last 2–3 years as companies scaled foundation models; the work of designing datasets, metrics, and tooling specifically to measure and improve AI model outputs at production scale did not exist as a standalone discipline before the LLM era.
Job Description
Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish.
The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.
In this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle.
You will help define how we measure the quality of AI experiences across the Commerce domain, including Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI, and Platform Self-Service.
This is an opportunity to work at the intersection of machine learning, software engineering, data, and product quality, helping ensure our AI experiences are accurate, relevant, grounded, reliable, and useful for users around the world.
Audit details(provenance, verification trail, raw fields)
Core fields
apple:200681825-0836Provenance
appleVerification trail
This posting hasn't been probed by our closure verifier yet. Stream C runs on a rolling schedule against postings approaching the close-decision threshold.
LLM enrichment
See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.
