Sr Manager, Siri Agentic Evaluation
Generate a McCoy IQ challenge in 30 seconds.
See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.
Key details
What makes this role novel
Agentic evaluation—systematically assessing the behavior, safety, and reliability of autonomous AI agents that take actions on behalf of users—is a specialized discipline that emerged only as LLMs and multimodal models became capable enough to act independently. This role focuses on evaluating agent behavior at scale, which is distinct from traditional ML evaluation or QA.
Job Description
Join the team redefining what a deeply personal and integrated assistant can be.
As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.
This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.
Audit details(provenance, verification trail, raw fields)
Core fields
apple:200673865-1251Provenance
appleVerification trail
This posting hasn't been probed by our closure verifier yet. Stream C runs on a rolling schedule against postings approaching the close-decision threshold.
LLM enrichment
See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.
