Lead, Product Content Engineering
Generate a McCoy IQ challenge in 30 seconds.
See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.
Key details
What makes this role novel
This role combines systematic human evaluation design for generative media with direct model training influence—turning creative judgment into reproducible training signals for image/video generation. The day-to-day work (building eval rubrics, running calibrated rater pools, arbitrating quality standards that shape model behavior) is a specialized function that emerged only as generative media models became central to product roadmaps in the last 2–3 years.
Job Description
Product Content Engineering is a horizontal function supporting initiatives across Instagram and Meta AI, and is seeking a content engineer to lead media generation evaluations. We partner closely with product, research and modeling teams to set standards of quality and define what “good” looks like for generative image and video experiences. This position is focused on model training for image and video generation capabilities. You will own the human evaluation layer that tells our models what quality means, designing the rubrics, running the eval rounds, and turning subjective creative judgment into reproducible, measurable signals that directly shapes model behavior and launch decisions. The ideal candidate has built video and image evaluation rubrics from scratch and has hands-on video production expertise — they can articulate why a generated clip fails on motion coherence, lighting continuity, subject consistency, pacing, framing or audio sync, and can write that judgment down so that dozens of raters apply it the same way. They are fluent in text-to-video, image-to-video and video-to-video generation, and in the failure modes specific to each. You will operate as the quality authority across multiple modeling workstreams: setting the taxonomy, arbitrating disagreement, calibrating rater pools, and translating eval results into clear guidance for researchers and product leaders. You will also drive where human evaluation should give way to automated or LLM-as-a-judge measurement, and prove out that transition with data.
Responsibilities
Own the evaluation rubrics for image and video generation models, define the quality dimensions, scales and decision rules, and maintain them as models and capabilities evolve. Design and run human eval rounds that produce training and fine-tuning signal, including prompt set construction, golden sets and side-by-side model comparisons. Set and defend the launch quality bar for generative media features, and make clear go / no-go quality recommendations. Calibrate and manage rater pools and vendor teams; monitor inter-rater reliability and drive measurable improvements in agreement. Partner with research and modeling teams to translate eval findings into model, data and post-training priorities. Scale human judgment into automated measurement (LLM-as-a-judge, auto-metrics) and validate automated scores against human ground truth. Lead through collaboration, mentoring other content engineers on eval craft and rubric design. Stay ahead of the generative media landscape, benchmarking against competitive and open-source model output. Deliver high-quality evaluation work on model and launch timelines in a fast-paced, dynamic environment.
Qualifications
10+ years of experience in video production, content strategy, editorial standards, media evaluation or relevant fields Experience authoring evaluation rubrics for video and/or image quality, including dimension design, rating scales, tie-breaking rules and annotator guidelines Hands-on video production expertise, direction, shooting, editing, post and/or VFX, with the craft vocabulary to diagnose quality failures precisely Experience evaluating generative media output (text-to-video, image-to-video, video-to-video, text-to-image) and diagnosing model failure modes Experience partnering with AI/ML research or modeling teams, translating evaluation results into training, data or model-behavior recommendations Experience running human evaluation programs at scale: rater calibration, inter-rater reliability, golden sets, quality audits Experience leveraging quantitative insights to inform content and quality decisions Proven ability to build influence and drive alignment across product, engineering, design, research and analytics partners Experience leading and/or mentoring others Bachelor's degree in a related field, or equivalent practical experience Experience building an evaluation function or quality framework from zero-to-one Background in film, television, OTT, commercial or VFX production Familiarity with generative media tooling and model families, and with prompt engineering for media generation Experience with annotation platforms, eval tooling and data pipelines supporting model training Experience working with product teams or programs from roadmapping through delivery Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies Experience working independently and navigating ambiguity in fast-paced environments
Audit details(provenance, verification trail, raw fields)
Core fields
meta:2867963683575119Provenance
metaVerification trail
This posting hasn't been probed by our closure verifier yet. Stream C runs on a rolling schedule against postings approaching the close-decision threshold.
LLM enrichment
See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.
