Product Manager - Post Training
Generate a McCoy IQ challenge in 30 seconds.
See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.
Key details
What makes this role novel
Post-training as a distinct product discipline—managing the intersection of RLHF, evals, model behavior, and inference optimization as a coherent product function—is a role that emerged only with large-scale LLM development in 2023–2024. The work of translating post-training research into product decisions and infrastructure priorities didn't exist as a standalone PM function before frontier models.
Job Description
About Thinking Machines
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
About the Role
This is a product role for someone who can operate inside frontier research without trying to turn research into a conventional software roadmap. The work requires judgment, technical fluency, user empathy, discretion, and the ability to create clarity without creating bureaucracy.
The Post-Training Product Manager will work as a high-trust partner to our post-training researchers. Your role is to understand the work deeply enough to ask the right questions, identify missing connections, surface implications, and help the team decide what matters next. The role sits at the seam between research, model behavior, data and environments, evaluations, training and inference infrastructure, safety, product, and the people using our models.
Inkling was post-trained across math, agentic code and tool use, audio, image, chat, and safety, with a large-scale asynchronous RL program that exceeded 30M rollouts. Inkling and Inkling-Small share a scalable post-training stack, and Inkling is available for customization on Tinker.
The next phase requires tight loops among research priorities, model behavior, evaluations, infrastructure constraints, user and customer learning, and product direction. The person in this role will help those loops compound instead of fragment — maintaining the bird's-eye view while researchers go deep: where work is converging, where teams are solving adjacent problems without enough shared context, what evidence is missing, what decisions are blocked, and how research becomes a stronger model and a useful product.
What You'll Do
Partner with post-training research leaders on priorities, sequencing, decision points, and the connection between research work and the broader model and product agenda
Maintain a clear view across SFT, RL, data and environments, evaluations, safety, model behavior, inference, training infrastructure, and product dependencies; identify gaps before they become blockers
Translate ambiguous research and product questions into concrete learning plans: what must be true, what evidence would change the decision, which experiments or user signals matter, and when the team should revisit the direction
Build lightweight operating mechanisms for critical work — owners, state, dependencies, decisions, risks, release criteria, and follow-through — without imposing a software-development process on research
Bring qualitative and quantitative evidence about model behavior and real workflows into research prioritization, ensuring user knowledge is represented accurately rather than flattened into feature requests
Connect research, product, infrastructure, safety, and leadership when a decision spans teams or when local optimization creates a broader product or model tradeoff
Support the path from research result to usable capability: internal adoption, evaluation, documentation, release readiness, product integration, and feedback after launch
Write clear narratives that explain what the team has learned, what remains uncertain, what decisions are needed, and why the work matters
Take on the unowned work that is necessary to move a critical research-product outcome forward
Skills and Qualifications
Minimum qualifications:
Experience as an early or first product leader in a technical startup, AI lab, research organization, or new product area where the role and operating model were not defined for you
Experience working closely with model training, post-training, RL, evaluations, data, safety, inference, developer platforms, or another technically demanding research-product area
Ability to understand research deeply enough to earn trust, ask sharp questions, and connect technical choices to model behavior and users without overstating your expertise
Strong track record identifying the missing connection, unresolved assumption, or cross-team decision that specialists may not see while deep in the work
Preferred qualifications:
Ability to represent user and product truth in a research environment without reducing research to a list of customer requests
Comfortable making progress when the goal, metric, or path is still evolving, and the correct next step may be a learning loop rather than a launch
Clear communicator who handles disagreement without ego and is willing to change a recommendation when the evidence changes
Motivated by senior IC ownership and proximity to the work more than a large PM team or a conventional product ladder
Background as a research product leader or early AI product leader at a frontier lab, model company, or technically ambitious startup; a technical founder, former engineer, or applied scientist who moved into product; a PM for model training, post-training, evaluation platforms, data systems, ML infrastructure, or developer platforms; or the first PM at a company that turned a novel technical capability into a product, category, or developer ecosystem
Logistics
Location: This role is based in San Francisco, CA.
Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $450,000 USD.
Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Audit details(provenance, verification trail, raw fields)
Core fields
thinking-machines:1b503b26-dd56-4496-8f74-2c8abb3b7e4bProvenance
thinkingmachinesVerification trail
This posting hasn't been probed by our closure verifier yet. Stream C runs on a rolling schedule against postings approaching the close-decision threshold.
LLM enrichment
See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.
