Staff Research Engineer, Agent Evals & Post-training
Generate a McCoy IQ challenge in 30 seconds.
See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.
Key details
What makes this role novel
Agent evals and post-training for hardware-reasoning systems is a specialized subcategory of AI safety/capability work that emerged only in the last 2–3 years as LLM agents became viable for domain-specific reasoning tasks; the day-to-day work—designing evaluations for agentic behavior in physical domains and iterating on model outputs for hardware contexts—did not exist as a distinct role before ~2022.
Job Description
About Nominal
About Hardware Intelligence
The Role:
💼 What You'll Do:
- Build eval suites for our agents, the MCP tool layer, and our internal company agent, grounded in real hardware tasks.
- Invent new benchmarks for agents working over physical engineering data, where none exist today.
- Benchmark our agents against frontier agents using our MCP, and show where and why ours win.
- Build the eval infrastructure: datasets, model-based and human graders, and regression gates in CI.
- Turn evals into reward signals and training data, and lead post-training (fine-tuning, RL, distillation) when we need models we can run anywhere, including air-gapped environments.
- Set Nominal's strategy for measurement and model improvement, working with hardware experts on what "correct" means.
🚀 What you'll bring:
- 8+ years in ML engineering or research, including evals or post-training work you led in production.
- Statistical rigor: you design evals that don't fool you, and you know when a difference is real.
- Deep experience evaluating LLM or agent systems, including model-graded evals and their limits.
- Hands-on post-training experience: fine-tuning, RL from feedback, or distillation on real tasks.
- The judgment to know when to measure, when to train, and when a better prompt or tool is the answer.
- A track record of setting technical direction across a team and raising the bar for the engineers around you.
- You build with modern AI coding agents (Claude Code, Cursor, Codex) every day, and stay curious and open to better ways of working. The tools keep changing, and so do we.
⚡️ Nice to have:
- You've built evals or post-training at a frontier lab or an AI-native company.
- You've published benchmarks or eval methods that others use.
- You've trained or served open-weight models in restricted, on-prem, or air-gapped environments.
- You've worked in test, reliability, or verification engineering for physical systems.
Benefits/Perks
- 🏥 100% coverage of medical, dental, and vision insurance
- 🏖️ Unlimited PTO and sick leave
- 🍽️ Free lunch, snacks, and coffee
- 🚀 Professional Development Stipend
- ✈️ Annual company retreat
Audit details(provenance, verification trail, raw fields)
Core fields
nominal:am9icG9zdDp7LtWwyt-ZLX3N-M7SHGRVProvenance
nominalVerification trail
This posting hasn't been probed by our closure verifier yet. Stream C runs on a rolling schedule against postings approaching the close-decision threshold.
LLM enrichment
See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.
