AI Engineer, Internal Systems
Generate a McCoy IQ challenge in 30 seconds.
See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.
Key details
What makes this role novel
Building verification and evaluation layers for autonomous AI agents to self-assess task completion is a distinct new discipline that emerged with recent advances in agentic AI systems; this work didn't exist as a standalone role before ~2023-2024.
Job Description
About Wispr
Wispr is an AI research and product company building the voice interface for computing. AI can now reason, code, and act. Yet, humans still do the work of the interface. We think that’s backwards.
Our first products are Flow, which lets you speak naturally in any application, and Notetaker, which builds context across conversations. We’re building toward an interface that can perceive, understand, and take action with earned trust. That means solving hard problems across models, systems, and product - and caring about the final human experience as deeply as the technology underneath it.
We’re a talent-dense team that holds strong opinions, tests them quickly, and builds technology that sparks joy. Our goal is to build the first voice interface used every day by a billion people.
About the role
Our agents already write code and respond to incidents around the clock, but they can't yet tell when they're done, so every loop still ends with a human checking. You'd be the first to build the verification and evaluation layer that lets them finish with confidence.
What you'll do
Build end-to-end verification loops that let agents determine when their work is actually complete
Create fast, reliable testing environments and turn failures and human feedback into durable evaluation signals
Improve how context, instructions, skills, and memory are delivered to agents—and measure what makes them perform better
Operate the agent fleet as a production system and measure its quality, adoption, and impact
You may be a fit if
You've built internal tools that other engineers adopted, and you can explain how you measured their impact
You use agentic coding tools deeply and have specific opinions about where they succeed and fail
You've built evaluation or verification infrastructure such as test harnesses, CI systems, eval pipelines, or benchmarks
You can work hands-on across application code, infrastructure, and unfamiliar systems
You are empirical about your own work, attentive to subtle failure modes, and comfortable owning ambiguous problems
Prior experience in voice, model training, or our product domains is not required.
Get to know us
Logistics
We sponsor H1B, O1, EB1, L1, STEM OPT, and more. We can't guarantee sponsorship for every role, but if we make you an offer we'll make every reasonable effort, with help from our immigration law firm.
We strongly encourage you to apply even if you don't meet every qualification. The strongest candidates we meet rarely do, so don't exclude yourself prematurely. We're rethinking how humans interact with AI, and doing that well demands diversity of perspective and experience.
We're committed to a fair and accessible interview process. If you need any accommodations or adjustments, please let us know.
Audit details(provenance, verification trail, raw fields)
Core fields
wispr_flow:96beb394-070d-4753-8cd5-5928a2657b3fProvenance
wispr-flowVerification trail
This posting hasn't been probed by our closure verifier yet. Stream C runs on a rolling schedule against postings approaching the close-decision threshold.
LLM enrichment
See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.
