Applied AI Researcher, Agent Systems & Evaluation
- Eval data collection. Where ground truth comes from. Mining production traces for labeled outcomes, capturing human accept/reject/edit signal as it happens, building task sets that reflect the real distribution of work rather than the tasks that are easy to score, and knowing when a model-based judge is trustworthy and when it is laundering an assumption.
- Eval loop construction. Turning a fuzzy objective into a measurement that runs on every change. Noise floors, statistical standards for acceptance, task suites that resist gaming, and experiments designed for production settings where clean randomization is not always available.
- Automated hill climbing. The payoff. Once a task has a trustworthy loop, improvement can be searched rather than hand-crafted — prompts, context strategies, tool sets, routing, reasoning budgets, eventually model choice. This only works if the first two stages are sound; done wrong, it optimizes hard against a metric that means nothing. Post-train models on data nobody else has. This role includes hands-on model work: supervised fine-tuning and RL on open-source vision-language models, using the proprietary driving data Nuro has collected across years of real-world autonomous operation. You would have a labeling workforce available to you, which means you can specify the data you need rather than making do with what exists. Very few researchers get to run this loop, form a hypothesis about model behavior, commission the exact data to test it, post-train, and evaluate against real driving performance. The fungibility of frontier models is precisely why this matters: the weights are rentable, the data and the labeling capacity behind them are not.
- Establish the evaluation foundation for the agent fleet we already run: eval data sources, task suites, noise floors, and the statistical standard the team uses to accept or reject a change.
- Take one high-volume workflow from unmeasured to automatically hill-climbing, end to end, as the template the rest of the system follows.
- Run a first post-training experiment on an open-weight VLM against our driving data, and establish whether the result justifies the pipeline.
- Put a defensible number on what the platform is worth: which workflows improved, by how much, with what confidence.
- Graduate degree in CS, ML, statistics, or a related field, or equivalent research experience. We care about demonstrated research judgment, not credentials.
- Fluent in the current literature and able to judge it. You read papers continuously, can tell a real result from a well-marketed one, and have opinions about which recent directions are overrated.
- Deep understanding of how LLMs work — pretraining through the post-training stack, and what actually happens at inference. You reason from mechanism, not just from published numbers.
- Firm grasp of the full evaluation pipeline: sourcing eval data, constructing the loop, automating the climb. Having done all three for a real system, rather than one in isolation, is the strongest signal for this role.
- Rigorous experimentalist. You design experiments that can fail, you understand variance and power, and you are comfortable saying an intervention didn't work.
- Hands-on post-training experience — SFT and RL, ideally on open-weight models — including the data curation and evaluation work required to know whether it actually helped. Vision-language model experience is a strong plus.
- A real engineering background. Strong Python, comfortable with production systems and data, able to stand up the infrastructure your own experiment needs.
- Direct experience with LLM agent systems — building them, evaluating them, or studying why they fail.
- You measure yourself in impact and in weeks. You want your work in front of hundreds of engineers this quarter.
- Published or applied work in agent evaluation, reasoning, test-time compute, RL, or verification.
- Experience with multimodal or vision-language models, and with data curation at scale.
- Online experimentation in production: A/B testing, causal inference from observational data, offline-to-online correlation.
- Experience building evaluation harnesses, task suites, or LLM-as-judge systems, including their failure modes.
- Familiarity with autonomous systems, safety cases, or verification-gated deployment.
Recommended Jobs
Licensed Claims Adjuster: Flexible Career & Training
A leading claims adjusting firm in California is seeking Independent Insurance Claims Adjusters. This role provides a rewarding career path where you help others recover from disasters. As a Licensed…
Retail Support Associate - Shoe Expeditor, Century City - Part Time
Be part of an amazing story Macy’s is more than just a store. We’re a story. One that’s captured the hearts and minds of America for more than 160 years. A story about innovations and traditions…abo…
News Photographer
OVERVIEW OF THE COMPANY Fox TV Stations FOX Television Stations owns and operates 29 full power broadcast television stations in the U.S. These include stations located in 14 of the top 15 larg…
Guest Service Shift Leader
: In accordance with California law, the expected salary range for this California position is between $20.00 and $22.00 an hour. The actual compensation will be determined based on experience and o…
Databricks Engineering Consultant
Our Deloitte AI & Engineering team to transform technology platforms, drive innovation, and help make a significant impact on our clients' success. You'll work alongside talented professionals reimagi…
Pathologists' Assistant (PA)
Overview Ansible Government Solutions, LLC (Ansible) is currently recruiting Pathologists' Assistants to support the VA San Diego Healthcare System located at 3350 La Jolla Village Dr, San Diego, …
Field Applications Engineer
Are you ready to power the future? At SolarEdge (NASDAQ: SEDG), we're a global leader in smart energy technology, with over 4,000 employees, offices in 34 countries, and millions of installations wor…
MOHS Physician opportunity with ...
Leading the future of health care Kaiser Permanente / The Permanente Medical Group The Permanente Medical Group, Northern California, (TPMG) is one of the largest medical groups in the nation wit…
Line Cook
Dinner Line Cook / Kitchen Team Hours: Tuesday - Saturday, 230pm - 10pm Starting wage: $20-23/hr, depending on experience The Dinner Line Cook will accurately and efficiently cook meats, fish…