
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("Market Average: Senior Level"), this listing's salary midpoint is about 89% lower. The offer sits below the benchmark range (€10,413–€13,916). The offer's range width is broadly in line with the benchmark. This benchmark is based on 1 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| Market Average: Machine Learning Engineer | €11,159/per month | €16,639/per month | €21,189/per month |
Job Description The Role Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better? That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing. This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments. What you get to do in this role: Judge design and calibration A shared base judge with per-item rubrics expressed as configuration next to the dataset — so eval authors express intent, rather than forking a prompt per eval Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy — was the clarifying question appropriate, was policy followed, was the path efficient Scoring that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team — they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job Fine-tuning a small judge model where an off-the-shelf one isn't good enough Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit Self-learning for the agent harness This is where the pillar is headed, and a large part of why the seat exists. A calibrated trajectory judge is, functionally, a reward model. Turning ours into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock Using that signal to optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success
Job Details
Responsibilities
- Design and calibrate judges for agent evaluation
- Develop deterministic validators for state checks
- Create scoring with confidence estimates
- Maintain calibration loops with human labeling
- Fine-tune small judge models
- Guard against blind spots
- Support offline/online divergence analysis
Requirements
- 8+ years in applied ML, data science, or ML-adjacent engineering
- Experience turning subjective human judgement into measurable signals
- Strong Python
- Ability to communicate complex problems
- Ownership mindset
- Experience in at least 3 domains listed in the posting
Skills & Technologies
Education Level
No degree required
| Location | Active listings |
|---|---|
| Remote - Global | 4 |
| Waltham, United States | 1 |
| Washington, United States | 1 |
| Helsinki, Finland | 1 |
| Nashua, United States | 1 |
| West Palm Beach, FL, United States | 1 |
| Role type | Active listings |
|---|---|
| Global Partner Manager | 2 |
| Senior Advisory Solution Consultant | 1 |
| Director, Product Innovation | 1 |
| Platform Architect | 1 |
| Role level | Active listings |
|---|---|
| Senior | 4 |
| Executive | 1 |
Related Opportunities
Discover more opportunities that match your interests and skills