
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("All roles in Remote - Germany"), this listing's salary midpoint is about 91% lower. The offer sits below the benchmark range (€4,851–€7,466). The listed pay band (€449–€602) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 4 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| All roles in Remote - Germany | €4,851/per month | €6,310/per month | €7,466/per month |
Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners. Mindrift, powered by Toloka — a leading enterprise AI and machine learning data partner since 2014 — connects top domain experts with cutting-edge AI initiatives. Backed by Toloka’s deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform. We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples — from problem structuring and work planning to analysis, synthesis, and client-ready recommendations. You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning. Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery. Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC). Who We’re Looking For Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in: Structuring ambiguous client problems into workable analytical plans Building financial models, market analyses, or synthesized findings from messy inputs Producing client-ready deliverables under time pressure Forming and defending recommendations under uncertainty No deep technical background is required — we will onboard you on the lightweight tools involved. What You’ll Do Build realistic consulting project environments — create detailed project scenarios grounded in real engagement dynamics: industry context, financials, constraints, conflicting inputs, and incomplete information. Design structured consulting tasks for AI agents — break projects into discrete tasks that mirror real consulting work: market sizing, commercial due diligence, cost optimization, growth strategy, operational diagnosis, benchmarking, and more. Define evaluation criteria and quality standards — develop grading frameworks, evaluation rubrics, and golden-answer solutions for each task, used to train and calibrate an LLM-based grading system that evaluates AI outputs at scale. This is a remote, project-based, individual-contributor role focused on analytical design and evaluation. Skill & Requirements 3+ years etc.
Job Details
Responsibilities
- Build realistic consulting project environments — create detailed project scenarios grounded in real engagement dynamics: industry context, financials, constraints, conflicting inputs, and incomplete information.
- Design structured consulting tasks for AI agents — break projects into discrete tasks that mirror real consulting work: market sizing, commercial due diligence, cost optimization, growth strategy, operational diagnosis, benchmarking, and more.
- Define evaluation criteria and quality standards — develop grading frameworks, evaluation rubrics, and golden-answer solutions for each task, used to train and calibrate an LLM-based grading system that evaluates AI outputs at scale.
Requirements
- 3+ years at McKinsey, BCG, Bain, Oliver Wyman, Roland Berger, Monitor Deloitte, EY-Parthenon, Kearney, or Strategy&
Skills & Technologies

| Location | Active listings |
|---|---|
| Remote - Belgium | 2 |
| Remote - Denmark | 2 |
| Remote - Sweden | 2 |
| Remote - Finland | 2 |
| Remote - Global | 2 |
| Remote - Germany | 2 |
| Remote - France | 2 |
| Role type | Active listings |
|---|---|
| Brand Designer | 1 |
| Algorithm Problem Setter | 1 |
| Presentation Designer | 1 |
| Python Data Scraping Engineer | 1 |
| Role level | Active listings |
|---|---|
| Senior | 2 |
| Intern | 1 |
| Mid-Level | 1 |
Related Opportunities
Discover more opportunities that match your interests and skills