Mindrift logo
Est. Monthly
Estimated €5,500 - €7,125
Posted August 24, 2026 · 5 days agoLast seen August 26, 2026Est. expiry September 28, 2026

Software Engineering Evaluation Specialist

How this salary compares
Salary Context: Software Engineering Evaluation Specialist

Hover or tap a row for full statistics (EUR / month on this chart).

Salary analysis

Compared with the selected benchmark ("Company in Remote - Germany"), this listing's salary midpoint is about 92% lower. The offer sits below the benchmark range (€6,945–€6,945). Range-width comparison is limited because one of the salary bands is incomplete. This benchmark is based on 1 comparable listings.

Monthly salary comparison for Software Engineering Evaluation Specialist
MarketLower bound (25th percentile)MedianUpper bound (75th percentile)
All roles in Remote - Germany€4,912/per month€6,313/per month€7,447/per month
Company in Remote - Germany€6,945/per month€6,945/per month€6,945/per month
About the role

Please submit your CV in English and indicate your level of English proficiency.Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.About the RoleYou’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.Responsibilities: - Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem. - Build a reproducible Docker environment with pinned dependencies. - Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix. - Write an instruction.md that reads like a Jira ticket a developer would receive. - Write a reference solve.sh proving the task is solvable. - Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time. - Iterate based on feedback from expert QA reviewers. - Later: review other authors’ tasks as a QA reviewer.Not in scope - Data labeling, prompt engineering.- Production code to ship — you design problems and verification for AI agents.- Leetcode puzzles — scenarios must look like real developer work.- Not every candidate task ships — quality over quantity.Requirements- 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth.- Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py.- Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.- Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.- AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it.- English — B2+ written.Not a fit- Data Science, ML, or Computer Vision engineers without backend-engineering output.- Manual QA testers without automation or test authoring.- Frontend-only, low-code / no-code, IT Support, or Business Analysts.- Engineers who have never written pytest from scratch.- Junior, intern, or assistant as the most recent role.Preferred qualifications- Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.- Modern Python tooling (uv, poetry, pyproject.toml).- Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov).- Fuzzing or property-based testing (Hypothesis).- Prior contribution to agent-evaluation benchmarks or related frameworks.Process- Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.Time commitment- Onboarding: ~10 hours per first task.- Steady state: ~5 hours per task, 2–4 parallel tasks per author.- Realistic weekly load: 8–20 hours. Higher volume available for top performers.- You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.Compensation:- Paid contributions, rates up to $35/hour*.- Task-based compensation equivalent to hourly rate, depending on performance and volume.- Some projects include incentive payments.*Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project.Apply- Submit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.

Job Details

Seen 3 days agoPartial Schema
Mindrift logo
Mindrift · 14 open roles
Top locations: Remote - Belgium · 2 · Remote - Denmark · 2 · Remote - Sweden · 2+4 other locations
View company
Current open roles at Mindrift on JobCrawls
LocationActive listings
Remote - Belgium2
Remote - Denmark2
Remote - Sweden2
Remote - Finland2
Remote - Global2
Remote - Germany2
Remote - France2
Current role mix at Mindrift on JobCrawls
Role typeActive listings
Brand Designer1
Algorithm Problem Setter1
Presentation Designer1
Python Data Scraping Engineer1
Current role-level mix at Mindrift on JobCrawls
Role levelActive listings
Senior2
Intern1
Mid-Level1

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.