
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("All roles in Berlin, Germany"), this listing's salary midpoint is about 3% lower. The offer still falls within the benchmark range (€1,667–€21,083). The listed pay band (€8,800–€13,200) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 10 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| All roles in Berlin, Germany | €1,667/per month | €8,010/per month | €21,083/per month |
About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity What You’ll Deliver Task Engineering: Evaluating Coding Agents. Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. Prompting & Translation: finding failure points where AI does not work, in your native language Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary). Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus). Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.
Job Details
Responsibilities
- Task Engineering: Evaluating Coding Agents.
- Asset Creation: Build realistic task environments using datasets and files in your native language.
- Prompting & Translation: finding failure points where AI does not work, in your native language
- Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).
- Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).
- Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.
Requirements
- Experience: 1+ years of industry experience in software or prompt engineering.
- Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
- Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency.
- Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing.
- Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.
- Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including: Encoding/decoding robustness and Unicode normalization; Locale-dependent conventions (collation, casing, non-Gregorian dates); Text I/O, toolchain interoperability, and safe string operations; (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.
Skills & Technologies
Recruitment Process
- 1Submit application in English
- 2Complete GenAI assessment
- 3Onboard in system and become eligible for Applied AI projects

| Location | Active listings |
|---|---|
| Remote - Global | 18 |
| New Delhi, India | 4 |
| Remote - Finland | 2 |
| Remote - Norway | 1 |
| Remote - Germany | 1 |
| Warsaw, Poland | 1 |
| Helsinki, Finland | 1 |
| Algiers, Algeria | 1 |
| Role type | Active listings |
|---|---|
| Subject Matter Expert | 8 |
| AI Training Contributor | 3 |
| Linguist | 2 |
| Medical Translator | 1 |
| Robotics Subject Matter Expert | 1 |
| Voice Talent | 1 |
| Project Coordinator | 1 |
| Subject Matter Expert in Mathematics | 1 |
| Gaming Project Coordinator | 1 |
| Project Manager | 1 |
| Role level | Active listings |
|---|---|
| Senior | 9 |
| Mid-Level | 5 |
| Intern | 4 |
| Expert | 2 |
Related Opportunities
Discover more opportunities that match your interests and skills