LILT (Production) logo
Monthly
€6,945 - €10,417
Posted February 23, 2026 · 188 days agoLast seen August 26, 2026Est. expiry March 30, 2026

Benchmark Engineer

AI Benchmark Engineer
How this salary compares
Salary Context: Benchmark Engineer

Hover or tap a row for full statistics (EUR / month on this chart).

Salary analysis

Compared with the selected benchmark ("All roles in Berlin, Germany"), this listing's salary midpoint is about 3% lower. The offer still falls within the benchmark range (€1,667–€21,083). The listed pay band (€8,800–€13,200) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 10 comparable listings.

Monthly salary comparison for Benchmark Engineer
MarketLower bound (25th percentile)MedianUpper bound (75th percentile)
All roles in Berlin, Germany€1,667/per month€8,010/per month€21,083/per month
About the role

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity What You’ll Deliver Task Engineering: Evaluating Coding Agents. Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. Prompting & Translation: finding failure points where AI does not work, in your native language Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary). Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus). Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

Job Details

Responsibilities

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build realistic task environments using datasets and files in your native language.
  • Prompting & Translation: finding failure points where AI does not work, in your native language
  • Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).
  • Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).
  • Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

Requirements

  • Experience: 1+ years of industry experience in software or prompt engineering.
  • Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.
  • Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency.
  • Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing.
  • Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.
  • Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including: Encoding/decoding robustness and Unicode normalization; Locale-dependent conventions (collation, casing, non-Gregorian dates); Text I/O, toolchain interoperability, and safe string operations; (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.

Skills & Technologies

PythonShell scriptingData processing

Recruitment Process

  1. 1
    Submit application in English
  2. 2
    Complete GenAI assessment
  3. 3
    Onboard in system and become eligible for Applied AI projects
Seen 3 days agoPartial Schema
LILT (Production) logo
LILT (Production) · 29 open roles
Top locations: Remote - Global · 18 · New Delhi, India · 4 · Remote - Finland · 2+5 other locations
View company
Current open roles at LILT (Production) on JobCrawls
LocationActive listings
Remote - Global18
New Delhi, India4
Remote - Finland2
Remote - Norway1
Remote - Germany1
Warsaw, Poland1
Helsinki, Finland1
Algiers, Algeria1
Current role mix at LILT (Production) on JobCrawls
Role typeActive listings
Subject Matter Expert8
AI Training Contributor3
Linguist2
Medical Translator1
Robotics Subject Matter Expert1
Voice Talent1
Project Coordinator1
Subject Matter Expert in Mathematics1
Gaming Project Coordinator1
Project Manager1
Current role-level mix at LILT (Production) on JobCrawls
Role levelActive listings
Senior9
Mid-Level5
Intern4
Expert2

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.