LILT (Production)-logo
Arvioitu kuukausipalkka
Arvioitu €5 505 - €8 500
Julkaistu 23. helmikuuta 2026 · 188 päivää sittenViimeksi nähty 26. elokuuta 2026Arvioitu päättymispäivä 30. maaliskuuta 2026

AI Benchmark-insinööri

Kuinka tämä palkka vertautuu muihin
Palkkayhteys: AI Benchmark-insinööri

Vie hiiri rivin päälle tai napauta riviä nähdäksesi täydelliset tilastot (EUR / kk tässä kaaviossa).

Tietoa tehtävästä

About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows. We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches. Note this is a remote, freelance opportunity What You’ll Deliver Task Engineering: Evaluating Coding Agents. Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. Prompting & Translation: finding failure points where AI does not work, in your native language Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary). Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus). Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity. Qualifications Experience: 5+ years of industry experience in software engineering. Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities. Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency. Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing. Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents. Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including: Encoding/decoding robustness and Unicode normalization. Locale-dependent conventions (collation, casing, non-Gregorian dates). Text I/O, toolchain interoperability, and safe string operations. (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts. Why Collaborate with Lilt? Your schedule, your rules. As an independent contractor, work when you want, as much or as little as you want. No fixed hours, no check-ins, no micromanaging. Get paid quickly and fairly. We respect your time and your expertise. Competitive rates, prompt payments, no chasing invoices. Work on projects that actually matter. Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate. Be part of something bigger. Join a global community of linguists, subject matter experts, and language professionals who are advancing human knowledge together. Grow without limits. As a Lilt contractor you get access to diverse, innovative projects that expand your portfolio and sharpen your skills across industries and domains. Have fun doing what you love. Bring your language skills to life on projects that are as interesting as they are impactful.We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

Työn tiedot

Seen 4 days agoPartial Schema
LILT (Production)-logo
LILT (Production) · 29 avointa tehtävää
Suosituimmat sijainnit: Etätyö - Globaali · 14 · New Delhi, Intia · 4 · Etätyö - Suomi · 2+8 muuta sijaintia
Näytä yritysprofiili
Tämänhetkiset avoimet roolit yrityksessä LILT (Production) JobCrawls-palvelussa
SijaintiAktiiviset ilmoitukset
Etätyö - Globaali14
New Delhi, Intia4
Etätyö - Suomi2
Etätyö - Globaalisti2
Algiers, Algeria1
Warsaw, Puola1
Etä1
Remote1
Etätyö - Saksa1
Etätyö - Norja1
Helsinki, Suomi1
Nykyinen roolien jakauma yrityksessä LILT (Production) JobCrawls-palvelussa
RoolityyppiAktiiviset ilmoitukset
AI Training Contributor3
Projektipäällikkö1
Subject Matter Expert1
Projektkoordinaattori1
Aiheen asiantuntija1
Asiantuntija1
Pelisprojektikoordinaattori1
Asiantuntija luonnontieteissä1
Matematiikan asiantuntija1
Asiantuntija – rahoitus1
Robotiikan asiantuntija1
Aineen asiantuntija1
Voice Talent1
Medical Translator1
Aihetta koskeva asiantuntija1
Aihealueen asiantuntija1
Kielenkääntäjä1
Kielitieteilijä1
Nykyinen roolitasojen jakauma yrityksessä LILT (Production) JobCrawls-palvelussa
RoolitasoAktiiviset ilmoitukset
Senior-taso9
Keskitaso5
Harjoittelija4
Asiantuntija2

Auta meitä parantamaan JobCrawlsia — kirjaudu sisään synkronoidaksesi tallennetut työpaikat laitteiden välillä, tai lähetä palautetta milloin tahansa.