
Vie hiiri rivin päälle tai napauta riviä nähdäksesi täydelliset tilastot (EUR / kk tässä kaaviossa).
Lightning AI is hiring an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production. You are a technical partner to ML teams—diagnosing failures, improving reliability, and guiding customers through complex distributed systems problems. The role is hybrid from our Seattle, San Francisco, or New York office hubs, with in-office requirements of at least 2 days per week and occasional team/offsite events. Working hours are 8:00 AM to 5:00 PM PST, Monday–Friday. We are not able to provide visa sponsorship for this role at this time. What You’ll Do - Work directly with ML engineers, diagnosing and resolving complex distributed systems and ML infrastructure issues, and acting as a technical advisor during high-impact incidents. - Debug ML infrastructure and distributed workloads, including distributed training, Kubernetes orchestration, GPU allocation, networking, and storage; troubleshoot PyTorch, CUDA, NCCL, and inference serving issues. - Analyze logs, metrics, traces, and system behavior to isolate root causes; debug containerized workloads on Kubernetes and bare metal GPUs; support scaling across multi-node GPU systems. - Improve reliability and platform operations by identifying recurring patterns, contributing to post-incident reviews, building internal tooling and runbooks, and collaborating with infrastructure, networking, and platform teams. - Enhance observability, troubleshooting workflows, and customer experience through better processes and guidance. What This Role Is Not - Not a traditional help desk or ticket routing role; not purely customer success or account management; not a backend engineer; not a passive escalation position. What You’ll Need - Strong software engineering and systems troubleshooting background; experience with Kubernetes and containerized environments; Linux networking, storage, and performance tuning. - Experience with cloud infrastructure, distributed systems, observability tools (Prometheus, Grafana, OpenTelemetry). - Hands-on experience operating ML workloads in production or research; distributed ML systems with PyTorch, CUDA, NCCL; GPU infrastructure and orchestration. - Excellent collaboration and communication skills; ability to work in fast-moving, ambiguous environments; enjoyment of solving complex technical problems with peers. - Ideal experience includes large-scale model training or distributed inference, Ray/Kubeflow/Slurm, InfiniBand/RDMA, bare metal infra, ML storage ecosystems, and Python automation. We offer competitive compensation with a discretionary bonus, equity, and benefits. The anticipated annual base salary range is $115,000 - $140,000 USD.
Työn tiedot
Vastuualueet
- Työ ML-insinöörien kanssa tuotantoon liittyvien ongelmien diagnosoinnissa ja ratkaisemisessa
- Vikahäiriöt hajautetuissa ML-järjestelmissä ja GPU-kuormissa
- Tutki virheitä, analysoi lokit, mittaukset ja jäljet sekä erota juurisyy
- Yhteistyö infrastruktuuri-, verkkoyhteys- ja alusta-tiimien kanssa
- Paranna monitorointia ja runbookeja; osallistu jälkikäteisarviointeihin
Vaatimukset
- Vahva ohjelmisto- ja järjestelmäkatselmointi tausta
- Kokemus Kubernetes- ja kontaineroitujen ympäristöjen kanssa
- Linuxin verkkostoraget, prosessinhallinta ja suorituskyky
- Kokemus pilvi-infrastruktuurista ja hajautetuista järjestelmistä
- Observability- ja debugging-työkalut (Prometheus, Grafana, OpenTelemetry)
- Pilottien ML-kuormien operointi tuotannossa tai tutkimuksessa
- Hajautetut ML-järjestelmät (PyTorch, CUDA, NCCL)
- Kokemus GPU-infrastruktuurista ja orkestroinnista
- Kyky diagnosoida suorituskykyyn, luotettavuuteen tai skaalautumiseen liittyviä ongelmia
- Erinomaiset viestintä- ja yhteistyötaidot
Taidot ja teknologiat
Edut ja etuudet

| Sijainti | Aktiiviset ilmoitukset |
|---|---|
| Lontoo, Yhdistynyt kuningaskunta | 1 |
| Etätyö - Globali | 1 |
| Seattle, Yhdysvallat | 1 |
| New York, Yhdysvallat | 1 |
| San Francisco, Yhdysvallat | 1 |
| Roolityyppi | Aktiiviset ilmoitukset |
|---|---|
| Tutkimusinsinööri | 1 |
| Roolitaso | Aktiiviset ilmoitukset |
|---|---|
| Keskitaso | 1 |
Aiheeseen liittyvät mahdollisuudet
Löydä lisää kiinnostuksen kohteisiisi ja taitoihisi sopivia mahdollisuuksia