Lightning AI-logo
Kuukausipalkka
€8 319 - €10 128
Julkaistu 28. elokuuta 2026 · 1 päivä sittenViimeksi nähty 27. elokuuta 2026Arvioitu päättymispäivä 2. lokakuuta 2026

AI Platform Support Engineer

AI Platform Support Engineer (US)
Kuinka tämä palkka vertautuu muihin
Palkkayhteys: AI Platform Support Engineer

Vie hiiri rivin päälle tai napauta riviä nähdäksesi täydelliset tilastot (EUR / kk tässä kaaviossa).

Tietoa tehtävästä

Lightning AI is hiring an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production. You are a technical partner to ML teams—diagnosing failures, improving reliability, and guiding customers through complex distributed systems problems. The role is hybrid from our Seattle, San Francisco, or New York office hubs, with in-office requirements of at least 2 days per week and occasional team/offsite events. Working hours are 8:00 AM to 5:00 PM PST, Monday–Friday. We are not able to provide visa sponsorship for this role at this time. What You’ll Do - Work directly with ML engineers, diagnosing and resolving complex distributed systems and ML infrastructure issues, and acting as a technical advisor during high-impact incidents. - Debug ML infrastructure and distributed workloads, including distributed training, Kubernetes orchestration, GPU allocation, networking, and storage; troubleshoot PyTorch, CUDA, NCCL, and inference serving issues. - Analyze logs, metrics, traces, and system behavior to isolate root causes; debug containerized workloads on Kubernetes and bare metal GPUs; support scaling across multi-node GPU systems. - Improve reliability and platform operations by identifying recurring patterns, contributing to post-incident reviews, building internal tooling and runbooks, and collaborating with infrastructure, networking, and platform teams. - Enhance observability, troubleshooting workflows, and customer experience through better processes and guidance. What This Role Is Not - Not a traditional help desk or ticket routing role; not purely customer success or account management; not a backend engineer; not a passive escalation position. What You’ll Need - Strong software engineering and systems troubleshooting background; experience with Kubernetes and containerized environments; Linux networking, storage, and performance tuning. - Experience with cloud infrastructure, distributed systems, observability tools (Prometheus, Grafana, OpenTelemetry). - Hands-on experience operating ML workloads in production or research; distributed ML systems with PyTorch, CUDA, NCCL; GPU infrastructure and orchestration. - Excellent collaboration and communication skills; ability to work in fast-moving, ambiguous environments; enjoyment of solving complex technical problems with peers. - Ideal experience includes large-scale model training or distributed inference, Ray/Kubeflow/Slurm, InfiniBand/RDMA, bare metal infra, ML storage ecosystems, and Python automation. We offer competitive compensation with a discretionary bonus, equity, and benefits. The anticipated annual base salary range is $115,000 - $140,000 USD.

Työn tiedot

Vastuualueet

  • Työ ML-insinöörien kanssa tuotantoon liittyvien ongelmien diagnosoinnissa ja ratkaisemisessa
  • Vikahäiriöt hajautetuissa ML-järjestelmissä ja GPU-kuormissa
  • Tutki virheitä, analysoi lokit, mittaukset ja jäljet sekä erota juurisyy
  • Yhteistyö infrastruktuuri-, verkkoyhteys- ja alusta-tiimien kanssa
  • Paranna monitorointia ja runbookeja; osallistu jälkikäteisarviointeihin

Vaatimukset

  • Vahva ohjelmisto- ja järjestelmäkatselmointi tausta
  • Kokemus Kubernetes- ja kontaineroitujen ympäristöjen kanssa
  • Linuxin verkkostoraget, prosessinhallinta ja suorituskyky
  • Kokemus pilvi-infrastruktuurista ja hajautetuista järjestelmistä
  • Observability- ja debugging-työkalut (Prometheus, Grafana, OpenTelemetry)
  • Pilottien ML-kuormien operointi tuotannossa tai tutkimuksessa
  • Hajautetut ML-järjestelmät (PyTorch, CUDA, NCCL)
  • Kokemus GPU-infrastruktuurista ja orkestroinnista
  • Kyky diagnosoida suorituskykyyn, luotettavuuteen tai skaalautumiseen liittyviä ongelmia
  • Erinomaiset viestintä- ja yhteistyötaidot

Taidot ja teknologiat

KubernetesGPUPyTorchCUDANCCLOpenTelemetryPrometheusGrafanaLinuxKubernetes-verkkoHajautetut järjestelmät
Seen 2 days agoContent Complete
Lightning AI-logo
Lightning AI · 5 avointa tehtävää
Suosituimmat sijainnit: Lontoo, Yhdistynyt kuningaskunta · 1 · Etätyö - Globali · 1 · Seattle, Yhdysvallat · 1+2 muuta sijaintia
Näytä yritysprofiili
Tämänhetkiset avoimet roolit yrityksessä Lightning AI JobCrawls-palvelussa
SijaintiAktiiviset ilmoitukset
Lontoo, Yhdistynyt kuningaskunta1
Etätyö - Globali1
Seattle, Yhdysvallat1
New York, Yhdysvallat1
San Francisco, Yhdysvallat1
Nykyinen roolien jakauma yrityksessä Lightning AI JobCrawls-palvelussa
RoolityyppiAktiiviset ilmoitukset
Tutkimusinsinööri1
Nykyinen roolitasojen jakauma yrityksessä Lightning AI JobCrawls-palvelussa
RoolitasoAktiiviset ilmoitukset
Keskitaso1

Auta meitä parantamaan JobCrawlsia — kirjaudu sisään synkronoidaksesi tallennetut työpaikat laitteiden välillä, tai lähetä palautetta milloin tahansa.