Lightning AI logo
Monthly
€8,319 - €10,128
Posted August 28, 2026 · 1 day agoLast seen August 27, 2026Est. expiry October 2, 2026

AI Platform Support Engineer

AI Platform Support Engineer (US)
How this salary compares
Salary Context: AI Platform Support Engineer

Hover or tap a row for full statistics (EUR / month on this chart).

Salary analysis

Compared with the selected benchmark ("Company in New York, United States"), this listing's salary midpoint is about 38% lower. The offer sits below the benchmark range (€11,936–€22,426). The listed pay band (€9,583–€11,667) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 1 comparable listings.

Monthly salary comparison for AI Platform Support Engineer
MarketLower bound (25th percentile)MedianUpper bound (75th percentile)
All roles in New York, United States€1,231/per month€12,913/per month€21,848/per month
Company in New York, United States€11,936/per month€17,181/per month€22,426/per month
About the role

Lightning AI is hiring an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production. You are a technical partner to ML teams—diagnosing failures, improving reliability, and guiding customers through complex distributed systems problems. The role is hybrid from our Seattle, San Francisco, or New York office hubs, with in-office requirements of at least 2 days per week and occasional team/offsite events. Working hours are 8:00 AM to 5:00 PM PST, Monday–Friday. We are not able to provide visa sponsorship for this role at this time. What You’ll Do - Work directly with ML engineers, diagnosing and resolving complex distributed systems and ML infrastructure issues, and acting as a technical advisor during high-impact incidents. - Debug ML infrastructure and distributed workloads, including distributed training, Kubernetes orchestration, GPU allocation, networking, and storage; troubleshoot PyTorch, CUDA, NCCL, and inference serving issues. - Analyze logs, metrics, traces, and system behavior to isolate root causes; debug containerized workloads on Kubernetes and bare metal GPUs; support scaling across multi-node GPU systems. - Improve reliability and platform operations by identifying recurring patterns, contributing to post-incident reviews, building internal tooling and runbooks, and collaborating with infrastructure, networking, and platform teams. - Enhance observability, troubleshooting workflows, and customer experience through better processes and guidance. What This Role Is Not - Not a traditional help desk or ticket routing role; not purely customer success or account management; not a backend engineer; not a passive escalation position. What You’ll Need - Strong software engineering and systems troubleshooting background; experience with Kubernetes and containerized environments; Linux networking, storage, and performance tuning. - Experience with cloud infrastructure, distributed systems, observability tools (Prometheus, Grafana, OpenTelemetry). - Hands-on experience operating ML workloads in production or research; distributed ML systems with PyTorch, CUDA, NCCL; GPU infrastructure and orchestration. - Excellent collaboration and communication skills; ability to work in fast-moving, ambiguous environments; enjoyment of solving complex technical problems with peers. - Ideal experience includes large-scale model training or distributed inference, Ray/Kubeflow/Slurm, InfiniBand/RDMA, bare metal infra, ML storage ecosystems, and Python automation. We offer competitive compensation with a discretionary bonus, equity, and benefits. The anticipated annual base salary range is $115,000 - $140,000 USD.

Job Details

Responsibilities

  • Work directly with ML engineers to diagnose and resolve issues in production
  • Troubleshoot distributed ML systems, Kubernetes, and GPU workloads
  • Investigate failures, analyze logs/metrics/traces, and isolate root causes
  • Collaborate with infrastructure, networking, and platform teams
  • Improve observability and runbooks; contribute to post-incident reviews

Requirements

  • Strong software engineering background
  • Experience with Kubernetes and containerized environments
  • Linux networking, storage, process management, and performance tuning
  • Experience with cloud infrastructure and distributed systems
  • Observability and debugging tools (Prometheus, Grafana, OpenTelemetry)
  • Hands-on experience with ML workloads in production or research
  • Distributed ML systems tooling (PyTorch, CUDA, NCCL)
  • Familiarity with GPU infrastructure and orchestration
  • Ability to diagnose performance, reliability, or scaling issues in ML infra
  • Excellent communication and collaboration skills

Skills & Technologies

KubernetesGPUPyTorchCUDANCCLOpenTelemetryPrometheusGrafanaLinuxKubernetes networkingDistributed systems
Seen 2 days agoContent Complete
Lightning AI logo
Lightning AI · 5 open roles
Top locations: London, United Kingdom · 1 · New York, United States · 1 · San Francisco, United States · 1+2 other locations
View company
Current open roles at Lightning AI on JobCrawls
LocationActive listings
London, United Kingdom1
New York, United States1
San Francisco, United States1
Seattle, United States1
Remote - Global1
Current role mix at Lightning AI on JobCrawls
Role typeActive listings
Research Engineer1
Current role-level mix at Lightning AI on JobCrawls
Role levelActive listings
Mid-Level1

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.