
AI Performance Engineer
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("All roles in Helsinki, Finland"), this listing's salary midpoint is about 93% lower. The offer sits below the benchmark range (€2,070–€5,000). The listed pay band (€219–€265) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 2534 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| All roles in Helsinki, Finland | €2,070/per month | €2,907/per month | €5,000/per month |
| Pay in our data — not quoted in ad (Senior) | €219/per month | €242/per month | €265/per month |
AMD is looking for a performance-obsessed engineer to drive AI inference performance to the absolute limit on AMD GPUs, with SGLang as the primary serving framework. You will lead a small, highly technical team and work end-to-end across the stack: profiling, diagnosing, and optimizing leading models running on SGLang across customer-relevant serving configurations (e.g. agentic coding, long-context, high-throughput serving). KEY RESPONSIBILITIES: - Drive performance optimization end-to-end on SGLang across leading models and customer-relevant serving configurations, closing competitive gaps through kernel and systems-level optimizations - Profile, diagnose, and resolve the hardest cross-stack performance bottlenecks in SGLang deployments, from GPU kernels and operator dispatch to the SGLang scheduler, RadixAttention/prefix caching, and multi-node communication - Diagnose kernel-level performance issues using profiling tools: identify occupancy limitations, L2 cache thrashing, register pressure, memory coalescing issues, etc, and translate findings into actionable optimizations - Lead customer-facing technical engagements: present findings, recommend optimizations, and deliver measurable performance uplifts on SGLang - Integrate and optimize custom kernels (Triton, Gluon, CK, PyDSL, ASM, AITER) within SGLang, understanding dispatch paths, shape extraction, and backend selection - Optimize multi-node distributed inference on SGLang: communication-compute overlap, parallelism strategies (TP/EP/DP), and scale-out performance - Develop and refine shared performance optimization methodology that raises the bar across the broader team - Leverage AI agents to accelerate daily work and define best practices for AI-assisted performance engineering - Upstream optimizations into SGLang and adjacent open-source frameworks such as vLLM and PyTorch PREFERRED EXPERIENCE: - 7+ years of software development experience in GPU computing, AI systems, or high-performance computing - Deep hands-on experience with SGLang internals; familiarity with vLLM, TensorRT-LLM, or similar is a plus - Strong background in end-to-end workload profiling and bottleneck diagnosis - Understanding of GPU kernel performance characteristics: occupancy, register and LDS pressure, memory coalescing, cache utilization, wavefront scheduling, and instruction-level bottlenecks - Understanding of model architectures (transformers, MoE, diffusion), inference paradigms (speculative decoding, prefill-decode disaggregation, continuous batching) - Experience with custom kernel development or integration (HIP, CUDA, Triton, CK, or similar) - Understanding of multi-GPU and multi-node distributed systems: RCCL/NCCL, RDMA - Strong proficiency in Python and C++ - Customer-facing technical leadership experience - Fluent in AI-assisted development - Strong Linux systems knowledge - Excellent written and verbal English communication skills ACADEMIC CREDENTIALS: Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent.
Job Details
Responsibilities
- Drive end-to-end performance optimization on SGLang across leading models and customer configurations
- Profile, diagnose, and resolve cross-stack performance bottlenecks from GPU kernels to the SGLang scheduler
- Identify and resolve kernel-level issues such as occupancy limitations and L2 cache thrashing
- Lead customer-facing technical engagements and deliver measurable performance uplifts
- Integrate and optimize custom kernels (Triton, Gluon, CK, PyDSL, ASM, AITER) within SGLang
- Optimize multi-node distributed inference, focusing on communication-compute overlap and parallelism strategies
- Develop shared performance optimization methodology for the team
- Leverage AI agents to accelerate workflows and define AI-assisted engineering best practices
- Upstream optimizations into SGLang, vLLM, and PyTorch
Requirements
- 7+ years of software development experience in GPU computing, AI systems, or high-performance computing
- Deep hands-on experience with SGLang internals
- Strong background in end-to-end workload profiling and bottleneck diagnosis
- Understanding of GPU kernel performance characteristics (occupancy, register/LDS pressure, memory coalescing, cache utilization)
- Understanding of model architectures (transformers, MoE, diffusion) and inference paradigms (speculative decoding, continuous batching)
- Experience with custom kernel development or integration (HIP, CUDA, Triton, CK, or similar)
- Understanding of multi-GPU and multi-node distributed systems (RCCL/NCCL, RDMA)
- Strong proficiency in Python and C++
- Customer-facing technical leadership experience
- Strong Linux systems knowledge
- Excellent written and verbal English communication skills
- Master's or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Skills & Technologies
Education Level
Masters
Related Opportunities
Discover more opportunities that match your interests and skills