
Hover or tap a row for full statistics (EUR / month on this chart).
Salary analysis
Compared with the selected benchmark ("All roles in Helsinki, Finland"), this listing's salary midpoint is about 93% lower. The offer sits below the benchmark range (€2,070–€5,000). The listed pay band (€219–€265) is tighter than the benchmark, which suggests lower salary variability. This benchmark is based on 2534 comparable listings.
| Market | Lower bound (25th percentile) | Median | Upper bound (75th percentile) |
|---|---|---|---|
| All roles in Helsinki, Finland | €2,070/per month | €2,907/per month | €5,000/per month |
| Pay in our data — not quoted in ad (Senior) | €219/per month | €242/per month | €265/per month |
AMD is looking for a performance-obsessed senior engineer to push AI inference performance to the limit on AMD GPUs, with vLLM as the primary serving framework. You will work end-to-end across the stack: profiling, diagnosing, and optimizing leading models running on vLLM across customer-relevant serving configurations (e.g. agentic coding, long-context, high-throughput serving). You will own hard performance problems on our most strategic customer engagements and leave behind measurable uplifts and reusable methodology. KEY RESPONSIBILITIES: - Drive performance optimization end-to-end on vLLM across leading models and customer-relevant serving configurations, closing competitive gaps through kernel and systems-level optimizations - Profile, diagnose, and resolve cross-stack performance bottlenecks in vLLM deployments, from GPU kernels and operator dispatch to the vLLM scheduler, PagedAttention/KV cache management, and multi-node communication - Diagnose kernel-level performance issues using profiling tools: identify occupancy limitations, L2 cache thrashing, register pressure, memory coalescing issues, etc, and translate findings into actionable optimizations - Contribute to customer-facing technical engagements: present findings, recommend optimizations, and deliver measurable performance uplifts on vLLM - Integrate and optimize custom kernels (Triton, Gluon, CK, PyDSL, ASM, AITER) within vLLM, understanding dispatch paths, shape extraction, and backend selection - Optimize multi-node distributed inference on vLLM: communication-compute overlap, parallelism strategies (TP/PP/EP/DP), and scale-out performance - Contribute to shared performance optimization methodology that raises the bar across the team - Leverage AI agents to accelerate daily work and help define best practices for AI-assisted performance engineering - Upstream optimizations into vLLM and adjacent open-source frameworks such as SGLang and PyTorch PREFERRED EXPERIENCE: - 5+ years of software development experience in GPU computing, AI systems, or high-performance computing - Hands-on experience with vLLM internals (V1 engine, scheduler, PagedAttention/KV cache manager, chunked prefill, ROCm backend integration); familiarity with SGLang, TensorRT-LLM, or similar is a plus - Strong background in end-to-end workload profiling and bottleneck diagnosis - Understanding of GPU kernel performance characteristics: occupancy, register and LDS pressure, memory coalescing, cache utilization, wavefront scheduling, and instruction-level bottlenecks - Ability to read and reason about kernel-level profiling data and translate it into concrete optimization actions - Understanding of model architectures (transformers, MoE, diffusion), inference paradigms (speculative decoding, prefill-decode disaggregation, continuous batching), and how they map to hardware and to vLLM's execution model - Experience with custom kernel development or integration (HIP, CUDA, Triton, CK, or similar) - Understanding of multi-GPU and multi-node distributed systems: scale-up and scale-out topologies, RCCL/NCCL, RDMA, and communication-compute overlap - Strong proficiency in Python and C++ - Ability to engage with customers, present findings, and support technical decisions - Fluent in AI-assisted development: daily user of AI agents and tools - Strong Linux systems knowledge - Excellent written and verbal English communication skills ACADEMIC CREDENTIALS: Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent. Advanced degree preferred but exceptional industry experience valued equally. LOCATION: Helsinki, Finland or Stockholm, Sweden.
Job Details
Responsibilities
- Drive end-to-end performance optimization on vLLM across leading models and serving configurations
- Profile, diagnose, and resolve cross-stack performance bottlenecks from GPU kernels to the vLLM scheduler
- Identify kernel-level issues like occupancy limitations and L2 cache thrashing using profiling tools
- Contribute to customer-facing technical engagements and deliver measurable performance uplifts
- Integrate and optimize custom kernels (Triton, Gluon, CK, PyDSL, ASM, AITER) within vLLM
- Optimize multi-node distributed inference, including communication-compute overlap and parallelism strategies
- Develop shared performance optimization methodologies for the team
- Leverage AI agents to accelerate daily work and define AI-assisted engineering best practices
- Upstream optimizations into vLLM, SGLang, and PyTorch
Requirements
- 5+ years of software development experience in GPU computing, AI systems, or high-performance computing
- Hands-on experience with vLLM internals (V1 engine, scheduler, PagedAttention/KV cache manager, chunked prefill, ROCm backend integration)
- Strong background in end-to-end workload profiling and bottleneck diagnosis
- Understanding of GPU kernel performance characteristics (occupancy, register/LDS pressure, memory coalescing, cache utilization)
- Ability to reason about kernel-level profiling data and translate it into optimizations
- Understanding of model architectures (transformers, MoE, diffusion) and inference paradigms
- Experience with custom kernel development or integration (HIP, CUDA, Triton, CK, or similar)
- Understanding of multi-GPU and multi-node distributed systems (RCCL/NCCL, RDMA)
- Strong proficiency in Python and C++
- Strong Linux systems knowledge
- Excellent written and verbal English communication skills
Skills & Technologies
Education Level
Masters
Related Opportunities
Discover more opportunities that match your interests and skills