Nebius B.V. logo
Monthly
€97,835 - €122,228
Posted August 29, 2026 · 0 days agoLast seen August 28, 2026Est. expiry October 3, 2026

IT Infrastructure Engineer

IT Infrastructure Engineer – RMA & Hardware Diagnostics
About the role

The role The IT Infrastructure Engineer – RMA & Hardware Diagnostics is responsible for advanced hardware troubleshooting and RMA lifecycle management within Nebius production data center environments. This role serves as the escalation point for complex server and firmware-related issues that impact system reliability and fleet availability. You will be performing deep diagnostics across enterprise server platforms, conducting structured root cause analysis, validating failed components, and managing end-to-end warranty replacement processes with OEM vendors. This is a hands-on technical role with direct impact on hardware reliability, SLA performance, and operational scalability. You will work on-site in Nebius data centers, collaborating closely with L1/L2 technicians, infrastructure engineers, and vendors to reduce repeat failures and improve hardware quality across the fleet. Minnesota is listed as a location. Responsibilities include: performing advanced firmware and hardware diagnostics on enterprise server platforms (CPU, memory, PCIe devices, GPUs, storage subsystems, power components); troubleshooting complex hardware failures using system logs, IPMI interfaces, BIOS diagnostics, and vendor-specific tooling; acting as the primary escalation point for L1 and L2 technicians on high-impact hardware incidents; conducting structured root cause analysis and documenting findings; owning the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution; interfacing directly with OEM vendors to escalate recurring hardware defects and drive corrective action; analyzing hardware failure trends and reporting metrics; developing and standardizing diagnostic playbooks, troubleshooting workflows, and hardware validation procedures; validating replacement components prior to redeployment into production environments; collaborating cross-functionally with data center operations, procurement, and engineering teams to improve hardware lifecycle processes; contributing to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing. We expect you to have 5+ years of hands-on experience with enterprise server hardware in a production data center environment; deep understanding of x86 server architecture; strong experience with firmware and BIOS/BMC diagnostics and upgrades; advanced Linux command-line troubleshooting; experience with remote management interfaces (IPMI, iDRAC, iLO); proven RMA management with OEM vendors; ability to perform structured root cause analysis and document findings clearly; familiarity with hardware monitoring systems and failure trend analysis; strong ownership mindset and ability to operate independently in mission-critical environments; high proficiency in spoken and written English. Bonus points: experience performing board-level diagnostics, data center networking exposure, GPU-dense HPC environments, and valid driver’s license. Compensation: We offer competitive salaries ranging from $112,700.00 - $140,800.00 OTE (On Target Earnings) based on experience and skills. Benefits & Perks include competitive compensation, career growth, flexibility and ownership, collaborative and innovative culture, opportunity to work on impactful AI projects, international environment and talented teams. What Nebius is like to work at: fast moving, bold thinking, constant growth, meaningful impact, trust and real ownership, opportunity to shape the future of AI. Equity and equal opportunity statement included.

Job Details

Responsibilities

  • Perform advanced firmware and hardware diagnostics on enterprise server platforms (CPU, memory, PCIe devices, GPUs, storage subsystems, power components)
  • Troubleshoot complex hardware failures using system logs, IPMI/BMC interfaces, BIOS diagnostics, and vendor-specific tooling
  • Act as the primary escalation point for L1/L2 technicians on high-impact hardware incidents
  • Conduct structured root cause analysis and document findings
  • Own the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution
  • Interface with OEM vendors to escalate recurring hardware defects and drive corrective action
  • Analyze hardware failure trends and report metrics such as repeat RMA rates and component reliability
  • Develop and standardize diagnostic playbooks, troubleshooting workflows, and hardware validation procedures
  • Validate replacement components prior to redeployment into production environments
  • Collaborate with data center operations, procurement, and engineering teams to improve hardware lifecycle processes
  • Contribute to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing

Requirements

  • 5+ years of hands-on experience with enterprise server hardware in a production data center environment
  • Deep understanding of x86 server architecture
  • Strong experience with firmware and BIOS/BMC diagnostics and upgrades
  • Advanced Linux command-line troubleshooting skills
  • Experience with remote management interfaces such as IPMI, iDRAC, iLO
  • Proven experience managing hardware RMA processes with OEM vendors
  • Ability to conduct structured root cause analysis and document findings clearly
  • Familiarity with hardware monitoring systems and failure trend analysis
  • Strong ownership mindset and ability to operate independently in mission-critical environments
  • High proficiency in spoken and written English

Skills & Technologies

IPMIiDRACiLOBMCBIOS diagnosticsLinuxServer hardwareRMA processesRoot cause analysisData center operations
Seen 20 hours agoPartial Schema
Nebius B.V. logo
Nebius B.V. · 790 open roles
Top locations: Remote - Global · 423 · Remote - Europe · 119 · Amsterdam, Netherlands · 56+57 other locations
View company
Current open roles at Nebius B.V. on JobCrawls
LocationActive listings
Remote - Global423
Remote - Europe119
Amsterdam, Netherlands56
London, United Kingdom22
Remote - United States20
Remote - Finland18
Berlin, Germany15
Mäntsälä, Finland12
Helsinki, Finland10
Lappeenranta, Finland9
Prague, Czech Republic6
Amsterdam5
Israel5
United Kingdom4
Canada4
Remote3
Tel Aviv, Israel3
Singapore3
Remote - United Kingdom2
Abu Dhabi2
Dubai2
Paris, France2
Remote - France2
New York City, United States2
Austin, United States2
France, Paris2
Remote - Germany2
London2
Remote - Netherlands2
Philadelphia, United States1
Béthune, France1
Abu Dhabi, Dubai1
Singapore, Singapore1
California, United States1
Abu Dhabi, United Arab Emirates1
Remote - EU1
Munich, Germany1
Minnesota, United States1
Alabama, US1
Prague, Czechia1
East London, United Kingdom1
Canada, Remote - United States1
Berlin1
Dallas, United States1
London, UK1
Oklahoma, United States1
New Jersey, US1
Austin, Texas1
Kansas City, United States1
San Francisco Bay Area, United States1
Czechia1
Finland1
Remote - Sweden1
UK1
Prague1
New Jersey, United States1
Remote - Czech Republic1
Netherlands1
Paris1
Béthune, Pas-de-Calais, France1
Current role mix at Nebius B.V. on JobCrawls
Role typeActive listings
Backend Engineer331
Software Engineer82
Account Executive58
Technical Project Manager6
Site Reliability Engineer4
Technical Product Manager4
System Engineer4
Sales Representative4
Product Manager3
Data Center Operations Technician3
Data Center Technician3
ML Engineer3
Technical Program Manager3
Delivery Manager2
Applied AI Researcher2
Hypervisor Engineer2
Product Designer2
IT Technician2
Backend engineers, Frontend engineers, Site reliability engineers2
Open Positions at Nebius2
VP of Strategic Sales1
Generalist1
Offensive Security Lead1
Head of Channel Marketing1
Principal1
Applied AI Solutions Engineer1
Field Technical Lead1
Compensation Analyst1
Solutions Architecture Leader1
Application Security Engineer1
Human Resources Specialist1
Solutions Architect1
Mechanical Data Center Technician1
Backend Developer1
Senior Research Scientist1
Cloud Solution Architect1
Data Center IT Technician1
Manager, ML Solutions Architecture1
ML Solutions Architect1
Network Planning Project Manager1
Data Center Logistics Specialist1
Technical Due Diligence Manager1
Data Center Operations Manager1
Senior Site Reliability Engineer1
IT Support Manager1
Technical Support Engineer1
Data Engineer1
Mechanical Engineer1
Data Center IT Manager1
Security Product Manager1
Data Center Electrical Lead1
GTM Recruiting Manager1
Security Solutions Engineer1
Physical Security Systems Technician1
Senior HPC Engineer1
Educational Content Author1
Customer Engineer1
Pricing Director1
MEP Engineer1
Instructional Designer1
Partner Solutions Architect1
Senior Software Developer1
Forward Deployment Engineer1
Financial Controller1
Internal Control Business Partner1
Data Center Facilities Manager1
Group Product Manager1
Structured Cabling Design Engineer1
Site Selection & Colocation Manager1
Product Growth Analytics Lead1
IT Risk and Control Manager1
Vulnerability Operations Center Lead1
Electrical Engineer1
Senior System Engineer1
Machine Learning Engineer1
ML Infrastructure Engineer1
Data Scientist1
Mechanical Design Engineer1
Applied ML Engineer1
Deal Initiation and Activation Manager1
Operations Specialist1
AI/ML Specialist Solutions Architect1
HPC Engineer1
Accountant1
VP of Developer Relations & Community1
Solutions Partner1
Senior Support Engineer1
Network Engineer1
Project Development Manager1
Backend Engineers1
Current role-level mix at Nebius B.V. on JobCrawls
Role levelActive listings
Mid-Level410
Senior63
Manager12
Executive2
Director1

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.