Nebius B.V. logo
Monthly
€97,835 - €122,228
Posted August 29, 2026 · 0 days agoLast seen August 28, 2026Est. expiry October 3, 2026

IT Infrastructure Engineer

IT Infrastructure Engineer – RMA & Hardware Diagnostics
About the role

The role The IT Infrastructure Engineer – RMA & Hardware Diagnostics is responsible for advanced hardware troubleshooting and RMA lifecycle management within Nebius production data center environments. This role serves as the escalation point for complex server and firmware-related issues that impact system reliability and fleet availability. You will be performing deep diagnostics across enterprise server platforms, conducting structured root cause analysis, validating failed components, and managing end-to-end warranty replacement processes with OEM vendors. This is a hands-on technical role with direct impact on hardware reliability, SLA performance, and operational scalability. You will work on-site in Nebius data centers, collaborating closely with L1/L2 technicians, infrastructure engineers, and vendors to reduce repeat failures and improve hardware quality across the fleet. Minnesota is listed as a location. Responsibilities include: performing advanced firmware and hardware diagnostics on enterprise server platforms (CPU, memory, PCIe devices, GPUs, storage subsystems, power components); troubleshooting complex hardware failures using system logs, IPMI interfaces, BIOS diagnostics, and vendor-specific tooling; acting as the primary escalation point for L1 and L2 technicians on high-impact hardware incidents; conducting structured root cause analysis and documenting findings; owning the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution; interfacing directly with OEM vendors to escalate recurring hardware defects and drive corrective action; analyzing hardware failure trends and reporting metrics; developing and standardizing diagnostic playbooks, troubleshooting workflows, and hardware validation procedures; validating replacement components prior to redeployment into production environments; collaborating cross-functionally with data center operations, procurement, and engineering teams to improve hardware lifecycle processes; contributing to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing. We expect you to have 5+ years of hands-on experience with enterprise server hardware in a production data center environment; deep understanding of x86 server architecture; strong experience with firmware and BIOS/BMC diagnostics and upgrades; advanced Linux command-line troubleshooting; experience with remote management interfaces (IPMI, iDRAC, iLO); proven RMA management with OEM vendors; ability to perform structured root cause analysis and document findings clearly; familiarity with hardware monitoring systems and failure trend analysis; strong ownership mindset and ability to operate independently in mission-critical environments; high proficiency in spoken and written English. Bonus points: experience performing board-level diagnostics, data center networking exposure, GPU-dense HPC environments, and valid driver’s license. Compensation: We offer competitive salaries ranging from $112,700.00 - $140,800.00 OTE (On Target Earnings) based on experience and skills. Benefits & Perks include competitive compensation, career growth, flexibility and ownership, collaborative and innovative culture, opportunity to work on impactful AI projects, international environment and talented teams. What Nebius is like to work at: fast moving, bold thinking, constant growth, meaningful impact, trust and real ownership, opportunity to shape the future of AI. Equity and equal opportunity statement included.

Job Details

Responsibilities

  • Perform advanced firmware and hardware diagnostics on enterprise server platforms (CPU, memory, PCIe devices, GPUs, storage subsystems, power components)
  • Troubleshoot complex hardware failures using system logs, IPMI/BMC interfaces, BIOS diagnostics, and vendor-specific tooling
  • Act as the primary escalation point for L1/L2 technicians on high-impact hardware incidents
  • Conduct structured root cause analysis and document findings
  • Own the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution
  • Interface with OEM vendors to escalate recurring hardware defects and drive corrective action
  • Analyze hardware failure trends and report metrics such as repeat RMA rates and component reliability
  • Develop and standardize diagnostic playbooks, troubleshooting workflows, and hardware validation procedures
  • Validate replacement components prior to redeployment into production environments
  • Collaborate with data center operations, procurement, and engineering teams to improve hardware lifecycle processes
  • Contribute to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing

Requirements

  • 5+ years of hands-on experience with enterprise server hardware in a production data center environment
  • Deep understanding of x86 server architecture
  • Strong experience with firmware and BIOS/BMC diagnostics and upgrades
  • Advanced Linux command-line troubleshooting skills
  • Experience with remote management interfaces such as IPMI, iDRAC, iLO
  • Proven experience managing hardware RMA processes with OEM vendors
  • Ability to conduct structured root cause analysis and document findings clearly
  • Familiarity with hardware monitoring systems and failure trend analysis
  • Strong ownership mindset and ability to operate independently in mission-critical environments
  • High proficiency in spoken and written English

Skills & Technologies

IPMIiDRACiLOBMCBIOS diagnosticsLinuxServer hardwareRMA processesRoot cause analysisData center operations
Seen 1 day agoPartial Schema
Nebius B.V. logo
Nebius B.V. · 772 open roles
Top locations: Remote - Global · 398 · Remote - Europe · 126 · Amsterdam, Netherlands · 57+54 other locations
View company
Current open roles at Nebius B.V. on JobCrawls
LocationActive listings
Remote - Global398
Remote - Europe126
Amsterdam, Netherlands57
London, United Kingdom22
Remote - United States20
Remote - Finland18
Berlin, Germany16
Mäntsälä, Finland12
Helsinki, Finland10
Lappeenranta, Finland10
Prague, Czech Republic8
Israel6
Amsterdam5
United Kingdom5
Canada4
Tel Aviv, Israel3
Singapore3
Remote3
Abu Dhabi2
New York City, United States2
London2
Paris, France2
Austin, United States2
Dubai2
France, Paris2
Canada, Remote - United States1
East London, United Kingdom1
Singapore, Singapore1
Minnesota, United States1
Remote - Israel1
Munich, Germany1
UK1
Berlin1
New Jersey, United States1
Remote - Germany1
Abu Dhabi, Dubai1
Prague1
New Jersey, US1
Remote - United Kingdom1
Oklahoma, United States1
Finland1
California, United States1
Abu Dhabi, United Arab Emirates1
Czechia1
Alabama, US1
Béthune, France1
Netherlands1
Austin, Texas1
San Francisco Bay Area, United States1
Kansas City, United States1
Béthune, Pas-de-Calais, France1
Philadelphia, United States1
Dallas, United States1
London, UK1
Israel, Israel1
Remote - EU1
Paris1
Current role mix at Nebius B.V. on JobCrawls
Role typeActive listings
Backend Engineer308
Software Engineer82
Account Executive54
Site Reliability Engineer6
Technical Product Manager5
Technical Project Manager5
Technical Program Manager5
Product Manager4
Sales Representative4
System Engineer4
Data Center Technician3
ML Engineer3
Data Center Operations Technician3
IT Technician2
Product Designer2
Hypervisor Engineer2
Delivery Manager2
Backend engineers, Frontend engineers, Site reliability engineers2
Open Positions at Nebius2
Applied AI Researcher2
Head of Channel Marketing1
Forward Deployment Engineer1
Site Selection & Colocation Manager1
Network Engineer1
DC IT Support Manager1
Application Security Engineer1
MEP Engineer1
Datacenter IT Technician1
Partner Solutions Architect1
Customer Engineer1
Internal Control Business Partner1
Solutions Partner1
Senior System Engineer1
Senior Technical Program Manager1
Solutions Architecture Leader1
Electrical Engineer1
Data Scientist1
Compensation Analyst1
Network Planning Project Manager1
Data Center IT Manager1
Technical Account Manager1
Security Solutions Engineer1
Cloud Solution Architect1
AI and ISV Partner Business Development Manager1
Structured Cabling Design Engineer1
ML Infrastructure Engineer1
Senior Applied AI Solutions Engineer1
VP of Developer Relations & Community1
Senior Support Engineer1
Operations Specialist1
Project Development Manager1
Backend Developer1
Data Center Operations Manager1
Data Center IT Technician1
Mechanical Design Engineer1
Generalist1
HPC Engineer1
Solutions Architect1
Data Center Electrical Lead1
Pricing Director1
Manager, ML Solutions Architecture1
Senior Technical Project Manager1
Accountant1
Product Growth Analytics Lead1
Group Product Manager1
Infrastructure Security Engineer1
Senior Research Scientist1
GTM Recruiting Manager1
Technical Due Diligence Manager1
Technical Support Engineer1
Data Engineer1
Machine Learning Engineer1
Mechanical Data Center Technician1
Sales Engineer1
Senior Software Developer1
AI/ML Specialist Solutions Architect1
Instructional Designer1
Vendor Security & Standards Manager1
ML Solutions Architect1
Physical Security Systems Technician1
Mechanical Engineer1
Vulnerability Operations Center Lead1
IT Risk and Control Manager1
VP of Strategic Sales1
Human Resources Specialist1
Educational Content Author1
Offensive Security Lead1
Data Center Logistics Specialist1
Financial Controller1
Applied ML Engineer1
Field Technical Lead1
Data Center Facilities Manager1
Principal1
Systems HPC Engineer1
Backend Engineers1
Current role-level mix at Nebius B.V. on JobCrawls
Role levelActive listings
Mid-Level387
Senior67
Manager14
Executive3

Help us improve JobCrawls — sign in to sync saved jobs across devices, or send feedback anytime.