
IT Infrastructure Engineer
The role The IT Infrastructure Engineer – RMA & Hardware Diagnostics is responsible for advanced hardware troubleshooting and RMA lifecycle management within Nebius production data center environments. This role serves as the escalation point for complex server and firmware-related issues that impact system reliability and fleet availability. You will be performing deep diagnostics across enterprise server platforms, conducting structured root cause analysis, validating failed components, and managing end-to-end warranty replacement processes with OEM vendors. This is a hands-on technical role with direct impact on hardware reliability, SLA performance, and operational scalability. You will work on-site in Nebius data centers, collaborating closely with L1/L2 technicians, infrastructure engineers, and vendors to reduce repeat failures and improve hardware quality across the fleet. Minnesota is listed as a location. Responsibilities include: performing advanced firmware and hardware diagnostics on enterprise server platforms (CPU, memory, PCIe devices, GPUs, storage subsystems, power components); troubleshooting complex hardware failures using system logs, IPMI interfaces, BIOS diagnostics, and vendor-specific tooling; acting as the primary escalation point for L1 and L2 technicians on high-impact hardware incidents; conducting structured root cause analysis and documenting findings; owning the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution; interfacing directly with OEM vendors to escalate recurring hardware defects and drive corrective action; analyzing hardware failure trends and reporting metrics; developing and standardizing diagnostic playbooks, troubleshooting workflows, and hardware validation procedures; validating replacement components prior to redeployment into production environments; collaborating cross-functionally with data center operations, procurement, and engineering teams to improve hardware lifecycle processes; contributing to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing. We expect you to have 5+ years of hands-on experience with enterprise server hardware in a production data center environment; deep understanding of x86 server architecture; strong experience with firmware and BIOS/BMC diagnostics and upgrades; advanced Linux command-line troubleshooting; experience with remote management interfaces (IPMI, iDRAC, iLO); proven RMA management with OEM vendors; ability to perform structured root cause analysis and document findings clearly; familiarity with hardware monitoring systems and failure trend analysis; strong ownership mindset and ability to operate independently in mission-critical environments; high proficiency in spoken and written English. Bonus points: experience performing board-level diagnostics, data center networking exposure, GPU-dense HPC environments, and valid driver’s license. Compensation: We offer competitive salaries ranging from $112,700.00 - $140,800.00 OTE (On Target Earnings) based on experience and skills. Benefits & Perks include competitive compensation, career growth, flexibility and ownership, collaborative and innovative culture, opportunity to work on impactful AI projects, international environment and talented teams. What Nebius is like to work at: fast moving, bold thinking, constant growth, meaningful impact, trust and real ownership, opportunity to shape the future of AI. Equity and equal opportunity statement included.
Job Details
Responsibilities
- Perform advanced firmware and hardware diagnostics on enterprise server platforms (CPU, memory, PCIe devices, GPUs, storage subsystems, power components)
- Troubleshoot complex hardware failures using system logs, IPMI/BMC interfaces, BIOS diagnostics, and vendor-specific tooling
- Act as the primary escalation point for L1/L2 technicians on high-impact hardware incidents
- Conduct structured root cause analysis and document findings
- Own the full RMA lifecycle, including validation of failed components, warranty claim creation, vendor coordination, tracking, and resolution
- Interface with OEM vendors to escalate recurring hardware defects and drive corrective action
- Analyze hardware failure trends and report metrics such as repeat RMA rates and component reliability
- Develop and standardize diagnostic playbooks, troubleshooting workflows, and hardware validation procedures
- Validate replacement components prior to redeployment into production environments
- Collaborate with data center operations, procurement, and engineering teams to improve hardware lifecycle processes
- Contribute to reducing MTTR and improving fleet-wide reliability through process improvements and knowledge sharing
Requirements
- 5+ years of hands-on experience with enterprise server hardware in a production data center environment
- Deep understanding of x86 server architecture
- Strong experience with firmware and BIOS/BMC diagnostics and upgrades
- Advanced Linux command-line troubleshooting skills
- Experience with remote management interfaces such as IPMI, iDRAC, iLO
- Proven experience managing hardware RMA processes with OEM vendors
- Ability to conduct structured root cause analysis and document findings clearly
- Familiarity with hardware monitoring systems and failure trend analysis
- Strong ownership mindset and ability to operate independently in mission-critical environments
- High proficiency in spoken and written English
Skills & Technologies

| Location | Active listings |
|---|---|
| Remote - Global | 398 |
| Remote - Europe | 126 |
| Amsterdam, Netherlands | 57 |
| London, United Kingdom | 22 |
| Remote - United States | 20 |
| Remote - Finland | 18 |
| Berlin, Germany | 16 |
| Mäntsälä, Finland | 12 |
| Helsinki, Finland | 10 |
| Lappeenranta, Finland | 10 |
| Prague, Czech Republic | 8 |
| Israel | 6 |
| Amsterdam | 5 |
| United Kingdom | 5 |
| Canada | 4 |
| Tel Aviv, Israel | 3 |
| Singapore | 3 |
| Remote | 3 |
| Abu Dhabi | 2 |
| New York City, United States | 2 |
| London | 2 |
| Paris, France | 2 |
| Austin, United States | 2 |
| Dubai | 2 |
| France, Paris | 2 |
| Canada, Remote - United States | 1 |
| East London, United Kingdom | 1 |
| Singapore, Singapore | 1 |
| Minnesota, United States | 1 |
| Remote - Israel | 1 |
| Munich, Germany | 1 |
| UK | 1 |
| Berlin | 1 |
| New Jersey, United States | 1 |
| Remote - Germany | 1 |
| Abu Dhabi, Dubai | 1 |
| Prague | 1 |
| New Jersey, US | 1 |
| Remote - United Kingdom | 1 |
| Oklahoma, United States | 1 |
| Finland | 1 |
| California, United States | 1 |
| Abu Dhabi, United Arab Emirates | 1 |
| Czechia | 1 |
| Alabama, US | 1 |
| Béthune, France | 1 |
| Netherlands | 1 |
| Austin, Texas | 1 |
| San Francisco Bay Area, United States | 1 |
| Kansas City, United States | 1 |
| Béthune, Pas-de-Calais, France | 1 |
| Philadelphia, United States | 1 |
| Dallas, United States | 1 |
| London, UK | 1 |
| Israel, Israel | 1 |
| Remote - EU | 1 |
| Paris | 1 |
| Role type | Active listings |
|---|---|
| Backend Engineer | 308 |
| Software Engineer | 82 |
| Account Executive | 54 |
| Site Reliability Engineer | 6 |
| Technical Product Manager | 5 |
| Technical Project Manager | 5 |
| Technical Program Manager | 5 |
| Product Manager | 4 |
| Sales Representative | 4 |
| System Engineer | 4 |
| Data Center Technician | 3 |
| ML Engineer | 3 |
| Data Center Operations Technician | 3 |
| IT Technician | 2 |
| Product Designer | 2 |
| Hypervisor Engineer | 2 |
| Delivery Manager | 2 |
| Backend engineers, Frontend engineers, Site reliability engineers | 2 |
| Open Positions at Nebius | 2 |
| Applied AI Researcher | 2 |
| Head of Channel Marketing | 1 |
| Forward Deployment Engineer | 1 |
| Site Selection & Colocation Manager | 1 |
| Network Engineer | 1 |
| DC IT Support Manager | 1 |
| Application Security Engineer | 1 |
| MEP Engineer | 1 |
| Datacenter IT Technician | 1 |
| Partner Solutions Architect | 1 |
| Customer Engineer | 1 |
| Internal Control Business Partner | 1 |
| Solutions Partner | 1 |
| Senior System Engineer | 1 |
| Senior Technical Program Manager | 1 |
| Solutions Architecture Leader | 1 |
| Electrical Engineer | 1 |
| Data Scientist | 1 |
| Compensation Analyst | 1 |
| Network Planning Project Manager | 1 |
| Data Center IT Manager | 1 |
| Technical Account Manager | 1 |
| Security Solutions Engineer | 1 |
| Cloud Solution Architect | 1 |
| AI and ISV Partner Business Development Manager | 1 |
| Structured Cabling Design Engineer | 1 |
| ML Infrastructure Engineer | 1 |
| Senior Applied AI Solutions Engineer | 1 |
| VP of Developer Relations & Community | 1 |
| Senior Support Engineer | 1 |
| Operations Specialist | 1 |
| Project Development Manager | 1 |
| Backend Developer | 1 |
| Data Center Operations Manager | 1 |
| Data Center IT Technician | 1 |
| Mechanical Design Engineer | 1 |
| Generalist | 1 |
| HPC Engineer | 1 |
| Solutions Architect | 1 |
| Data Center Electrical Lead | 1 |
| Pricing Director | 1 |
| Manager, ML Solutions Architecture | 1 |
| Senior Technical Project Manager | 1 |
| Accountant | 1 |
| Product Growth Analytics Lead | 1 |
| Group Product Manager | 1 |
| Infrastructure Security Engineer | 1 |
| Senior Research Scientist | 1 |
| GTM Recruiting Manager | 1 |
| Technical Due Diligence Manager | 1 |
| Technical Support Engineer | 1 |
| Data Engineer | 1 |
| Machine Learning Engineer | 1 |
| Mechanical Data Center Technician | 1 |
| Sales Engineer | 1 |
| Senior Software Developer | 1 |
| AI/ML Specialist Solutions Architect | 1 |
| Instructional Designer | 1 |
| Vendor Security & Standards Manager | 1 |
| ML Solutions Architect | 1 |
| Physical Security Systems Technician | 1 |
| Mechanical Engineer | 1 |
| Vulnerability Operations Center Lead | 1 |
| IT Risk and Control Manager | 1 |
| VP of Strategic Sales | 1 |
| Human Resources Specialist | 1 |
| Educational Content Author | 1 |
| Offensive Security Lead | 1 |
| Data Center Logistics Specialist | 1 |
| Financial Controller | 1 |
| Applied ML Engineer | 1 |
| Field Technical Lead | 1 |
| Data Center Facilities Manager | 1 |
| Principal | 1 |
| Systems HPC Engineer | 1 |
| Backend Engineers | 1 |
| Role level | Active listings |
|---|---|
| Mid-Level | 387 |
| Senior | 67 |
| Manager | 14 |
| Executive | 3 |
Related Opportunities
Discover more opportunities that match your interests and skills