TensorWave
Hardware Diagnostics Engineer - Infrastructure
Las Vegas, Nevada
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 1 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About TensorWave
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
About the Role
We are looking for a Hardware Diagnostics Engineer to run burn-in, triage what fails, work servers out-of-band, and own RMAs end to end. If you like hardware that misbehaves in ways that take real work to explain, this is a good seat.
Before any GPU server carries a customer workload, it has to prove it works — under load, at temperature, for hours. When it doesn't, somebody has to figure out why, get replacement hardware in, and send the failed part back to the vendor.
What You’ll Do
Run server and GPU burn-in and stress testing, interpret the results, and decide whether hardware is production-ready
Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling: reproduce the failure, isolate the faulty component, and document what proved it
Work servers out-of-band through IPMI and Redfish for power control, boot configuration, BIOS settings, and sensor and event log collection
Apply firmware updates across the fleet following the team's qualified baselines and rollout process
Drive RMAs with vendors from ticket through replacement, installation, and return of the failed part
Keep asset, serial, and replacement history accurate in NetBox so we know what's actually in every rack
Track failure patterns across the fleet and raise them when the same part or firmware version keeps turning up
Improve the runbooks you work from, and script the steps you find yourself repeating
Partner with datacenter operations on hands-on work during turn-ups and expansions
Take part in an on-call and escalation rotation for hardware issues
Who You Are
Required Qualifications
3–6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering
Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, and reading hardware behavior at the rack
Practical experience with BMCs and out-of-band management: IPMI, Redfish, iDRAC, iLO, or equivalent
Strong Linux troubleshooting: boot process, driver and device issues, and diagnostic tools such as {{dmesg}}, {{lspci}}, {{ipmitool}}, and SMART
Comfort reading sensor data, event logs, and thermal and power telemetry well enough to tell a real failure from noise
Working scripting ability in Bash or Python — enough to automate a repetitive task and read someone else's tooling
Experience running hardware RMAs with vendors, or a clear track record of driving issues to closure with outside parties
A methodical troubleshooting habit: you isolate variables, you don't change three things at once, and you can say what evidence led to your conclusion
Clear written communication for tickets, runbooks, and vendor cases
Preferred Qualifications
GPU server experience, especially AMD GPUs and ROCm
Burn-in, stress testing, or node validation tooling in a GPU or HPC environment
Familiarity with firmware update processes and why fleet-wide changes get staged
NetBox or other DCIM and IPAM tooling
Ansible, or Python against REST APIs
Prior work in a high-volume hardware environment: hyperscaler, colo, integrator, or manufacturing test
First Six Months
By 90 days you'll run burn-in cycles and triage failures independently from our runbooks, and you'll have driven at least one RMA to closure. By six months you're the person who spots the pattern before anyone else — this batch, this firmware, this part — and you've automated at least one step you used to do by hand.
What We Offer
Stock Options
100% paid Medical, Dental, and Vision insurance for Employees
Company Health Savings Account Contributions
100% paid Short Term and Long Term Disability Insurance for Employees
Life and Voluntary Supplemental Insurance Options
Other Insurance Options, such as Pet & Legal Insurance
Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
Flexible Spending Account
401(k)
Employee Assistance Program
Flexible PTO
Paid Holidays
Parental Leave
Other In-Office Perks
Equal Employment Opportunity
TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.
Reasonable Accommodations
TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].
Employment Eligibility
All offers of employment are contingent upon verification of identity and authorization to work in United States, as required by law.
Background Checks
Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.
Data Privacy Notice
By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.
Similar jobs
- High-Enthalpy Diagnostics EngineerAmainc · Mountain View, CAFirst seen yesterday
- Americas Monitoring & Diagnostics Engineering LeaderGevernova · 2 LocationsFirst seen 4d ago
- Prognostics & Diagnostics EngineerEmerson · Marshalltown, IA, United StatesFirst seen 4d ago
- Vehicle Diagnostics Engineer (ECU / CAN Network)Sonyhondamobilityofamerica · Torrance, CAFirst seen 15d ago
- Diagnostics EngineerNexthopai · SF Bay Area, CaliforniaFirst seen 29d ago
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job