hirly

Xai

Hardware Failure Analysis Engineer - Memphis

Southaven, MS · Memphis, TN

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Xai first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
13 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."

RESPONSIBILITIES:

Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.

Run recurrence analysis and fleet-wide defect clustering; detect systemic patterns before they become fleet-scale loss.

Build vendor technical escalation packages with evidence quality that forces action; partner with OEM/ODM/component manufacturers on firmware, bring-up, and manufacturing escapes through committed CAPA.

Feed findings into RMA policy, spare strategy, and "do not reseat forever" stop-rules.

Work the FA intake queue on rotation; keep queue age within SLA

BASIC QUALIFICATIONS:

Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).

2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.

Proven expertise in firmware analysis, hardware specifications review, and release validation.

Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.

Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.

Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.

Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.

Excellent problem-solving skills with a data-driven approach to reliability engineering.

Ability to work collaboratively with cross-functional teams, including operations technicians.

PREFERRED SKILLS AND EXPERIENCE:

Experience in AI/ML infrastructure or supercomputing environments.

Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.

Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).

Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

Original posting on Xai's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job