hirly

CoreWeave

Senior Hardware Engineer, Server Infrastructure

New York, NY/ Bellevue, WA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at CoreWeave first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Senior
Country
US
Work mode
Remote-friendly
First seen by hirly
30 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com .

About the Role

CoreWeave is seeking a highly skilled and motivated engineer to join our Hardware Engineering team. In this role, you will help design, develop, and optimize our server hardware infrastructure. You’ll collaborate closely with cross-functional teams, external vendors, and key stakeholders to deliver performant, reliable, and scalable hardware solutions that power CoreWeave’s rapidly growing infrastructure.

You will own server hardware from provisioning through decommission. This hands-on role combines engineering and operational support. Engineering work includes automation across the hardware lifecycle, hardware and firmware management services, monitoring and alerting, and qualification and bring-up of new platforms. Operational support includes acting as a senior point of contact for hardware escalations, driving deep root-cause analysis across hardware and firmware, and working quality and RMA issues through to resolution with server vendors and OEMs.

Both engineering and operational support are core responsibilities. You will support the systems you build and use what you learn from production failures to improve automation, telemetry, and platform design. You will work closely with data center operations teams, hardware technicians, and engineering teams to bring new regions online and keep existing infrastructure healthy.

What You’ll Do

Design and develop server hardware infrastructure to support CoreWeave’s high-performance workloads.

Automate all aspects of the server hardware lifecycle, from provisioning and configuration through firmware management, monitoring, and decommissioning.

Develop and maintain hardware and firmware management services that ensure reliability at scale.

Develop and implement monitoring and alerting for server hardware health, improving alert quality to support reliable on-call response.

Serve as a senior point of contact for hardware escalations, performing deep troubleshooting and root-cause analysis across hardware and firmware to drive long-term fixes.

Participate in an on-call rotation for hardware escalations and improve the runbooks, alerts, and tooling that make the rotation sustainable.

Collaborate with cross-functional teams to define hardware requirements, specifications, and system architecture.

Work with server vendors and OEMs to evaluate, qualify, and deploy new platforms and resolve firmware, quality, and RMA issues.

Support new data center region bring-up and hardware qualification.

Analyze hardware system performance, identify bottlenecks, and implement improvements to efficiency and resilience.

Establish and continuously refine processes for internal hardware testing, deployment, and performance optimization.

Create and maintain accurate documentation of hardware designs, specifications, test procedures, and results.

Support data center operations teams and hardware technicians with troubleshooting guidance, runbooks, and training so common issues can be resolved without engineering escalation.

Turn recurring production failures into automation, better telemetry, and improvements to platform design and vendor solutions.

Communicate status, trade-offs, and risks clearly to engineering, operations, and customer-facing stakeholders, including during active incidents.

Who You Are

Deep understanding of server hardware, components, and management technologies.

Proficiency in Ansible or Python, with hands-on experience programmatically interacting with server BMCs using Redfish or IPMI; Redfish preferred.

Experience collaborating with hardware vendors and OEMs to evaluate, qualify, and deploy server solutions.

Demonstrated experience supporting and troubleshooting production infrastructure, participating in on-call or escalation rotations, and driving incidents through to root-cause resolution.

Comfort with both building and automating systems and supporting infrastructure already in production.

Proven ability to stay current with technologies and trends in server and data center hardware.

Strong interest in automation and infrastructure scalability, with a commitment to continuous improvement.

Excellent technical documentation skills and attention to detail.

Strong analytical and problem-solving abilities, with a bias toward systematic, data-driven decisions.

Excellent written and verbal communication skills in English, with the ability to work effectively with technical teams and cross-functional stakeholders.

Preferred Qualifications

Experience bringing up new data center regions or standing up infrastructure in a new geography.

Experience with GPU platforms and rack-scale systems such as NVIDIA GB200 or GB300.

Experience with BMC technologies and management interfaces such as Redfish or IPMI at fleet scale.

Experience with firmware lifecycle management or hardware qualification programs in large-scale environments.

Strong Linux systems administration and debugging skills at fleet scale.

Familiarity with Kubernetes-based services, observability tools such as Prometheus and Grafana, and distributed production environments.

Experience designing alerting, runbooks, and self-service tooling that enable operations teams to resolve common hardware failures without engineering escalation.

Experience supporting external customers or partners in a technical troubleshooting capacity.

Wondering if You’re a Good Fit?

We believe in investing in our people and value candidates who bring diverse experiences to our teams—even if you aren’t a 100% skill or experience match. If some of this describes you, we’d love to talk.

You enjoy solving problems at the boundary of hardware, firmware, and software.

You want to both build systems and support them, and you use operational experience to improve what you build.

You enjoy tracing intermittent hardware failures to their root cause and validating that a fix addresses the underlying problem.

You look for opportunities to automate repetitive hardware procedures.

You work effectively with operations, engineering, and vendor teams to resolve complex technical issues.

You bring structure, accountability, and momentum to fast-moving environments.

What We Offer

The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:

Medical, dental, and vision insurance - 100% paid for by CoreWeave

Company-paid Life Insurance

Voluntary supplemental life insurance

Short and long-term disability insurance

Flexible Spending Account

Health Savings Account

Tuition Reimbursement

Ability to Participate in Employee Stock Purchase Program (ESPP)

Mental Wellness Benefits through Spring Health

Family-Forming support provided by Carrot

Paid Parental Leave

Flexible, full-service childcare support with Kinside

401(k) with a generous employer match

Flexible PTO

Catered lunch each day in our office and data center locations

Original posting on CoreWeave's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job