hirly

Coreweaveu

Operations Engineer, MetalDev

Warsaw, Poland

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Coreweaveu first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Mid level
Country
PL
Work mode
Remote-friendly
First seen by hirly
14 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com .

We're proud to be a Living Wage accredited Employer.

About the Role

CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management.

You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work.

You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation.

Key Responsibilities

Production Support and Troubleshooting

Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures.

Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately.

Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data.

Perform and validate approved remediation, ensuring services and devices return to a healthy state.

Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions.

Participate in the team's on-call rotation after completing onboarding and a readiness review.

Observability and Reliability

Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems.

Maintain and improve dashboards, alerts, and operational KPIs for team-owned services.

Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements.

Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent.

Validate software, firmware, or process changes in test environments and support controlled production rollouts.

Documentation and Collaboration

Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation.

Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence.

Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution.

Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes.

Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain.

Minimum Qualifications

Two or more years of experience in technical support, systems administration, cloud operations, site reliability engineering, infrastructure operations, or a related technical field, or equivalent practical experience.

Working knowledge of Linux system administration, including logs, processes, services, filesystems, networking fundamentals, and command-line troubleshooting.

Working knowledge of Kubernetes, containers, and at least one public cloud platform or comparable distributed infrastructure environment.

Experience using monitoring and observability tools such as Prometheus and Grafana; familiarity with PromQL or a similar query language.

Scripting experience in Bash, or another shell scripting language.

Experience troubleshooting issues across software services, operating systems, networks, or physical infrastructure.

A methodical approach to problem solving, strong documentation skills, and clear written and verbal communication.

Experience participating in an on-call rotation or supporting time-sensitive production issues.

Preferred Qualifications

Experience supporting Kubernetes applications or distributed production services.

Experience with incident response, escalation, root-cause analysis, or post-incident reviews.

Familiarity with servers, BMCs, Redfish, IPMI, server provisioning, or hardware lifecycle management.

Exposure to GPU infrastructure, DPUs, high-performance computing, power systems, cooling systems, or data center operations.

Experience improving dashboards, alerts, PromQL queries, run-books, or operational automation.

Experience collaborating with hardware or firmware vendors to investigate and validate fixes.

Bachelor's degree in computer science, engineering, information technology, or a related discipline, or equivalent practical experience.

To fulfill our obligation to protect client data, successful applicants offered employment with CoreWeave will be required to complete a basic criminal record check, conducted in compliance with GDPR. Employment offers are conditional upon receiving satisfactory check results.

What We Offer

In addition to a competitive salary, we offer a variety of benefits to support your needs, including:

Family-level Medical Insurance

Family-level Dental Insurance

Generous Pension Contribution

Life Assurance at 4x Salary

Critical Illness Cover

Employee Assistance Programme

Tuition Reimbursement

Work culture focused on innovative disruption

Benefits may vary by location.

Equal Opportunity

CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.

Recruitment Agencies

CoreWeave does not accept speculative CVs. Any unsolicited CVs received will be treated as the property of CoreWeave and your Terms & Conditions associated with the use of CVs will be considered null and void.

Any unsolicited CVs sent by your company to us – that is to say, in any situation where we have not directly engaged your company in writing to supply candidates for a specific vacancy – will be considered by us to be a “free gift”, leaving us liable for no fees whatsoever should we choose to contact the candidate directly and engage the candidate’s services, and will in no way establish any prior claim by your company to representation of that candidate should the candidate’s details also be submitted by any other party.

Export Control Compliance

This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card

Original posting on Coreweaveu's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job