hirly

Weights And Biases

Senior Software Engineer, Infrastructure Operations - Weights & Biases

Livingston, NJ / New York, NY / San Fransisco, CA / Sunnyvale, CA / Bellevue, WA

Apply through hirly

hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
US
Work mode
On-site / unstated
First seen by hirly
13 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

CoreWeave, the AI Hyperscaler™, acquired Weights & Biases to create the most powerful end-to-end platform to develop, deploy, and iterate AI faster. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe, and was ranked as one of the TIME100 most influential companies of 2024. By bringing together CoreWeave’s industry-leading cloud infrastructure with the best-in-class tools AI practitioners know and love from Weights & Biases, we’re setting a new standard for how AI is built, trained, and scaled.

The integration of our teams and technologies is accelerating our shared mission: to empower developers with the tools and infrastructure they need to push the boundaries of what AI can do. From experiment tracking and model optimization to high-performance training clusters, agent building, and inference at scale, we’re combining forces to serve the full AI lifecycle — all in one seamless platform.

Weights & Biases has long been trusted by over 1,500 organizations — including AstraZeneca, Canva, Cohere, OpenAI, Meta, Snowflake, Square,Toyota, and Wayve — to build better models, AI agents and applications. Now, as part of CoreWeave, that impact is amplified across a broader ecosystem of AI innovators, researchers, and enterprises.

As we unite under one vision, we’re looking for bold thinkers and agile builders who are excited to shape the future of AI alongside us. If you're passionate about solving complex problems at the intersection of software, hardware, and AI, there's never been a more exciting time to join our team.

What You'll Do

The Infra Ops team is a new team in the Execute pillar of Weights & Biases infrastructure. Our mission is to keep our platform engineers building: we absorb, reroute, and automate incoming infrastructure requests — incidents, how-to questions, troubleshooting requests, and one-off asks — so that other pillars can focus on their deliverables. Build and release engineering is part of the Execute pillar — an organized, safe, and robust release process is the natural way to reduce post-release incidents, rollbacks, and toil. We work across the full W&B /CoreWeave infra stack — Kubernetes, Go, Terraform, ClickHouse, CircleCI, ArgoCD, Argo Rollouts, GitHub Actions, and others — running on GCP, AWS, and Azure, across multi-tenant SaaS, dedicated cloud, and on-prem deployments.

About the Role

We are seeking an Infrastructure Operations Engineer to be part of the first line of support for the infrastructure org. This is a support-oriented, interrupt-driven role: your days are shaped by incoming requests from our internal customers — Solutions Engineers, security, compliance, and product developers on other teams — rather than by a single long-running project. With guidance from senior teammates, you'll triage and resolve many requests end to end, and help convert recurring issues into documentation and automation to prevent repeats. It's a fast way to learn the entire infrastructure surface, and a natural "landing pad" into infrastructure engineering. You'll write code most days — automation and scripts in Python, Bash, and Go — and take part in a first-responder rotation covering the daily request peak in Slack and office hours. Your customers are primarily internal, though you'll occasionally work a customer escalation when our merchant-support and SA teams need help.

In this role, you will:

Triage and resolve incoming infrastructure requests across Slack, Jira, and office hours — resolving well-scoped ones yourself and escalating others with a clear, reproducible hand-off.

Troubleshoot problems across the stack — dig through logs, systems, and configuration to work out what is actually happening, asking for help when a problem runs deep.

Write scripts and other automation in Python, Bash, and Go that reduce repetitive work.

Help maintain and improve build and release pipelines, along with the runbooks and self-service docs the team depends on.

Partner and collaborate closely with our SA, security, and product engineers on one side, and with the other infra pillars on the other.

Take part in a 24/7 escalation on-call rotation

Document recurring how-to and configuration questions so they can be answered once and reused.

Who You Are

Bachelor’s degree or foreign equivalent in Computer Science, Computer Engineering, Data Science, Information Systems, or a closely related technical or quantitative field involving substantial coursework in Engineering.

Experience: 2+ years in an infrastructure, SRE, operations, DevOps, or technical-support role — or equivalent experience from a computer-science (or related) degree, a bootcamp, or substantial personal/open-source projects. We hire on demonstrated ability, not credentials.

Scripting: Comfortable scripting in at least one of Python, Bash, or Go, and eager to use code to eliminate repetitive work.

Systems foundation: Experience with Linux, containers, and at least one major public cloud (GCP, AWS, or Azure); infrastructure-as-code (Terraform) is a must.

Able to drive well-scoped problems to resolution, and to recognize when to pull in a more senior teammate.

Strong written and verbal communication and a service mindset — you communicate effectively with our internal customers.

A genuine desire to help people and unblock them — you take satisfaction in solving someone else's problem, not just your own.

Patience and composure under pressure, with the ability to stay positive and constructive.

Comfortable with a support-driven, high-context-switching, and sometimes repetitive workload — you can hold several small threads at once without losing the plot, and you drive the automation of recurring patterns.

Aim to make yourself unnecessary — automate everything to free time for larger infrastructure projects at the edges of other infrastructure pillars. We have much ground to cover.

Preferred

Exposure to CI/CD systems (GitHub Actions, Argo, or similar) and observability tooling (Datadog).

Applied AI: Hands-on experience designing, building, and maintaining complex, production-grade AI agents — including continuous multi-source data ingestion and indexing, and RAG / context augmentation.

Hands-on experience with production databases (PostgreSQL, ClickHouse, or similar).

Prior internship or role on a platform, developer-experience, or support-engineering team.

Wondering if you're a good fit?

We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk.

You love helping and unblocking other people, and feel the satisfaction of a support queue going quiet.

You're curious about turning repetitive toil into automation and self-service rather than absorbing it forever.

You're eager to learn a large, real-world infrastructure stack quickly.

Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

Be Curious at Your Core

Act Like an Owner

Empower Employees

Deliver Best-in-Class Client Experiences

Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

The base salary range for this rol

Browse similar roles

Is this role actually a fit for you?

hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.

Score it against my resume
Senior Software Engineer, Infrastructure Operations - Weights & Biases at Weights And Biases — hirly