hirly

Vast.ai

Systems Operations Support Engineer — Linux

Los Angeles

Apply through hirly

Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.

Apply with hirly

hirly's read of this role

Role family
Customer support
Seniority
Mid level
Stated salary
$90,000 – $150,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

About Us

Vast.ai 's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This is a systems operations support role focused on deep-diving into escalated infrastructure issues that go beyond frontline triage. You’ll be the engineering resource our L1 support team relies on when tickets become complex, investigating and resolving issues across the full infrastructure stack—including hardware, BIOS and firmware, networking, Ubuntu, Docker, NVIDIA CUDA and GPUs, and KVM virtual machines.

You’ll own complex escalations end-to-end: gathering evidence, reproducing issues, identifying the root cause, proposing solutions, and working with the appropriate teams to bring each issue to resolution. The best engineers in this role don’t just resolve individual tickets—they identify recurring patterns, improve operational tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic infrastructure issues.

Strong Linux systems knowledge, technical depth, and support experience are the primary requirements. You should be comfortable working autonomously in Ubuntu environments, troubleshooting hardware, networking, containers, virtual machines, and GPU workloads, and clearly communicating your findings and proposed solutions to both technical and non-technical audiences.

Vast.ai users or hosts strongly preferred.

Location and Schedule

This is a full-time position based in our Westwood, Los Angeles office.

Available schedules:

Monday–Friday: Fully on-site

Sunday–Thursday: Four days on-site and one day working from home

Key Responsibilities

Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting

Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads

Provide coverage for L1 support overflow during peak periods or incidents

Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and KVM virtualization environments

Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

Investigate performance issues involving GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks

Advise suppliers on installation best practices, including hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance

Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues

You Are

Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line

Someone who enjoys debugging difficult problems and fixing broken systems

Methodical and focused on finding root causes, not just temporary fixes

Able to manage complex tickets independently

A clear written communicator with an interest in AI infrastructure and GPU computing

Must-Haves

Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions

Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting

Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting

Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting

Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting

Python and Bash scripting skills for automation and diagnostic tooling

Strong written English communication that is clear, professional, and technically precise

Experience providing technical support in a customer-facing or internal help desk environment

Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice-to-Haves

Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers

Monitoring and observability experience (Prometheus, Grafana)

Relevant certifications: RHCSA, CompTIA Linux+, or similar

Knowledge of the Vast.ai platform as a client or infrastructure supplier

Interview Process (~1 week)

After you submit your application, our technical team will review your experience and qualifications. Selected candidates will proceed through the following stages:

15 minutes — Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role

45 minutes — Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience

2 hours — Meet and Greet and Technical Assessment (On-site): Meet the team and complete an LLM-assisted Linux systems operations assessment

Annual Salary Range

$90,000 – $160,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits

Comprehensive health, dental, vision, and life insurance

401(k) with company match

Meaningful early-stage equity

Onsite meals, snacks, and close collaboration with founders/tech leaders

Ambitious, fast-paced startup culture where initiative is rewarded

Original posting on Vast.ai's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job