Avride
Software Engineer – ML Platform
Austin, Texas
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Mid level
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 10 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About the team
The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use — scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle.
About the role
As an ML Platform Engineer at Avride, you'll own critical pieces of the ML stack: workflow orchestration, distributed execution, resource governance, performance.You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience.
What you will do
Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration
Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance — scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and IO
Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention
Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes
Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs
What you will need
Strong proficiency in Python or Go; C++ is a plus
Track record of designing and building scalable, maintainable systems and services
Experience operating production services end-to-end: APIs, reliability practices, observability
Deep knowledge of Kubernetes: how scheduling, resource management, controllers, and pod lifecycle actually behave under pressure
Solid Linux and systems debugging skills: performance investigation, networking, storage/IO
Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution
Nice to have
Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling
Hands-on experience building or operating large-scale ML training systems: GPU scheduling, distributed training, training data pipelines
Track record of optimizing resource usage and performance in distributed environments
#LI-MS1
Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.
Avride is an equal opportunity employer and committed to providing reasonable accommodations to qualified applicants and employees with disabilities to ensure they have equal access to employment opportunities. Avride complies with the Americans with Disabilities Act (ADA), if you need a reasonable accommodation to assist with the application or hiring process, or to perform the essential functions of a job, please email [email protected] .
Similar jobs
- Systems Software Engineer IACCO Engineered Systems · Pleasanton, CA, United StatesFirst seen today
- Software EngineerCapgemini · New York, USFirst seen today
- Software EngineerSpatial Front, Inc · College Park, MDFirst seen today
- Quantitative Software EngineerPmsi · Henderson, NevadaFirst seen today
- Software Engineer II – Java/Spring boot Engineer (Microservice Systems)WorldpayFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job