hirly

Veeda AI

Member of Technical Staff - ML Operations

Toronto · Zürich · Seattle · California

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Veeda AI first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Lead / management
Countries
CA, CH, US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

ABOUT US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

RESPONSIBILITIES

Experiment Lifecycle Tracking and Tooling: Design, build and deploy tools that own how a run is defined, launched, resumed, and killed. Build and operate experiment databases for code version tracking, data version tracking, reproducibility, and checkpoint ancestry.

Inference Fleet Orchestration: Design, build, and operate the serving control plane that accepts large volumes of concurrent client requests and assigns them across inference clusters. Develop cache-aware admission, routing, batching, and scheduling policies that improve cache locality, balance workload, protect tail latency, and keep the fleet highly utilized and reliable. Partner with ML Performance on model runtime, kernel, and per-worker throughput optimization.

Model Evaluation in CI: Design, build, and operate automatic model checkpoint evaluation systems on seeded rollout and policy-success suites, run per-change and nightly.

Data Pipeline Operations : Design, build, and operate high-performance, fault-tolerant, distributed backend services and event-driven systems for our large scale data processing pipeline.

Visualization Platforms: Design, build, and deploy experiment observability and dataset visualization platform( s) that provides interactive data visualization, progress tracking, search, and comparison.

End-to-End Ownership: Lead projects through the complete software lifecycle, including technical specs, implementation, CI/CD, on-call support, and production observability.

REQUIREMENTS

Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.

Proficiency in at least one scripting (Python or Bash) and one compiled (Java, Rust, or Go) languages.

Experience in shipping production-quality developer tools.

Proficient in CICD automations (pipelines, runners, deployment)

Experience in building reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.

One of the following:

Proven understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns.

Experience building or operating large-scale inference control planes or distributed serving infrastructure, with hands-on work in traffic management, admission control, request scheduling, routing, batching, or cache-aware load balancing; able to reason clearly about cache locality, queueing, tail latency, availability, and fleet utilization.

Experience working with multi-node workloads and building around slurm based scheduling systems.

Experience working on in-production model evaluation frameworks, in particular regression testing of large multimodal models.

NICE TO HAVE

You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars.

You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail.

Full stack development experience with web-based front-end.

You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks.

You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams.

You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.

Original posting on Veeda AI's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job