Veeda AI
Member of Technical Staff - ML Performance
Toronto · Zürich · Seattle · California
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Lead / management
- Countries
- CA, CH, US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Member of Technical Staff - ML Performance
About Us
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
Responsibilities
Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default.
Inference Throughput Optimization: Profile end-to-end latency, GPU utilization, memory pressure, and kernel efficiency with PyTorch Profiler, Nsight Systems, Nsight Compute, and torch.utils.benchmark, then improve throughput through torch.compile/Inductor, CUDA Graphs, mixed and low precision, quantization, operator fusion, and multi-GPU serving.
Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live.
Kernels & Compilation: Write and tune the CUDA and Triton kernels PyTorch does not give us, driving FlashAttention-4, FlexAttention, and torch.compile integration so quadratic attention over long video sequences stops setting step time.
Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric, using the NCCL flight recorder to turn a watchdog timeout into a named rank and collective, not a restart.
Fault Diagnostics & Recovery: Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs, plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day.
Requirements
Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.
Deep hands-on experience with PyTorch and at least one large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed) on real multi-node jobs, not single-node approximations.
Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling.
Experience profiling live training and inference runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.
Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis, and credibility in the others.
Nice to Have
Experience optimizing training or inference workloads across very large GPU clusters, including topology-aware placement, scaling efficiency, performance isolation, and diagnosing failures that emerge only at fleet scale.
Experience writing Triton, CUTLASS, or CuTe-DSL kernels, or contributing to open-source kernel libraries.
Experience implementing context or sequence parallelism for long-horizon video or high-token-count models.
Experience running or porting large training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA).
Experience building fault-tolerant training with elastic world size, dynamic node re-queueing, or asynchronous distributed checkpointing.
Experience optimizing generative inference for interactive rollouts, including few-step samplers, distillation, and KV or latent caching.
Publications or presentations on machine learning systems, compilers, or high-performance kernels.
Similar jobs
- Sr. Member Technical Staff - ESD and Latch-Up - HBMMicron · Folsom, CAFirst seen 4d ago
- Member Technical StaffPirros · Los Angeles OfficeFirst seen 29d ago
- Member Technical Staff - Applied AI Engineer (US Timing) Composio · BangaloreFirst seen 23d ago
- Member TechnicalBroadridge · Bengaluru-EPIP Industrial AreaFirst seen 6d ago
- Senior Member TechnicalBroadridge · Hyderabad-Hi-Tec CityFirst seen 8d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job