Wayve
Staff ML Performance Engineer (Inference Optimisation)
London, United Kingdom
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Lead / management
- Country
- GB
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Before the detail, here's the challenge you'd help us solve.
We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that.
Here’s what this particular role covers.
The role
As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus of this team is to run large transformer-based models efficiently on low-cost, low-power edge devices to enable Wayve’s first driving product.
You’ll help set the technical direction for turning these models into production systems that run reliably on in-vehicle compute. This is a hands-on role working across ML systems, compilers, runtimes, kernels, and embedded deployment, contributing to several early-stage, high-impact projects at Wayve.
Key responsibilities:
Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory movement) and deliver measurable improvements.
Implement and validate optimisations in compilers, runtimes, and/or kernels (e.g. operator fusion, scheduling, quantisation-aware performance, custom kernels).
Build robust benchmarking and regression testing to ensure performance improvements hold across models, devices, and software releases.
Optimise for multiple targets (e.g. NVIDIA Orin/Thor, Qualcomm) and work with teams to support these in a maintainable way
Collaborate with model developers to influence architecture and training/deployment decisions that affect on-device performance.
Contribute to technical roadmaps and tooling and help raise the standard of performance engineering across the team
About you
Essential
Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly.
Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution.
Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code).
Clear communicator and collaborative teammate; able to align multiple stakeholders on performance trade-offs and priorities.
Desirable
Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints.
Experience with NVIDIA and/or Qualcomm SoCs and performance tooling.
Python and C++ proficiency.
Experience mentoring others and/or driving technical direction in a small, fast-moving team.
This is a full-time role based in our office in London. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home.
#LI-HH1
A quick, honest note before you apply.
Wayve is not a mature, fully-structured place with the playbook already written. Much of how we work is still being written, and if you join, you’ll help write it. That suits people who want real ownership more than people who need a settled structure from day one.
If that sounds like the kind of problem you want to spend your time on, we’d really like to hear from you.
Similar jobs
- Performance Engineering ManagerGresearch · London, UKFirst seen 2d ago
- Member of Technical Staff, Training Performance EngineerCohere · LondonFirst seen 27d ago
- Performance Engineer source.dev · London First seen yesterday
- NLP Performance EngineerGresearch · London, UKFirst seen 2d ago
- Cloud Performance EngineerDrweng · LondonFirst seen 15d agoremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job