hirly

Genmo

GPU Performance Engineer

San Francisco HQ

Apply through hirly

Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.

Apply with hirly

hirly's read of this role

Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
10 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

We are Genmo, a research lab developing the world’s most sophisticated video world models to understand, simulate, and interact with the physical world. Our mission is to unlock the right brain of AGI. Join us in advancing physical intelligence and enabling robots to learn and act in a changing world.

We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.

The Role

You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.

Key Responsibilities

Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation

Write high-performance CUDA and Triton kernels for critical model operations

Optimize cold start latency from seconds to milliseconds for our serving infrastructure

Tune memory access patterns, kernel fusion, and GPU utilization

Collaborate with ML engineers to optimize model implementations

Debug performance issues across the full stack from application to hardware

Implement custom memory pooling and allocation strategies

Share optimization techniques and build performance culture across teams

Qualifications

Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field

5+ years systems programming experience with 3+ years focused on GPU optimization

Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)

Strong CUDA programming skills with production kernel development

Deep understanding of GPU architecture (memory hierarchy, SMs, warps)

Track record of achieving significant performance improvements (5-10x)

Experience with Python and C++ in production environments

We Value

Experience with Triton kernel development

Knowledge of CUTLASS or similar high-performance libraries

Background in ML-specific optimizations (attention, transformers)

RDMA/InfiniBand optimization experience

Contributions to GPU libraries or frameworks

Low-level debugging skills (PTX/SASS reading)

Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish .

Original posting on Genmo's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job