hirly

Gimlet

Member of Technical Staff - Kernels & GPU Performance

San Francisco, CA

Apply through hirly

Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.

Apply with hirly

hirly's read of this role

Seniority
Lead / management
Country
US
Work mode
Remote-friendly
First seen by hirly
10 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

About us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

About the role

As a Member of Technical Staff, you will build and optimize the low-level execution primitives that turn accelerator performance into production inference performance.

Rather than optimizing for one hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks. Your work will shape the latency, throughput, and efficiency Gimlet can achieve across established and emerging hardware architectures.

You will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation. You will develop optimizations that account for differences between accelerator architectures and partner with compiler, ML systems, and distributed systems engineers to improve performance across the full execution stack.

What success looks like

In the first 12-18 months, you will:

Build and optimize kernels that improve latency, throughput, and hardware utilization for production AI workloads

Develop execution strategies that unlock performance across both established and emerging accelerator architectures

Improve memory efficiency, scheduling behavior, and execution characteristics across the inference stack

Partner with compiler, runtime, and distributed systems engineers to ensure end-to-end performance optimization

Influence how heterogeneous hardware is deployed and utilized within the next generation of AI infrastructure

Help establish performance engineering standards that shape the future of Gimlet's execution platform

You may be a good fit if you have

Strong software engineering fundamentals

Experience working on performance-critical systems close to hardware

Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs

Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have

Experience with CUDA, Triton, CUTLASS, or other accelerator programming models

Deep understanding of GPU execution models (warps/wavefronts, blocks, grids)

Experience optimizing memory access patterns (coalescing, shared memory, cache behavior)

Familiarity with occupancy, latency hiding, and instruction-level parallelism

Experience using profiling and performance analysis tools

Familiarity with multi-GPU or distributed execution is a plus

Why join now?

Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.

Solve hard problems.

Own meaningful work.

Build for production.

Help define what’s next.

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Original posting on Gimlet's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Member of Technical Staff - Kernels & GPU Performance at Giml… | hirly