hirly

Gimlet

Member of Technical Staff - ML Systems & Inference

San Francisco, CA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Gimlet first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Lead / management
Stated salary
$150,000 – $390,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
10 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About us

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

About the role

As a Member of Technical Staff focused on ML Systems, you will build the inference systems that execute models end-to-end in production.

You will work on the systems that determine how inference executes across that pipeline: how requests are batched and scheduled, how stages are placed and scaled, how KV cache and intermediate state move between accelerators, and how the system balances latency, throughput, and utilization across different hardware characteristics.

You will work across model serving, batching, scheduling, concurrency, KV cache management, and memory placement. You will help bring up models on novel hardware. You will support new model architectures and inference techniques, improve performance under real production workloads, and partner with compiler, kernel, networking, and distributed systems engineers to optimize the full execution path.

What success looks like

In the first 12-18 months, you will:

Improve the latency, throughput, and efficiency of production inference workloads

Design execution strategies across batching, scheduling, concurrency, and resource utilization

Improve KV cache management, memory efficiency, and execution under load

Enable new models, accelerator architectures, and inference techniques to run efficiently in production

You may be a good fit if you have

Strong software engineering fundamentals

Experience building or operating ML inference or model serving systems

Comfort reasoning about performance, memory usage, and system behavior under load

Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have

Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems

Deep understanding of modern model architectures and attention mechanisms

Experience with batching, scheduling, and concurrency control in inference systems

Familiarity with KV cache management and memory placement strategies

Experience profiling and tuning latency- and throughput-critical systems

Software development experience in Python and C++

Why join now?

Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.

Solve hard problems.

Own meaningful work.

Build for production.

Help define what’s next.

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Original posting on Gimlet's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job