hirly

Nscale

Staff AI Product Engineer

Houston · New York · San Francisco · Seattle

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Nscale first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Product management
Seniority
Lead / management
Stated salary
$220,000 – $330,000 per year
Country
US
Work mode
Remote-friendly
First seen by hirly
3 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About the Role

Nscale is looking for a Staff AI Engineer (Specialised) to set technical direction for part of the inference platform at the core of our AI cloud: dedicated and serverless inference, bring-your-own-model deployments, and the post-training services that sit alongside them.

You’ll be the technical authority for one or more focus areas, working across the teams that build serving, post-training, and platform. Your decisions shape the latency, throughput, and cost of the tokens Nscale serves, and whether we can say with evidence that the models we run are the right ones. You’ll take on open questions, such as how KV cache should move between GPUs, nodes, and storage, or where hand-written kernels beat the open-source defaults, and your answers become the standards others build on.

We run state-of-the-art GPU systems, and much of the open-source ecosystem hasn’t caught up with them yet. This role is for people who want to close that gap hands-on while growing the engineers around them.

How We Work

Dog years. We move quickly and compress a lot of learning into a short time.

Don’t let perfect be the enemy of good. Ship, measure, iterate.

Be relentless. Own the problem end to end and see it through.

One team, one mission. Outcomes over process, and no “not my job”.

Focus Areas

We’re hiring for depth. You don’t need all of these; we want people who are exceptional in one or more , and we’re deliberately hiring people with different focus areas:

KV cache offloading and inference performance: KV cache orchestration across GPU, host memory, and storage; cross-instance cache sharing (LMCache or similar); KV-aware routing; disaggregated prefill/decode; speculative decoding; quantisation (FP8, NVFP4, INT8/4); MoE serving

GPU kernels and performance: writing and tuning kernels in CUDA, Triton, CUTLASS or ROCm; attention, MoE and GEMM optimisation; multi-GPU and multi-node communication on NVLink-scale systems

Evals and benchmarking: designing evaluation frameworks, turning customer requirements into custom benchmarks, and producing objective model comparisons we can stand behind

Post-training and RL: fine-tuning and preference optimisation as a service; RL for LLMs (PPO/GRPO-style, reward modelling, multi-turn and tool-use RL); rollout generation, weight synchronisation, and sharing infrastructure between inference and training

Serving platform and APIs: serving engines (vLLM, SGLang, TensorRT-LLM), model onboarding, and OpenAI-compatible APIs that other engineers and customers build on

Responsibilities

Set technical direction for your focus areas and turn it into work that multiple teams can deliver

Lead the resolution of systemic performance and reliability problems across the serving stack, from kernel bottlenecks to fleet-level capacity and multi-tenant isolation

Own the trade-offs between cost, latency, throughput, and model quality in your area, and back them with measurement

Build reusable frameworks, benchmarks, and tooling that make other AI engineers at Nscale more effective

Evaluate emerging serving engines, kernel libraries, RL frameworks, and accelerators, and make clear build/adopt/contribute recommendations

Coach and grow engineers across teams, and raise engineering quality broadly

Work with research, product, and infrastructure leadership so the platform tracks customer demand

Represent your area in cross-team technical reviews and planning

Requirements

8–12 years of engineering experience, with significant depth in production AI systems or ML research at scale

4+ years of hands-on work with LLMs in at least one of the focus areas above, in production or research

Deep, demonstrated expertise in at least one of the focus areas above

Enough working knowledge of the full serving stack to debug across it: API, router, scheduler, engine, kernel, and cluster

A track record of setting technical direction and creating standards or tools adopted beyond your own team

Strong Python and PyTorch, with a track record of maintainable, well-tested, production-grade systems

Experience operating AI workloads in containerised, distributed environments (Kubernetes, large GPU clusters)

Solid understanding of GPU performance: memory bandwidth, interconnect, and parallelism strategies (data, tensor, pipeline, expert)

Deep understanding of transformer and LLM architectures and how they behave under production load

Preferred

Contributions to widely used open-source inference, kernel, or RL projects (vLLM, SGLang, TensorRT-LLM, LMCache, FlashInfer, Triton, verl, OpenRLHF, TRL, DeepSpeed, etc.)

Experience at an AI lab, hyperscaler AI team, or leading ML infrastructure company

Published research or technical writing on inference, kernels, evals, or RL (MLSys, NeurIPS, ICLR, blog posts)

Hands-on custom kernel work (CUDA, Triton, CUTLASS) if kernels aren’t already your focus area

Experience designing developer-facing APIs and SDKs used by external customers

Experience with control plane / data plane separation and cell-based architecture for scale-out and blast-radius isolation

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000 — $330,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Original posting on Nscale's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job