hirly

Inferact

Member of Technical Staff, Inference

Remote

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Inferact first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Lead / management
Country
US
Work mode
Remote-friendly
First seen by hirly
6 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Skills and Qualifications

Minimum qualifications:

Bachelor's degree or equivalent experience in computer science, engineering, or similar.

Deep understanding of transformer architectures and their variants.

Strong programming skills in Python with experience in PyTorch internals.

Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).

Ability to read and implement model architectures and inference techniques from research papers.

Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.

Preferred qualifications:

Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.

Familiarity with RL frameworks and algorithms for LLMs.

Experience with multimodal inference (audio/image/video/text).

Contributions to open-source ML or system infrastructure projects.

Bonus points if you have:

Implemented core features in vLLM or other inference engine projects.

Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).

Written widely-shared technical blogs or side projects on vLLM or LLM inference.

Logistics

Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.

Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.

Visa sponsorship: We sponsor visas on a case-by-case basis.

Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.

Original posting on Inferact's site ↗

Listed on hirly, a job board. hirly is not the employer: Inferact is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job