Inferact
Member of Technical Staff, Inference
Remote
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 6 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Deep understanding of transformer architectures and their variants.
Strong programming skills in Python with experience in PyTorch internals.
Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
Ability to read and implement model architectures and inference techniques from research papers.
Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.
Preferred qualifications:
Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
Familiarity with RL frameworks and algorithms for LLMs.
Experience with multimodal inference (audio/image/video/text).
Contributions to open-source ML or system infrastructure projects.
Bonus points if you have:
Implemented core features in vLLM or other inference engine projects.
Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
Written widely-shared technical blogs or side projects on vLLM or LLM inference.
Logistics
Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.
Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.
Listed on hirly, a job board. hirly is not the employer: Inferact is hiring for this role.
Similar jobs
- Sr. Member Technical Staff - ESD and Latch-Up - HBMMicron · Folsom, CAFirst seen 7d ago
- Member Technical StaffPirros · Los Angeles OfficeFirst seen 32d ago
- Member Technical Staff - Applied AI Engineer (US Timing) Composio · BangaloreFirst seen 26d ago
- Senior Member TechnicalBroadridge · Bengaluru-EPIP Industrial AreaFirst seen today
- Senior Member TechnicalBroadridge · Hyderabad-Hi-Tec CityFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job