Near AI
LLM Inference Engineer
San Francisco or Remote
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 3 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Locations: San Francisco or Remote
About The Role
The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.
What You'll Be Doing
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
What We're Looking For
Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
Strong problem-solving skills and ability to communicate technical ideas clearly.
We'd Love If You Have
Experience with Trusted Execution Environments (TEE).
Active contributor to open-source LLM inference engines.
Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.
Similar jobs
- Distributed Training and Inference EngineerSciforium · San Francisco, CAFirst seen 5d ago
- INFERENCE ENGINEERMakerMaker · San FranciscoFirst seen 5d ago
- Inference Engineering, Co-opInferact · San FranciscoFirst seen 15d ago
- SRE, AI Inference EngineerFfive · 2 LocationsFirst seen 20d ago
- Distributed LLM Inference EngineerAnyscale · San FranciscoFirst seen 30d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job