Hyperbolic
Member of Technical Staff - Inference
San Francisco, CA
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 9 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Who We Are
Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.
About the Role
We're looking for an Inference Engineer to build inference capabilities on top of Forge, our unified control plane, so customers can consume model tokens without managing GPUs and our NeoCloud partners get a full-stack path to their own token-factory offering. You'll own how models get deployed and served across clusters distributed around the world, on heterogeneous hardware.
Deployment comes first: serving models on Forge and our Kubernetes offering, evaluating inference frameworks, and standing up the monitoring, gateways, and endpoints that make a deployment production-ready. From there the work expands into optimization, autoscaling, KV-cache orchestration, and customer inference debugging. This is the primary seat for inference at Hyperbolic \u2014 you'll build it end to end, with real influence over where the scope lands.
Who You Are
Strong general inference background with a broad, high-level command of the stack rather than a narrow specialty — you can reason about the whole path from request to token
Deep Kubernetes experience, including hands-on ability to operate clusters in production, not just deploy to them
Solid grasp of the concepts that govern inference performance: TTFT, disaggregated inference, speculative decoding, and KV cache and its inner workings
Familiarity with modern inference frameworks and serving engines, and the judgment to evaluate and select among them for a given workload
Working knowledge of NVIDIA Dynamo and how it fits into a distributed serving architecture
Experience setting up monitoring, gateways, and endpoints for production inference services
Proven ability to build a product end to end — you've taken something from nothing to serving real traffic
Strong self-initiative and comfort operating as the primary owner of an area with minimal direction
Generalist instincts: you're willing to pick up adjacent work when it's what the product needs
Preferred Qualifications
Experience spanning both inference deployment and inference optimization
Hands-on model optimization work — quantization, batching strategies, kernel-level tuning, or similar
Understanding of RDMA and high-performance networking as they apply to distributed serving
Experience deploying inference across heterogeneous accelerators
Background supporting customers directly on inference debugging and performance issues
Experience at a GPU cloud, inference provider, or AI infrastructure company
Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Listed on hirly, a job board. hirly is not the employer: Hyperbolic is hiring for this role.
Similar jobs
- Sr. Member Technical Staff - ESD and Latch-Up - HBMMicron · Folsom, CAFirst seen 7d ago
- Member Technical StaffPirros · Los Angeles OfficeFirst seen 32d ago
- Member Technical Staff - Applied AI Engineer (US Timing) Composio · BangaloreFirst seen 26d ago
- Senior Member TechnicalBroadridge · Bengaluru-EPIP Industrial AreaFirst seen today
- Senior Member TechnicalBroadridge · Hyderabad-Hi-Tec CityFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job