ChipAgents
ML Systems Engineer
San Jose · Santa Barbara
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Stated salary
- $150,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About ChipAgents
ChipAgents is redefining the future of chip design and verification with agentic AI workflows. Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating chip development. Founded by experts in AI and semiconductor engineering, we partner with top semiconductor firms, cloud providers, and innovative startups to build intelligent AI agents. The company is a Series A company backed by tier-1 VC firms. ChipAgents is deployed in production to companies that have shipped 16B chips.
Position Overview
We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.
Key Responsibilities
Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
Implement and benchmark concrete inference optimizations.
Profile and analyze inference bottlenecks at the systems level—from GPU kernel execution to memory bandwidth constraints.
Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.
Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.
Qualifications
B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.
Strong proficiency in Python and C++/CUDA; hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.
Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.
Strong systems-level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic.
Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.
Self-directed problem solver who is interested in working on ambitious optimization challenges.
Why Join Us
Work on cutting-edge LLM inference optimization problems with real-world production impact.
Access to substantial GPU compute resources for experimentation and benchmarking.
Collaborate with a world-class team spanning AI research, systems engineering, and EDA.
Shape the performance characteristics of AI systems used by leading semiconductor companies.
What we offer
$150K/yr – $350K/yr + Offers Equity. We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.
Unlimited PTO and full benefits (medical, vision, dental, 401k).
Two engineering-centric offices with free parking, private gym, and free lunch, drinks and snacks.
Similar jobs
- Network Systems EngineerVITAS Healthcare · Jacksonville, FL, United States; St Augustine, FL, United States; Macclenny, FL, United States; Fernandina Beach, FL, United StatesFirst seen today
- Network Systems EngineerVITAS Healthcare · Miramar, FL, United StatesFirst seen today
- Site IT Systems Engineer, Ford EnergyFord Global · Glendale, KY, United StatesFirst seen today
- Systems EngineerJaxon Engineering and Maintenance · Colorado Springs, COFirst seen today
- Off-Grid Power Systems EngineerChartis Consulting Corporation · Ashburn, VAFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job