This role has closed. Clera has taken the posting down.
hirly last saw it live on 28 September 2026. See similar open roles below, or browse the live board.
Clera
ML Infrastructure Engineer
San Mateo
Similar open jobs
- Infrastructure EngineerVarick Agents · SFFirst seen today
- Critical Infrastructure Engineer - Flexential Talent CommunityFlexential Corp. · GA - Norcross; Portland, ORFirst seen today
- Data Center - Critical Infrastructure Engineer IIFlexential Corp. · VA - RichmondFirst seen today
- Oracle Cloud Infrastructure EngineerLightfeatheriollc · Washington, DC First seen todayremote
- Audio Video Infrastructure EngineerUber · San Francisco, CA, United StatesFirst seen today
- ICT Infrastructure EngineerSAIC · San Diego, CA, United StatesFirst seen todayremote
- AI Infrastructure EngineerPropio · Overland Park, KS ( Hybrid)First seen today
- Grand Bank - Infrastructure EngineerContfinco · Hattiesburg, MississippiFirst seen today
- Infrastructure EngineerWf · STERLING, VAFirst seen today
- Cloud Infrastructure EngineerQualys · RaleighFirst seen yesterday
- Infrastructure Engineer, On-Prem SystemsLumafield · San Francisco, CAFirst seen yesterday
- Hybrid Infrastructure EngineerSimspace Corporation · Boston OfficeFirst seen yesterday
- Senior Infrastructure EngineerUpstart · United States | RemoteFirst seen todayremote
- Senior Infrastructure EngineerTurnkeycareers · United States (Remote)First seen todayremote
- Cloud Infrastructure Engineer IIIJPMorgan Chase · Columbus, OH, United StatesFirst seen today
hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting.
the posting
About the Role
This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI company building a context and data governance layer that makes AI agents reliable in production. You will own the inference and model-serving infrastructure end to end, ensuring agents run fast and reliably at increasing concurrency. The work is squarely production-focused with real-world impact across regulated industries like insurance, banking, healthcare, and asset management.
What You'll Do
Design, build, and scale inference and model-serving infrastructure from the ground up through production deployment.
Optimize systems for latency, throughput, and reliability under high concurrency.
Collaborate closely with ML and infrastructure teams to ensure seamless integration and surface performance bottlenecks.
Drive solutions to infrastructure challenges across a fast-moving, cross-functional team.
What We're Looking For
5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
Hands-on experience designing and scaling inference-serving systems using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.
Strong distributed systems fundamentals, including containerization and orchestration with Docker and Kubernetes.
Proficiency with monitoring and observability tooling for production systems, such as Prometheus, Grafana, or distributed tracing frameworks.
Experience deploying and managing ML workloads on cloud platforms (AWS, GCP, or Azure).
Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.
Comfort collaborating across both ML and infrastructure disciplines in a fast-paced environment.
Nice to have: experience with knowledge graphs, semantic search, or graph databases; real-time or low-latency inference systems; agentic or multi-step AI pipelines; enterprise data integration or pipeline infrastructure.
Location
On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.