Elicit
Infrastructure Engineer
Oakland, CA (or remote within US timezones)
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 10 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Elicit
Elicit is an AI research assistant that uses language models to help researchers figure out what’s true and make better decisions, starting with common research tasks like literature review.
What we're aiming for:
Elicit radically increases the amount of good reasoning in the world.
For experts, Elicit pushes the frontier forward.
For non-experts, Elicit makes good reasoning more affordable. People who don't have the tools, expertise, time, or mental energy to make well-reasoned decisions on their own can do so with Elicit.
Elicit is a scalable ML system based on human-understandable task decompositions, with supervision of process, not outcomes . This expands our collective understanding of safe AGI architectures.
Visit our Twitter to learn more about how Elicit is helping researchers and making progress on our mission.
Why we're hiring for this role
Elicit is an AI research platform used by scientists, pharma companies, and decision-makers for high-stakes evidence synthesis. A single session could trigger hundreds of thousands of language model invocations across multiple providers, which means our infrastructure decisions directly impact cost, reliability, and the quality of research outcomes for users making decisions worth millions of dollars.
Our infra is well-architected using best practices — Terraform, Kubernetes, Argo CD, GitHub Actions. But we're at an inflection point: enterprise contracts are getting larger, single-tenant deployments are multiplying, and the surface area that needs dedicated attention has outgrown what our current team can cover part-time. This is the first dedicated infrastructure hire and you'll define how this function works at Elicit.
James (Head of Engineering, ex-Square) set up the original infrastructure and will be your close partner. This role will own and evolve the infrastructure platform that underpins Elicit's product. You will ensuring it is reliable, secure, cost-efficient, and ready for the demands of a growing enterprise customer base. Under your ownership, our systems will scale gracefully across single-tenant deployments, our SLAs will be backed by real engineering rigor rather than best intentions, and our compliance posture will be a selling point rather than an afterthought.
What you'll own
Own our cloud infrastructure across AWS and GCP — Kubernetes clusters, networking, databases (Aurora PostgreSQL, Redis, MongoDB Atlas), Cloudflare, and our CI/CD pipeline.
Scale single-tenant deployments from a handful to many — each with distinct data retention, geographic, monitoring, and compliance requirements. Make a private cloud deployment a repeatable, low-overhead operation.
Build our observability and incident response practice — proactive monitoring, alerting, SLA tracking, and structured post-mortems that make the whole team better at diagnosing and resolving issues.
Drive compliance and security operations — ensure we follow through on the policies we've written (SOC 2, NIST AI framework, EU Cyber Resilience). Own disaster recovery exercises, database restoration drills, and security event monitoring (SIEM).
Manage infrastructure cost and capacity — make smart decisions about where we run workloads (AWS, CoreWeave, Parasail), optimize spend, and plan capacity as usage grows.
Improve developer experience — CI/CD pipeline performance, preview environments, local development tooling, and deployment confidence.
Contribute to backend systems where infrastructure and application intersect — circuit breakers, inference routing, data connector infrastructure for enterprise customers bringing their own data.
What success will look like (6-12 months)
A private cloud deployment is a ~1-day turnkey operation. Playbooks and templated Terraform make standing up Elicit in a customer's cloud routine, which opens up 8-figure enterprise deals.
Our observability signal:noise ratio improves 10-fold. Health monitors cover every endpoint and job, and an alert firing means something needs attention.
Disaster recovery is practiced. We run database restoration drills and provider-outage dry runs on a schedule, with post-mortems that make the whole team better at diagnosis.
Our SLAs are backed by engineering rigor. We follow through on SOC 2, NIST AI framework, and EU Cyber Resilience commitments, and enterprise security reviews go faster because of it.
Inference is faster and cheaper. You've found and executed opportunities like shifting load between providers to cut p95 latency and cost at the same time.
What we're looking for
5+ years of hands-on infrastructure/SRE/platform engineering experience.
An AI-native way of working . Agentic coding tools (Claude Code, Cursor, Devin, etc.) are how we build at Elicit, and infrastructure is no exception: agents help us write IaC and investigate incidents. You should be an enthusiastic practitioner who uses AI to multiply your impact, and excited to find new places agents can safely take on infrastructure work. Bonus: you've written about, spoken about, or built projects demonstrating this.
Solid Terraform experience . this is our primary infrastructure-as-code layer and the most important technical requirement.
Strong Kubernetes expertise . You've operated production clusters, not just deployed to them. Comfortable with EKS, networking, autoscaling (Karpenter), and debugging cluster-level issues.
AWS experience (primary), with GCP familiarity a plus.
GitOps and CI/CD fluency . Argo CD, GitHub Actions, or equivalent. You understand deployment automation, rollback strategies, and change management.
SRE mindset . You've built or significantly improved observability stacks (DataDog or equivalent), incident response processes, and on-call practices.
Security and compliance awareness. Experience with SOC 2 or similar frameworks, SIEM tooling, and translating compliance requirements into engineering practice.
Ability to write software . You can contribute to our backend codebases where infrastructure meets application logic.
Am I a good fit?
Strong applicants will find it easy to answer these questions:
Can you describe a time you designed and executed a multi-tenant or single-tenant deployment architecture for enterprise customers?
How have you approached disaster recovery planning and testing at a previous company?
Walk me through how you'd evaluate whether to build vs. buy for a new infrastructure component at a ~30-person startup.
Have you owned compliance follow-through (not just policy writing) for a framework like SOC 2?
Who will I work with?
James (Head of Engineering) : set up the original infrastructure and will be your closest partner. You'll own execution, with James as sounding board and advocate.
Panda : the engineer who has been covering infrastructure part-time, with deep context on our cluster bootstrapping, Cloudflare setup, and inference providers.
Product: PMs covering the core product, ML, and evals. They carry the customer side of enterprise deployments, so you'll work together to turn requirements like data residency, compliance commitments, and SLAs into architecture, and to weigh the cost and latency tradeoffs behind product decisions. Eval infrastructure is a shared surface with Ben, from CI integration to inference capacity.
The whole engineering team : we're ~30 people company-wide, so you'll work directly with the engineers whose developer experience you're improving, and pair with them where infrastructure meets application code.
Andreas and Jungwon (cofounders) : you'll meet both during the interview process, and infrastructure decisions with strategic weight (enterprise deployments, compliance posture) get their direct attention.
On-call & incident expectations
We don't have a formal on-call rotation yet. Incidents today are handled by the engineers closest to the affected system. Part of this role is building the incident respo
Listed on hirly, a job board. hirly is not the employer: Elicit is hiring for this role.
Similar jobs
- EDA Infrastructure EngineerUsc · Marina Del Rey, CAFirst seen today
- Infrastructure Engineer - Network Automation (On-site Dallas, TX)Truist · Dallas, TXFirst seen today
- Associate Infrastructure Engineer - Network Automation (On-site)Truist · 2 LocationsFirst seen today
- Regional Infrastructure EngineerInternationalschools · ISP US East and Canada, States of America, DallasFirst seen today
- Infrastructure Engineer – Azure Virtual Desktop / End User ComputingHcsc · IL - ChicagoFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job