Edison Scientific
Member of Technical Staff, Principal Infrastructure Engineer
San Francisco · New York · Boston
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Stated salary
- $200,000 – $350,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Us
Edison Scientific builds and deploys AI scientist agents to accelerate science and the development of new medicines. We are an ambitious team run by scientists and engineers from leading institutions across biology, physics, chemistry, and AI.
About the Role
As a Principal Member of Technical Staff you'll play a key role in designing, scaling, and operating the core platform infrastructure that powers autonomous scientific discovery. Your primary focus will be the orchestration for our agents at scale, building and managing clusters that orchestrate thousands of persistent, stateful workloads, developing custom resource definitions (CRDs) and operators, and ensuring the reliability and efficiency of our compute layer at scale.
Our mission is to build an AI scientist, and you'll own the infrastructure foundation it runs on. AI agents performing long-running scientific research demand resilient scheduling, lifecycle management, and resource orchestration far beyond typical cloud-native workloads. This role will influence platform architecture, establish infrastructure best practices, and partner closely with backend engineers, ML engineers, and researchers to deliver a production-grade environment that lets science move faster.
This position is part of the Platform Infrastructure team.
Key Responsibilities
Architect, implement, and operate Kubernetes clusters that support thousands of concurrent, persistent resources (agents, jobs, services) with high availability and efficient resource utilization.
Design and develop custom resource definitions (CRDs) and Kubernetes operators to model and manage domain-specific workloads such as AI agent lifecycles, research pipelines, and long-running compute tasks.
Drive the strategy for cluster scaling, node pool management, autoscaling policies, and resource quota frameworks to handle rapid workload growth.
Build and maintain infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible, version-controlled environment management.
Design and implement robust scheduling, placement, and affinity strategies to optimize cost, performance, and fault tolerance for heterogeneous workloads (CPU, GPU, memory-intensive).
Establish and uphold best practices around observability, monitoring, alerting, and incident response for infrastructure systems (Prometheus, Grafana, Datadog, or similar).
Own storage and networking strategy within Kubernetes — including persistent volume management, CSI drivers, service mesh, network policies, and ingress architecture.
Troubleshoot complex, cross-system infrastructure issues and guide others through effective debugging and remediation in distributed environments.
Collaborate closely with backend, ML, and research teams to understand workload requirements and translate them into reliable infrastructure patterns.
Required Qualifications
10+ years of professional infrastructure or platform engineering experience, with deep hands-on Kubernetes expertise in production environments.
Experience designing and implementing custom resource definitions (CRDs) and Kubernetes operators (using frameworks such as Kubebuilder, Operator SDK, or controller-runtime).
Track record of operating and scaling Kubernetes clusters supporting thousands of persistent or long-lived resources (stateful workloads, persistent pods, long-running jobs).
Deep understanding of Kubernetes internals — API server, etcd, scheduler, controller manager, kubelet — and how they behave at scale.
Expertise with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated networking, storage, and IAM primitives.
Proficiency in at least one systems or backend language for operator development and infrastructure tooling.
Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or Crossplane) and GitOps workflows.
Strong working knowledge of container networking (CNI plugins, service mesh, network policies), storage (CSI, persistent volumes, StatefulSets), and security (RBAC, Pod Security Standards, secrets management).
Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production.
Preferred Qualifications
Experience with data-intensive platforms, scientific computing, or ML/AI infrastructure.
Prior experience in startups or small teams with significant architectural ownership and ambiguity.
Experience scaling systems, teams, or platforms through periods of rapid growth.
Why join us?
We're a fast-moving, mission-driven culture where smart people do their best work and actually enjoy doing it. We also offer full-time employees the following:
Competitive salary and equity
Full healthcare coverage; we pay 100% of premiums for you and your dependents
Support for growing families, including a yearly new parent stipend and fertility coverage through Carrot
Mental health support through Rula, our in-network therapist and psychiatrist network with fast availability
12 weeks of paid parental leave for maternity, paternity, and adoption
Pet care support with a yearly employer-funded stipend for your animal companions
Commuter benefits so you can pay for transit and parking with pre-tax dollars
401(k) company matching
$300 health and wellness benefit quarterly
Lunch is on us every day you're in the office, and dinner is on us when you're working late
Regular team off-sites and company events
Edison Scientific is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.
Listed on hirly, a job board. hirly is not the employer: Edison Scientific is hiring for this role.
Similar jobs
- Sr. Member Technical Staff - ESD and Latch-Up - HBMMicron · Folsom, CAFirst seen 7d ago
- Member Technical StaffPirros · Los Angeles OfficeFirst seen 32d ago
- Member Technical Staff - Applied AI Engineer (US Timing) Composio · BangaloreFirst seen 26d ago
- Senior Member TechnicalBroadridge · Bengaluru-EPIP Industrial AreaFirst seen today
- Senior Member TechnicalBroadridge · Hyderabad-Hi-Tec CityFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job